Radar

9 public scores published

LMArena Text

Asks people to compare two anonymous model responses and vote, measuring which models users actually prefer.

Latest score: · 20 brands

LiveBench Agentic Coding

Uses coding tasks that require autonomous planning and execution to measure a model's ability to solve code problems independently.

Latest score: · 10 brands

τ³ Banking Knowledge

Uses knowledge retrieval, long conversations and tool calls in banking support to measure whether agents can complete realistic requests.

Latest score: · 8 brands

GPQA-Diamond

Uses difficult biology, physics and chemistry questions written by PhD-level experts to measure advanced scientific knowledge and reasoning.

Latest score: · 13 brands

OTIS Mock AIME 2024-2025

Uses challenging mathematics-competition problems to measure mathematical reasoning and problem solving.

Latest score: · 13 brands

AIME 2025 · OpenCompass

Uses challenging mathematics-competition problems to measure mathematical reasoning and problem solving.

Latest score: · 8 brands

SWE-bench Verified

Uses GitHub issues and tests from real open-source projects to measure whether models can understand a repository and implement correct fixes.

Latest score: · 8 brands

LiveCodeBench V6 · OpenCompass

Uses newly collected programming-contest problems to measure test-passing code generation while reducing training-data contamination.

Latest score: · 8 brands

Terminal-Bench 2.1

Uses software, system and data tasks in real terminal environments to measure agents' ability to operate computers independently.

Latest score: · 6 brands