LMArena Text
Asks people to compare two anonymous model responses and vote, measuring which models users actually prefer.
Popular public benchmark scores, all in one page.
9 public scores published
Asks people to compare two anonymous model responses and vote, measuring which models users actually prefer.
Uses coding tasks that require autonomous planning and execution to measure a model's ability to solve code problems independently.
Uses knowledge retrieval, long conversations and tool calls in banking support to measure whether agents can complete realistic requests.
Uses difficult biology, physics and chemistry questions written by PhD-level experts to measure advanced scientific knowledge and reasoning.
Uses challenging mathematics-competition problems to measure mathematical reasoning and problem solving.
Uses challenging mathematics-competition problems to measure mathematical reasoning and problem solving.
Uses GitHub issues and tests from real open-source projects to measure whether models can understand a repository and implement correct fixes.
Uses newly collected programming-contest problems to measure test-passing code generation while reducing training-data contamination.
Uses software, system and data tasks in real terminal environments to measure agents' ability to operate computers independently.