MONA explainer 11 min read

The Leaderboard Illusion: Style Bias, Vote Rigging, and the Technical Limits of LLM ELO

Visual diagram showing LLM leaderboard rankings shifting as style bias and vote manipulation are applied

ELI5

ELO ratings for LLMs rank models by aggregated human votes in head-to-head battles. The problem: voters consistently prefer longer, better-formatted answers — not more accurate ones. A few hundred injected votes can meaningfully shift top-model rankings.

The first time someone publishes a new LLM and it jumps ten positions on the Arena leaderboard overnight, the obvious explanation is that the model improved. But there’s a competing explanation, and it has the equations to back it up.

Arena (arena.ai), the platform formerly known as LMSYS Chatbot Arena, has become the closest thing the ELO Rating for LLMs community has to a ground truth for model quality. With 6.8M+ votes across 360+ models as of 2026, it carries a gravitational authority that static benchmarks like SWE Bench have largely lost — partly because it uses real humans, partly because it escapes the memorization problem that contaminates fixed datasets. But authority and accuracy are different properties. The mechanism underneath the leaderboard has structural weaknesses that deserve precise examination, because the same teams treating Arena scores as ground truth are making architectural decisions and product bets based on numbers that partially measure something other than model quality.

The Geometry of a Preference Vote

The ratings system Arena uses is often called “ELO” by journalists and researchers, but the underlying mechanism is technically the Bradley Terry Model — a pairwise comparison framework that Arena migrated to in late 2023. The distinction matters. Classic ELO Rating assumes dynamic entities: a chess player’s skill changes, so their rating should update rapidly to track improvement. LLMs deployed behind an API have static skill after training ends. Applying a dynamic update model to a fixed-skill entity violates one of the system’s core design assumptions, as Boubdir et al. documented in a 2023 study finding that reliability and transitivity axioms are not always satisfied in LLM Elo computations.

The Bradley-Terry model estimates the probability that model A beats model B in a pairwise comparison. Votes accumulate, the model’s parameters update, and a stable ranking emerges over thousands of comparisons. The mathematical intuition is clean: more votes, more signal, more reliable rank. What the equation cannot see is what the voter was responding to.

What Are the Limitations of Using ELO Ratings to Rank LLMs?

Three structural limits compound each other.

The first is style confounding. In August 2024, the LMSYS research team published a controlled analysis of what drives Arena votes. The finding was specific: the response length coefficient was 0.249 — meaning longer responses systematically receive higher preference ratings, independent of their information content. Markdown headers, lists, and bold formatting contributed additional but smaller effects (LMSYS Blog). When they applied statistical style controls to the leaderboard, the rankings shifted measurably. GPT-4o-mini dropped from rank 6 to rank 11. Grok-2-mini fell from rank 6 to rank 18. Claude 3.5 Sonnet, whose outputs tend toward precision over decoration, climbed to tie for first in the Hard Prompts category it had previously ranked second in (LMSYS Blog).

This is not a calibration error. It is a structural fact about the data-generating process. Voters are not evaluating a quality metric; they are expressing a preference, and preferences systematically correlate with presentation cues that are orthogonal to reasoning quality. The Arena score is a preference signal, not a capability measurement — and the gap between those two things is widest precisely where it matters most, in hard technical domains where verbose-but-wrong is easy to prefer over terse-but-correct.

The second limit is non-transitivity. Pairwise comparison assumes that preferences form a consistent ordering: if voters prefer A over B, and B over C, they should prefer A over C. Real voters don’t behave this way. Xu et al. (ICML 2025 Spotlight) found that applying round-robin tournament structures with Bradley-Terry fitting improved Spearman rank correlation with Arena from 95.0% to 96.4% over baseline pairwise methods — a modest but real gap attributable to the inconsistencies in human preference cycles. Research into LLM as a Judge systems reveals the same pattern: Qwen2.5-Max produced preference judgments that contained logical inconsistencies in 67.96% of cases (Yu et al.). That figure is specific to Qwen2.5-Max as an evaluator and should not be generalized as a universal rate — but it illustrates how deeply non-transitivity can run in preference-based ranking systems.

The third limit is the static-dynamic mismatch identified by Boubdir et al. One practical consequence appeared in March 2026: Arena methodology updates shifted some Elo score distributions by 20 to 40 points without any underlying change in model quality (Maginative). Any historical ELO comparisons made against pre-March 2026 Arena scores are not directly comparable to post-update values — a gap that compounds the existing interpretability problems.

The Architecture of Bias in a Vote

Before the style control analysis existed, the working theory was that Arena votes were noisy but unbiased. Noise averages out over millions of votes. Bias doesn’t.

The difference is systematic direction. If length preference is consistent across voters — which the 0.249 coefficient suggests it is — then aggregating more votes doesn’t reduce the error; it hardens it. A model optimized to produce longer outputs would climb the leaderboard by exploiting a preference signal rather than improving on any capability axis the Rubric Design community would recognize as meaningful.

The style-controlled leaderboard, introduced as a separate view in November 2024, attempts to partial out these effects statistically. But it exists as an alternate view, not the default headline ranking — meaning the number most teams cite when comparing models is the one that includes the style premium.

The Social Engineering Layer

Style bias operates passively — voters don’t know they’re rewarding formatting. Vote injection and preferential data access are active vulnerabilities, and both exploit the same underlying structure: the Bradley-Terry model’s sensitivity to the full graph of pairwise comparisons.

Can Chatbot Arena ELO Rankings Be Gamed or Rigged?

The answer is yes, and the mechanism is more subtle than it sounds. Min et al. (accepted ICML 2025) demonstrated that hundreds of injected votes — against a platform running over 1.7 million total votes at the time — can meaningfully shift top model rankings. The magnitude depends on attack strategy and target position, and the paper does not provide a single clean threshold that applies universally.

The more counterintuitive finding is the attack vector. The “omnipresent strategy” Min et al. document does not require the attacker to vote directly in battles involving the target model. Because the Bradley-Terry model estimates ratings from the full graph of pairwise comparisons, votes in battles between other models can indirectly alter the rating of the target. The ranking is a function of the entire comparison graph, not just a model’s direct win-loss record.

A separate structural vulnerability involves preferential data access. Singh et al. (NeurIPS 2025 Datasets and Benchmarks Track) documented a striking asymmetry: Google had participated in 19.2% of all Arena battles, OpenAI in 20.4%, while 83 combined open-weight models shared only 29.7% of total Arena data. Meta tested 27 private LLM variants through the platform before the Llama-4 public release, per Singh et al. The authors named this paper “The Leaderboard Illusion.” Their claim was that large labs with early access to Arena-distribution data — the specific prompts and comparison patterns that Arena uses — could optimize against that distribution and achieve a form of benchmark contamination within the live evaluation system.

Diagram showing how style bias and data access asymmetry distort LLM ELO rankings on Chatbot Arena
Two pathways to leaderboard distortion: voter style preferences and differential data access among model providers.

Arena’s rebuttal is worth reading alongside the critique. The Arena Blog response disputes the characterization of open-weight model representation (citing 40.9% of the leaderboard, not 8.8% as claimed), and estimates the pre-release testing boost at approximately +11 Elo points after 50 tests and 3,000 votes — not the hundreds of points the “Leaderboard Illusion” framing implies. The debate is live and unresolved; both papers are making methodological arguments about the same system from different vantage points.

Arena’s testing and privacy policy has been publicly available since March 1, 2024 (Arena Blog) — so the pre-release testing access is disclosed, even if the distributional effects of that access are disputed.

When the Rating Breaks Down

The implications of these structural limits follow a consistent pattern: the ranking is most informative where the biases matter least, and least informative where the stakes are highest.

For conversational general tasks — the kind where longer, more formatted responses genuinely do reflect higher quality — Arena ELO tracks something real. A model that is better at following complex instructions, reasoning through multi-step problems, and generating coherent long-form text probably does deserve the length premium.

For narrow technical evaluation — coding correctness, mathematical reasoning, factual precision — the style correlation is actively misleading. A model that outputs an incorrect proof in well-formatted markdown will consistently outperform a model that outputs the correct answer in plain text, at least until the voter understands the mathematics involved.

The Human Evaluation for AI literature offers a partial fix: Cohen's Kappa and inter-annotator agreement measures can quantify how much of the variance in votes is attributable to genuine quality assessment versus idiosyncratic presentation preferences. The Label Studio and annotation community has developed rubric design frameworks precisely to constrain what annotators evaluate. Arena’s open-ended “which is better” question is, by design, maximally permissive — and therefore maximally susceptible to the style confound.

If you are building systems that rely on Arena rankings to make architectural choices, the practical prediction is this: a model’s Arena rank tells you its performance in the human-preference-at-scale distribution, which correlates with general assistant quality and partially decorrelates from narrow technical capabilities. If your production task is a coding assistant, a math solver, or a factual retrieval system, the rank deserves a discount — and you should weight it against task-specific benchmarks where outputs can be verified against ground truth.

When it breaks: ELO-based leaderboards produce rankings that are unreliable for any task where response quality is objectively verifiable but voters lack the domain knowledge to verify it — specifically, technical domains where superficial formatting quality and actual reasoning quality diverge.

The Data Says

The Arena ELO system measures human preference at scale, and preference is not the same as capability. A length coefficient of 0.249 means the ranking system has a measurable thumb on the scale in favor of verbosity. Vote rigging is feasible at surprisingly low injection volumes, and pre-release data access creates distributional advantages the rating system cannot directly detect. The style-controlled leaderboard is a real improvement — but it exists as an alternate view, not the default number, which means most comparative citations are citing the biased version.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors