MONA explainer 11 min read

SWE-bench Verified, Lite, Full, and Multimodal: Choosing the Right Benchmark Split

MONA studying five overlapping code evaluation grids, each labeled with a different SWE-bench split variant

ELI5

SWE-bench converts real GitHub pull requests into AI coding tests: given a repository and a bug report, the model must produce a patch. Five official splits — Full, Lite, Verified, Multimodal, Pro — trade off cost, accuracy, and language coverage.

In February 2026, OpenAI deprecated a benchmark it had created. Not a legacy standard from before the frontier-model era. The benchmark that OpenAI had built eighteen months earlier to fix the original SWE-bench’s reliability problems.

The Preparedness team had spent that intervening time doing what you do when a measurement tool concerns you: auditing it. What they found concerned them enough to retire it.

A benchmark built to repair a benchmark, retired for the same class of reasons. That recursion is not ironic; it is precisely how measurement science is supposed to work. The question it leaves is practical: which split are you actually running, and what does its score tell you?

How SWE-bench Builds Its Test Cases

Every task in SWE-bench is sourced from a real, closed GitHub pull request — specifically, one that resolves an issue and includes at least one modified test file. This construction distinguishes the benchmark from collections of synthetic problems or manually authored challenges. The difficulty is not calibrated by a committee; it is inherited from real software maintainers who found a problem hard enough to open an issue about.

What do you need to know about pull requests and unit tests to understand SWE-bench?

A SWE-bench task instance has a fixed structure. The model receives the full repository at the commit immediately before the pull request was merged, plus the plain-text description of the GitHub issue the PR was created to resolve. Its job is to produce a valid git patch — a diff that transforms the pre-PR repository into something that closes the issue.

Grading is automated and deterministic. The patch runs inside an isolated Docker container against two categories of unit tests:

  • FAIL_TO_PASS: tests that existed in the PR but were failing before the fix. A correct patch makes these pass.
  • PASS_TO_PASS: tests that were already passing before the PR. A correct patch must not break them.

This construction means Rubric Design never enters the evaluation pipeline the way it does for prose or reasoning benchmarks. There is no LLM as a Judge scoring output against subjective criteria, and unlike ELO Rating for LLMs systems where models are ranked through pairwise comparisons, SWE-bench produces an absolute score: the proportion of tasks where the generated patch passes all graded tests.

Not judgment. Evidence from an automated harness.

The unit test is the grader. The benchmark never asks whether the model’s code resembles the human’s original patch; it asks whether the code makes the tests go green. A solution that differs completely from the original still counts as correct, provided the tests pass. That grading choice is the elegant part and the fragile part at once — it inherits a quiet assumption: that the tests are a fair description of the task. When they are not, the grade lies.

The source repositories in SWE-bench Full are exclusively Python. The twelve repositories were selected for maturity and test-coverage density; django contributes 850 of the 2,294 task instances, sympy 386, and scikit-learn 229 (SWE-bench paper). Tasks lean toward medium complexity — not trivial one-liners, but not sweeping architectural refactors either.

The Five Official Splits

Running 2,294 tasks in isolated Docker containers — each requiring full repository checkout and test harness execution — is expensive. The official splits exist for different reasons: some to reduce cost, some to improve signal quality, some to extend domain coverage beyond Python.

What is the difference between SWE-bench Verified, Lite, Full, and Multimodal?

SplitSizeLanguageValidationCurrent Status
Full2,294PythonAutomatedActive (2023–)
Lite300PythonAutomated, curatedActive (2023–)
Verified500PythonHuman-annotatedDeprecated Feb 2026
Multimodal619JavaScript, TypeScriptAutomated + visualActive (ICLR 2025)
Pro1,865Python, Go, TS, JSAutomated + held-outEmerging standard (2025–)

SWE-bench Full (2,294 instances) is the original dataset, constructed by Princeton NLP from twelve Python repositories and published as an ICLR 2024 oral (SWE-bench paper). Most expensive to run; the comprehensive baseline for Python software engineering tasks.

SWE-bench Lite (300 instances) is a curated subset of Full, favoring self-contained functional bugs — tasks where the fix doesn’t depend on external systems or wide repository context not captured in the issue text (SWE-bench Docs). Lite is the standard for preliminary comparisons when infrastructure cost constrains evaluation.

SWE-bench Verified (500 instances) was a human-validated subset of Full, intended to remove the annotation quality problems discovered in the original set. Deprecated in February 2026.

SWE-bench Multimodal (619 instances) is a separate dataset, not a subset of Full. Published at ICLR 2025, it covers 17 JavaScript and TypeScript repositories — user-facing applications including UI design systems, web apps, and syntax highlighters (SWE-bench M paper). The dataset contains 517 test instances across 12 repositories and 102 development instances across 5 repositories. Every task includes at least one image in its problem statement or tests, compared to only 5.6% of original SWE-bench tasks. It is the only split suitable for evaluating visual understanding alongside code correctness.

SWE-bench Pro (1,865 instances) was created by Scale AI and announced September 19, 2025 (Scale AI Blog). It covers Python, Go, TypeScript, and JavaScript across 41 repositories, includes a private held-out test set designed for contamination resistance, and is intended as the replacement for Verified. Top scores on Pro’s public set clustered near 23% in the February 2026 independent evaluation, compared to the 70%+ scores multiple systems achieved on Verified at the same time (Simon Willison) — a gap that reflects harder tasks, reduced contamination, and the private held-out set obscuring the test boundary.

What is SWE-bench Verified and why did OpenAI create it?

Verified emerged from a systematic quality audit of Full. The Human Evaluation for AI process involved annotators reviewing each original task and applying two binary judgments: whether the problem statement was sufficiently specified to resolve without repository-level context unavailable in the issue text, and whether the unit tests were fair — meaning a functionally correct solution would pass them, and a functionally incorrect one would not.

The annotation workflow drew on structured tooling comparable to what Label Studio pipelines enable for raw data labeling. Cohen's Kappa agreement between annotators was measured to verify the flagging was consistent across reviewers rather than dependent on a single annotator’s judgment.

Consistent. Not necessarily accurate.

The statistics from that audit were striking. Annotators flagged 38.3% of samples for underspecified problem statements and 61.1% for unfair unit tests; combined, 68.3% of original Full samples were removed, leaving 500 validated instances (OpenAI Blog). Verified was released on August 13, 2024, and became the de-facto standard for frontier coding capability comparisons.

The reason so many original tasks were problematic is structural, not accidental. Real developer code doesn’t always separate “the fix” from “the surrounding context” in a way that maps cleanly onto an isolated benchmark task. A unit test can pass due to a property of the implementation unrelated to the intended fix. An issue description can be clear to a developer with repository familiarity but ambiguous to a model without it. These problems don’t surface until you test explicitly for them — which is precisely what the Verified annotation process did.

Then, less than two years after Verified’s release, OpenAI deprecated it.

The stated reasons involved two compounding problems: training data contamination — models’ training corpora may include SWE-bench tasks or close derivatives, making high scores reflect memorization rather than general coding capability — and flawed evaluations, with approximately 39.4% of audited Verified problems found to be rejecting functionally correct submissions (OpenAI Blog). The benchmark that fixed the benchmark had the same flaw.

Benchmark compatibility notes:

  • SWE-bench Verified (deprecated February 2026): OpenAI deprecated this split citing contamination and flawed evaluations. Scores reported against Verified are not reliably comparable across time or model versions. The recommended replacement is SWE-bench Pro.
  • Epoch AI evaluation infrastructure: Epoch AI upgraded its SWE-bench evaluation harness to v2.0.0 in February 2026. Scores produced before this upgrade are not directly comparable to post-upgrade scores (Epoch AI leaderboard).
Diagram showing five SWE-bench splits arranged by size and validation rigor, with Verified marked deprecated and Pro emerging as replacement
From Full's 2,294 Python tasks to Pro's multilingual held-out corpus — each split measures a different scope, under different contamination risk.

What Choosing the Wrong Split Costs You

The structure of SWE-bench — unit tests as automated graders, real PRs as task sources — generates specific, predictable failure modes. Understanding them changes how you read a leaderboard number.

If the unit tests in a given split are over-constrained or poorly specified — a known limit of automated evaluation metrics — a model that produces a genuinely correct fix will still fail. This was the finding that motivated Verified; the recursive finding that motivated its deprecation. Benchmark quality requires auditing the auditor.

If a model’s training data includes benchmark tasks — directly, or through repositories that have since entered pretraining corpora — its score measures exposure rather than capability. Scores above 70% on Verified warrant this reservation; no externally verifiable method exists to separate genuine generalization from contamination at that level.

If your evaluation requires coverage beyond Python, Lite and Full provide no signal. Multimodal extends to JavaScript and TypeScript but adds a visual requirement that purely text-based models cannot satisfy. Pro adds Go and TypeScript without the visual constraint, with its private held-out set limiting the contamination surface.

If you are comparing results across time, verify the evaluation infrastructure version. The February 2026 Epoch AI upgrade to v2.0.0 means pre- and post-upgrade scores on the same model, same split, are not directly comparable — the number can change for reasons unrelated to the model itself.

Rule of thumb: A leaderboard number without a named split, an infrastructure version, and a date is not a measurement; it is an assertion.

When it breaks: SWE-bench cannot distinguish between a model that understood the problem and a model that memorized the answer. When training data contamination is present, the benchmark grades familiarity, and the score looks identical to genuine capability from the outside. The execution-based grading also narrows the measurable world to bugs a unit test can express; engineering problems that resist test specification never enter the benchmark at all.

The Data Says

SWE-bench Verified moved from benchmark to deprecated standard in less than two years — a recursion that reveals something important about measurement: every ruler eventually needs its own ruler. As of February 2026, SWE-bench Pro, with its multilingual scope, private held-out set, and top public-set scores near 23%, is the contamination-aware option for research comparisons. Full and Lite remain stable for reproducibility on established Python baselines. The split you choose determines what you are actually measuring; treating all leaderboard numbers as interchangeable is not rigor, it is misattribution.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors