MONA explainer 13 min read

Inter-Annotator Agreement, Rubric Design, and Calibration in Human AI Evaluation

Abstract diagram showing annotator agreement scores converging around a rubric scoring matrix

ELI5

Inter-annotator agreement (IAA) measures how consistently different raters score the same AI output. Without it, disagreement could reflect rater subjectivity, not model quality. Cohen’s kappa corrects for chance; rubric anchors and calibration keep raters aligned.

Two annotators sit down with the same LLM output. One rates it a 4 out of 5 for factual accuracy. The other gives it a 2. Both are following the same rubric. Both are experienced evaluators.

This is not an anomaly. It is the default state of any human evaluation study that skips the infrastructure phase. The disagreement does not disappear when you average the scores — it migrates into your conclusions, quietly corrupting every comparison you draw between models.

The machinery that prevents this has three components: a statistical agreement metric, a carefully engineered rubric, and a calibration protocol. Each one fails in a specific, predictable way when misapplied. Understanding those failure modes is how you build an evaluation study that tells you something true.

The Agreement Problem: Why Percentage Alone Misleads

Human Evaluation for AI is precise in theory and messy in practice. The practical problem starts the moment you try to measure how much your raters agree.

The naive approach is raw percentage agreement: count the cases where all raters gave the same score, divide by total cases, report the number. A 90% agreement sounds rigorous. It is often meaningless.

Consider a safety classification task where 95% of responses are labeled “safe.” Two raters who both blindly label everything “safe” achieve 95% agreement — without reading a single response. The percentage captures their shared laziness, not their shared judgment.

This is why agreement metrics that correct for chance are the standard in serious annotation work.

What is inter-annotator agreement and how is Cohen’s kappa calculated?

Inter Annotator Agreement comes in several forms, but Cohen's Kappa remains the most widely cited for two-rater categorical tasks. The formula is:

κ = (P_a − P_e) / (1 − P_e)

where P_a is the observed proportion of agreement and P_e is the expected agreement by chance. The denominator normalizes against the theoretical ceiling — perfect agreement beyond chance — so κ = 1 means perfect agreement, κ = 0 means raters perform no better than chance, and negative values mean systematic disagreement (raters are actively anti-correlated).

The Landis & Koch thresholds (1977) give the interpretive scale: below 0.00 is poor, 0.00–0.20 slight, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and 0.81–1.00 almost perfect. A κ ≥ 0.8 is widely cited as the practical target for high-quality annotation work (Claru AI Glossary). Treat these as guidelines, not hard rules — several alternative interpretive scales exist, and the thresholds were not derived empirically for any specific domain.

One important failure mode: the kappa paradox. When label distributions are highly imbalanced, the expected-by-chance term P_e can be so large that two raters who agree on nearly all cases still produce a low κ. The statistic is technically correct — it tells you agreement provides little information beyond label prevalence — but it reads as failure when the evaluation is sound. This matters in LLM safety evaluation, where most outputs clear a baseline bar and genuine disagreements cluster in a small, important fraction.

Cohen’s kappa handles exactly two raters. Scale to three or more raters using Fleiss’ kappa, an extension for fixed-panel categorical annotation. For the more common case in NLP evaluation — multiple raters, ordinal scales, and missing ratings — a different tool is more appropriate.

Choosing the Right Metric for Ordinal and Multi-Rater Scenarios

The choice between kappa variants and Krippendorff’s alpha is not stylistic. It is structural, and the wrong choice produces numbers that do not say what you think they say.

Krippendorff’s alpha is defined as:

α = 1 − (D_o / D_e)

where D_o is observed disagreement and D_e is expected disagreement. The difference from kappa is in how disagreement is measured: alpha uses a distance metric appropriate to the measurement scale. For categorical data, any disagreement is distance 1. For ordinal data — like a 5-point quality scale — the distance between adjacent scores is smaller than the distance between extremes. Raters who score 3 versus 4 on a fluency scale are closer than raters who score 1 versus 5, and alpha encodes that (Label Studio Blog).

This matters because LLM evaluation rubrics are almost always ordinal. Treating a 1-vs-5 disagreement identically to a 3-vs-4 disagreement inflates observed disagreement and systematically underestimates alignment between raters who roughly agree.

Krippendorff’s alpha also handles missing ratings — cases where not every rater annotated every item — without requiring you to discard incomplete rows. For large-scale LLM evaluation batches where raters work in parallel on different subsets, this is practically essential. The interpretive thresholds are: α < 0.667 indicates unreliable annotation, 0.667–0.799 allows tentative conclusions, and α ≥ 0.800 indicates reliable, definitive results (Encord Blog). As with kappa thresholds, these originate from Krippendorff’s own writing rather than domain-specific empirical derivation.

A 2026 survey of NLP annotation methodology (“Counting on Consensus”) recommends Krippendorff’s alpha with an ordinal distance metric as the default for multi-rater NLP evaluation. Human agreement between two annotators using the same evaluation method on the same material typically reaches roughly 80%, per the same source — a ceiling that is informative when you are also benchmarking automated judges against human judgments.

Diagram comparing Cohen's kappa formula versus Krippendorff's alpha, showing scale types and rater count requirements for each metric
Choosing the right agreement metric depends on the number of raters, the measurement scale, and whether missing ratings need to be accommodated.

Building a Rubric That Raters Can Actually Agree On

A rubric is not a description of what good output looks like. It is a decision procedure. The distinction matters because descriptions leave interpretation gaps that expand into disagreement under load.

Rubric Design divides into two structural approaches: holistic and analytic. A holistic rubric asks raters to produce a single overall quality score. An analytic rubric decomposes quality into criteria — factual accuracy, coherence, tone — scored independently. Analytic rubrics consistently produce higher inter-rater reliability because they force raters to separate concerns that human judgment tends to conflate. A response that is technically accurate but poorly written receives different scores on different dimensions, rather than a confused holistic average (Autorubric, 2026).

How do you design a scoring rubric for evaluating LLM responses?

The scale width is a more consequential design decision than most practitioners expect. The instinct is to use a 10-point or 7-point scale for “granularity.” In practice, rater reliability degrades as the scale widens, because the behavioral distinctions between adjacent points become ambiguous. Scales of 3–5 levels with explicit behavioral anchors — specific descriptions of what a response at each level must do — significantly outperform wider scales on inter-rater agreement. LLM as a Judge systems show a related pathology: LLM judges applying broad numeric scales tend toward the middle ratings (central tendency bias), compressing the distribution and reducing the signal in the scores (Autorubric, 2026).

The most reliable criterion type is binary: MET or UNMET. Binary criteria yield the highest inter-rater agreement because they force a single categorical decision rather than a relative placement. Not every criterion can be binary — some genuinely require gradient measurement — but where the underlying question is “did the response do X?”, a binary criterion is strictly preferable to a Likert-style one.

Behavioral anchors are the single most underinvested part of rubric construction. A criterion labeled “factual accuracy” with a description “responses should be factually accurate” is functionally useless. A criterion that defines Level 3 as “all factual claims are verifiable against the provided context, no claims contradict the context, no claims extend beyond what the context supports” is actionable. The difference between these is not the label but the decision boundary.

Well-anchored rubrics also reduce the overhead of rater training: a new rater who encounters an ambiguous case can consult the anchor rather than asking a human. At scale, this matters.

What Rater Calibration Actually Fixes

Rubric design handles the structural sources of disagreement. Calibration handles the interpretive ones.

Even the best rubric leaves room for individual rater tendencies. Some raters are lenient by default. Others are strict. Some drift over a long annotation session, applying increasingly consistent interpretations that diverge from the original rubric intent. These are not character flaws — they are predictable features of human cognitive labor. Calibration protocols counteract them.

What is rater calibration and why does it matter in human evaluation?

The standard calibration procedure begins before annotation starts. Raters work through a set of gold-standard items — responses with established correct labels, ideally determined by expert consensus or domain specialists — in a structured session. They discuss disagreements explicitly, with reference to the rubric criteria. The goal is not to achieve agreement by social pressure but to identify ambiguous rubric language and resolve it.

The calibration session surfaces cases where two raters apply different rules to the same situation while both following the rubric as written. These cases indicate anchor gaps. Closing the gaps before live annotation prevents systematic divergence at scale.

The calibration work does not end at the pre-annotation session. Inserting gold-standard items periodically into live annotation batches — without rater awareness — allows continuous monitoring of individual drift (Autorubric, 2026). A rater whose accuracy on gold items declines is recalibrated before the drift propagates across a large batch.

5-shot calibration — showing raters five pre-labeled examples before each session — achieved 80% accuracy on a chemistry grading benchmark in one 2026 study (Autorubric, 2026). Iterative rubric refinement, where the rubric itself is updated based on calibration findings and rater disagreement patterns, reached 89% human-level agreement in a separate case study from the same research. The mechanism in both cases is the same: reducing the interpretive gap between what the rubric says and what raters infer from it.

Calibration for automated judges follows similar logic. An LLM-as-a-judge configuration that includes few-shot exemplars with explicit scoring rationales performs more consistently than one calibrated only with a scoring rubric. The few-shot examples serve as behavioral anchors for the model’s probability distribution, not just instructions.

What Agreement Numbers Predict About Your Evaluation

The agreement statistics from your calibration phase are diagnostic, not decorative. They tell you which parts of your evaluation are trustworthy and which conclusions you should not draw.

When κ or α falls below the reliability threshold, the annotation task itself is underspecified — not the raters. The measurement tool is not measuring a coherent construct. Any model comparison built on that data inherits the underspecification; the ranking it produces reflects rater variance more than model quality.

The practical consequence flows in two directions. First: if you are comparing two models on a task where your raters show poor agreement, your comparison is noise. The model ranked higher may have genuinely higher quality, or it may have just happened to match the idiosyncrasies of whichever raters annotated its outputs. Second: tasks where humans reliably disagree are interesting in their own right. They identify dimensions of LLM output quality where the evaluation question itself is not well-formed.

Reporting practice matters here. The 2026 LREC guidance (“Counting on Consensus”) recommends including confidence intervals alongside agreement statistics and analyzing disagreement patterns explicitly — not just the summary statistic but where disagreement concentrates. Whether disagreements cluster in a specific output category, a specific rater pairing, or specific rubric criteria changes the corrective action.

The relationship between human evaluation for AI and automated evaluation is also clarified by agreement data. LLM-as-a-judge systems like those evaluated on SWE Bench are typically validated against human judgments by computing agreement between the automated judge and a human panel. An agreement statistic in that comparison tells you something specific: how closely the automated judge approximates human consensus, and under what conditions it diverges. Without the agreement measurement, the validation claim is undefined.

Rule of thumb: any evaluation study that reports model rankings without reporting IAA statistics is incomplete. The ranking is a function of the measurement — and the measurement’s reliability is the number you need to trust the ranking.

When it breaks: IAA statistics assume raters are making independent judgments. When raters work in teams that discuss cases before annotating, or when gold-standard items are visible during annotation rather than blinded, agreement inflates artificially — the statistic measures social alignment, not measurement reliability.

The Data Says

Agreement measurement, rubric design, and calibration are not procedural add-ons to a human evaluation study — they are the study’s validity argument. A κ in the “poor” range tells you the instrument is not measuring a coherent construct. A rubric with no behavioral anchors tells you raters are applying different decision procedures while believing they are applying the same one. A calibration protocol that ends at the pre-annotation session tells you drift is uncontrolled. Each failure mode produces evaluation data that looks precise while being structurally unreliable — the most dangerous kind of measurement error, because it is invisible in the outputs.

AI-assisted content, human-reviewed. Images AI-generated. Editorial Standards · Our Editors