Scoring turns a benchmark from “eyeball the JSON” into an objective number. Each model’s result is graded against a gold reference — the expected output — producing completeness, correctness, and an overall quality score you can sort on.
Scoring needs something to score against. Each scenario carries a reference output: the correct answer for its one fixed entity. Build it by generating with strong models (web search + a source-of-truth document), by pasting a known-good result, then editing it by hand — and mark it verified once you trust it. A verified reference is required to benchmark the scenario at all, so there’s always something to grade against. If you later edit the reference — or change the scenario’s scoring config — existing scores are flagged stale until you re-score.
The reference also updates itself. A verified reference is still a hand-corrected draft, and the models you benchmark are its reviewers: wherever the judge finds a candidate’s answer better than the reference’s (or the reference’s wrong), wherever the scenario’s own samples prove the reference wrong, and wherever a model fills a value the reference left empty that the judge confirms, the result records a finding. At the end of every scoring pass the findings of all scored models are folded — what the samples prove first, then the answer most models gave — and written into the reference. An edit a pass already made is only replaced by stronger evidence (the samples, a verdict that the value is wrong, or more agreeing models — counted across passes), never by one more model’s opinion, so benchmarking models one at a time cannot drift the reference toward the last one scored. These automatic edits never mark scores stale (only your own saves do), and every one of them is logged with what it replaced and which models raised it.
Review them as track changes: the scenario’s Reference view, next to its results, shows the reference itself with every automatic edit highlighted where it sits — the previous value struck through, the new one, how many models stand behind it, and why. Everything is accepted unless you reject it; a rejection restores the previous value and keeps that path out of future passes. Filter by attribute or evidence (one-model edits are the ones worth a look), select what is shown, and reject in bulk.
The core problem: two correct answers can be written differently. A model that names an actor “R. Downey Jr.” instead of “Robert Downey Jr.” isn’t wrong. So each field is compared with a tiered ladder — cheapest and most certain first, escalating only when needed:
Identical values match. So do values that differ only in case, surrounding whitespace, or numeric precision ("Acme" = "ACME", 4.0 = 4). Free and fully deterministic.
For text, the candidate and reference are embedded and compared by cosine similarity. Above the threshold they count as the same — so a valid alternate spelling like "R. Downey Jr." vs "Robert Downey Jr." is a match, not an error. Dates are the exception: they are compared as calendar values, never by similarity, so a near-but-wrong date ("1972-03-14" vs "1972-03-24") is a clean mismatch rather than a deceptively high cosine. Booleans are likewise exact-or-nothing.
Values too close to call by similarity — all free-text fields like summaries and descriptions, every non-identical number, and a clearly different value that your documents or most other models back — are sent to a judge model. The judge is blind: it sees the two values as A and B, with the field’s place in the schema (its parents and their descriptions, its type, which list item it belongs to) and your source documents when the scenario has any, and says which serves the field better, that both are correct, or that one is wrong. A candidate judged equivalent or better than the reference — or a reference judged wrong — earns full credit; a weaker answer earns partial credit, a wrong one little or none. A number gets partial credit when the field tolerates it (a molecular weight of 273.37 vs 273.35, a half-life of 12 vs 15) and fails where exactness matters (a release year of 2020 vs 2023). A value the reference left empty is asked about on its own: confirmed correct, it counts as a value the reference should have held and the model found — the score goes up, not merely unpunished — and the reference adopts it.
A strictness setting controls the embedding threshold: higher means two differently-written values must be more similar to count as the same. The strictness, the optional judge model, and the embedding model are all set on the scenario — not chosen each time you score — so every model is graded identically and scores stay comparable.
Lists — a film’s cast, a drug’s side effects — are where models differ most: a small model might find 4 actors where a strong one finds 15. Order doesn’t matter, and finding more correct items should win. So arrays are scored as a set, not position by position:
Expand a result row to see exactly which items were matched, missed, or hallucinated.
A single number hides too much, so every result carries sub-scores:
The expandable row shows the per-field breakdown: candidate vs reference, which rung of the ladder decided it, and the similarity where relevant.
Quality is only one third of the story. Where a benchmark feeds model selection — as a scoring source — the model’s standing comes from blending quality, speed and cost, and that ratio is yours: set it per scenario type under Settings → Organization → Defaults, each summing to 100 and split evenly by default. Weight cost heavily and a cheap, decent model will outrank an excellent expensive one — which is the right answer for some workloads and the wrong one for others, so the platform declines to decide it for you. Speed and cost are read against the scenario’s other results on a log scale: the fastest or cheapest model scores 100 and one ten times slower or dearer scores 0, but a tight field is never stretched to fill the range — two models at $10 and $12 land at 100 and 92, not 100 and 0.
When a scenario runs a model more than once (repetitions), each run is scored on its own and the row shows the mean quality plus a consistency spread (lowest–highest of the runs) — so a model that is right on average but erratic is easy to spot. The visible output is the median-by-quality run.
Benchmarks aren’t limited to enrichment: a scenario can also test sample generation (each model invents an example JSON for the same free-text request) or schema generation (each model converts a fixed sample into a schema). Each has its own scoring rules:
Either way the familiar columns keep their meaning — hover a column header for the type-specific definition, and expand a row for the full breakdown.
Scoring is a separate pass over already-saved results — it never re-enriches, so it never re-pays for the models under test. It does embed text to compare values (and run the judge, if the scenario has one), which deducts credits based on usage. This happens automatically during every run — each model is scored as soon as its runs finish — and again whenever you re-score. If your organization has no embedding model configured (and the scenario sets no override), scoring still runs but falls back to exact matching only (alternate spellings then count as mismatches), and says so. A broken judge is different: if a judge call fails, scoring stops with an explicit error and the affected model keeps no partial scores — a result is either fully scored or not scored at all, and stays re-scorable later. Judge answers are cached per scenario by their content: the same question raised by another model, repetition or pass is never paid twice, and a re-score after the reference updated itself re-asks only the questions at the edited fields. Nothing the judge does is hidden: every scoring pass over a model leaves a scoring record in History — one row per judge call with its prompt, answer, tokens and cost, failed calls included, plus how many questions the cache answered for free — and the results table shows each model’s judge cost, call count and scoring time next to a link to it. When picking the judge, note that LLM judges can favor their own model family — prefer a judge from a provider you are not benchmarking.
In Model Management → Benchmarks, set and verify a reference in the scenario editor (and pick its judge model, embedding model, and strictness there). From then on, every run auto-scores its successful results — a sortable Quality column fills in with no extra step. Use Re-score results (the header button or the ··· menu) to re-grade after you edit the reference or the scoring config. The reference updates badge next to the reference status opens the log of what the scoring passes changed, with a revert per entry.