Methodology

How assert(llm) decides what counts as correct

Every number on this site comes from a rule you can read. This page is that rule set: what is in each test set, how an answer is scored, what the confidence interval means, and what this method cannot tell you.

v1.0 · October 2026 · scoring engine scoring.js

1. Datasets

Five sets ship with the tool. Every item carries an expected answer written before any model saw it, and the type of that expected answer decides which grader runs. All five are at version 1.0.

SetItemsTypeGraderSource
ISTQB Foundation15mcq deterministic ISTQB-style questions written against Foundation-level syllabus topics. Not the official exam; see limitations.
Python coding20code deterministic Function tasks written for this tool, each with a required signature and a list of forbidden constructs.
SQL fundamentals25code deterministic Query tasks written for this tool, scored on required clauses rather than on execution.
General knowledge5short deterministic Facts with a single defensible answer, plus accepted aliases.
Hard problems5short ×3, code ×2 deterministic Maths and algorithm problems whose answers were each verified two independent ways before being pinned.

You can also write or paste your own prompts in the browser. They are scored by exactly the same rules, and they never leave the tab.

Item fields

Changelog

A change to any item is a change to the set version. Results from different versions are not comparable, which is why the version is shown on every run.

2. Scoring

Each answer is scored on its own, by a rule that does not involve a model. The same reply always produces the same score. An item scores 1 or 0; the pass threshold is 0.7, so with binary graders a pass means the rule matched.

Multiple choice (mcq)

The grader pulls the letters out of the reply and compares them as a set with the expected letters. It reads an answer whether the model writes b, B) or a sentence ending in the letter. Two rules earn their keep here: a trailing article — “the answer is a” — is not read as option a, and a letter must stand on a word boundary, so a refusal containing the word “can” does not count as choosing c.

Code (code)

The grader checks the shape of the answer, not its runtime. Nothing is executed, in your browser or anywhere else:

  1. the required function name and parameter count are present;
  2. the body returns or yields something;
  3. every require string appears;
  4. no forbid construct appears, and no globally forbidden one either.

The global rule is a single pattern: an answer that imports anthropic, openai, groq or google.generativeai fails. It exists because a model asked for a CSV parser answered by importing an SDK and delegating the parsing to another model — an answer that passed every item-level check while doing none of the work.

Short answer (short)

The expected phrase has to appear inside the reply, on word boundaries, after both sides are normalised: Unicode folded to a canonical form, case dropped, whitespace collapsed, and digit grouping removed so that 299 792, 299,792 and 299792 are the same number. Listed aliases are accepted too.

Why phrase matching, not exact equality

Models answer in sentences. Demanding an exact string would fail “The capital of Australia is Canberra” against Canberra. The cost of this choice is the opposite error: a long enough reply can contain the right phrase by accident, which is why short answers are kept to a key term.

3. Confidence intervals

A run here is 15 to 50 questions. On that few, a single percentage is close to meaningless on its own, so every accuracy is shown with a Wilson 95% interval beside it.

Wilson score interval
        p̂ + z²/2n                z         ⎧ p̂(1-p̂)   z²  ⎫
centre = ───────────   margin = ───────  √ ⎨ ─────── + ──── ⎬
         1 + z²/n               1 + z²/n   ⎩    n      4n²  ⎭

p̂ = correct / n      n = scored answers      z = 1.96  (95%, two-sided)
Implemented as wilsonConfidenceInterval(correct, total, z); the result is clamped to 0–100%.

Why Wilson and not the normal approximation

The textbook interval, p̂ ± z·√(p̂(1-p̂)/n), assumes a large sample and a proportion away from the ends. A 15-question run breaks both. At 14 out of 15 it puts the upper bound above 100%, and at 15 out of 15 it collapses to zero width — claiming certainty from fifteen answers. Wilson stays inside 0–100% and stays wide when the sample is small, which is the honest answer.

ResultAccuracyWilson 95%Reading
14 / 1593%70–99%Fifteen questions cannot tell 93% from 75%.
22 / 2588%70–96%Still overlaps the row below it.
20 / 2580%61–91%Overlaps 22/25, so the two are a tie.
40 / 5080%67–89%Same accuracy as above, narrower: more questions, more signal.

The last two rows are the argument for longer runs. The accuracy is identical; only the interval changes.

4. Statistical ties

A model card carries a ≈ statistical tie flag when its interval overlaps the leader's. The flag is on the challenger, never on the leader, and it never appears when only one model ran.

What it means: on this many questions, the difference between the two could be sampling noise. It is not a claim that the models are equally good — only that this run cannot separate them. The remedy is more questions, not a different reading of the same ones.

Worked example

22/25 is 88% with an interval of 70–96%. 20/25 is 80% with 61–91%. The ranges overlap between 70% and 91%, so the eight-point gap in accuracy is not something 25 questions can establish. Both get the flag relative to a leader they overlap.

5. LLM-as-judge

A judge is optional. When you pick one, it reads every answer and returns a score from 0 to 5 with a one-line reason. It is advisory: the judge never changes pass or fail. The verdict always comes from the deterministic grader, and the judge's score sits next to it.

The prompt

Sent to the judge, verbatim
You are evaluating an LLM's answer. Score it 0-5 and explain in one sentence.

Question: {prompt}
Expected answer: {expected}
Model's answer: {actual}

Respond in this exact format:
SCORE: [0-5]
REASON: [one sentence]
The reply is parsed with a strict pattern. A judge that answers in any other shape is recorded as unscored rather than guessed at.

Known biases, and what is done about them

Not yet done

Judge-versus-human agreement has not been measured on these sets. Until it is, the judge average is an opinion with known biases, not a calibrated metric — treat it as a reading aid for failures, which is what it is designed for.

6. Cost and tokens

The tool reports tokens in and tokens out, summed from the usage figures each provider returns with its response. It does not convert those into money, and no screen shows a price.

The reason is that a number in dollars would be a guess. Prices differ per provider, per model and per tier; they change without notice; and a run through an OpenRouter key is billed differently from the same model called directly. Reporting tokens means reporting what was actually measured.

Not implemented

“Cost per correct answer” does not exist in this tool yet. The OpenRouter catalogue does publish per-token prices, so the figure is computable for models listed there — but it would be wrong for a model called directly on a provider key, which is the other half of how this tool runs. Until that is resolved, tokens are what you get.

7. Limitations and known issues

Reasons to distrust a number from this tool, in roughly the order they matter.

Contamination

These questions are public. Public questions end up in training data, and a model that has seen an item is not being tested on it, it is being asked to recall it. A held-out set — questions never published, used only for the ranking — is the standard defence, and this tool does not have one yet. Treat high scores on the public sets as an upper bound.

Sample size

Fifteen to fifty questions is a small sample, and the intervals in section 3 show how small. Most differences between frontier models on these sets will come out as a tie. That is the honest result, not a defect in the measurement.

The ISTQB set is not official

The items are written against Foundation-level syllabus topics in the style of the exam. They are not the official ISTQB sample exam, are not endorsed by ISTQB, and should not be used as exam preparation. The name describes the subject matter, not the provenance.

Graders check shape, not behaviour

Code answers are never executed. A function with the right name, the right parameter count, a return statement and no forbidden construct passes — even if the logic inside is wrong. This catches the failure modes worth catching at this scale and misses everything else; a passing code item means “plausibly shaped”, not “correct”.

Short answers can match by accident

Phrase matching inside a long reply will occasionally reward an answer that contains the right words for the wrong reason. The mitigation is keeping expected answers to a key term, which narrows but does not close the gap.

One run is one sample

Models are non-deterministic. The same questions, the same model and the same settings can give a different score an hour later. Nothing here is averaged across runs.

The judge is uncalibrated

See section 5. Its biases are known and unmeasured on these sets.

Found something wrong here, or a case where the scoring disagrees with what you would call correct? That is the most useful kind of bug report. The engine is scoring.js, and its behaviour is pinned by the test suite in the repository.