AI

No model judge

What is recorded

D-1 recorded

Evalset grades with deterministic graders, not LLM-as-judge. What was cut to get there: LLM-as-judge grading. The reason is not recorded yet.

Source: Bryan, product-facts kickoff request, 2026-09-21. Last verified 2026-09-21.

What that rules out

A model judge is the standard answer to grading open-ended output. You give a second model the input, the output and a rubric, and ask it to score. It is fast, it scales to criteria no rule can express, and it is extremely easy to stand up.

Ruling it out means the eval suite can only test what a rule can check. Exact matches, parseable shapes, required fields present, values in range, forbidden strings absent, ordering preserved. Anything that needs a reader's judgement is out of scope for the suite by construction.

That is a real loss and I am not going to write around it. Tone, whether an explanation explains, whether a summary is faithful in the sense a person means: none of that is reachable by a deterministic grader, and a suite that cannot see them will happily go green on output that is technically correct and useless.

What I have not written down

Why.

The ledger line above has the decision and the cut. Where the reasoning should be, it has an open TODO. I could write a persuasive paragraph here about reproducibility, about a judge that drifts between model versions, about grading a model with a model being a measurement instrument made of the thing being measured. It would read well. It would also be me reconstructing someone's reasoning from the outside and presenting it as a record.

This desk's whole claim is that each tool states what it refuses to be. Writing an invented rationale would make it a desk that states what it would be nice to have refused.

Where the judgement goes instead

Ruling out a model judge does not remove the need for judgement; it moves it to a human with a written rubric. That is what Rubric is for: criteria on a 1 to 5 scale with a description per level, blind mode so the source label is hidden until the score is in, and weighted kappa between two scorers so you find out whether the levels mean anything before you trust the numbers.

That app is built but not live yet, so that link will not resolve until its row says Live. I am leaving the sentence in because the relationship between the two decisions is the point: the judgement did not disappear, it got a scale and a second scorer.

What the finished memo needs

  1. The specific failure with a model judge that made this not worth trying.
  2. Which criteria the suite now cannot test at all, named.
  3. The condition under which I would revisit it, written down before I want to revisit it.
  4. The date the decision was made, which is not

T-3 recorded

2026-09-21, company/product-facts.md drafted for Bryan's review.

Source: the product-facts drafting task. Last verified 2026-09-21.

, the date it was written down.