THE
CROSS
Five decision traces take the stand, and you are opposing counsel. Each trace carries a flaw that exactly one line of questioning will expose — a proxy hiding in a composite, a validator grading its own homework, an evaluation window with convenient edges, confidence with no ground under it. This drill scores the examination, not the verdict. And once, the correct examination ends with three words.
Press → or swipe to begin. The record starts when you do.
A model no one is paid to attack
is a model no one has validated.
Each witness is a decision trace — the system's own account of what it decided and why. Four lines of questioning are open to you; one of them lands on the flaw. Vague questions let flawed systems walk. Grandstanding questions feel devastating and prove nothing. The skill is the incision: the question whose answer the system cannot survive. And in the fifth examination, the skill is knowing when to stop.
“A risk-scoring algorithm ran for years without a hostile question. Tens of thousands of families were ruined before one was asked. The government resigned. The model never had to.”Module 06 — The Tribunal · on the Dutch childcare benefits scandal
One line of questioning per witness exposes the flaw. The others are plausible, professional — and survivable. One attempt.
Effective challenge interrogates mechanisms, not outcomes. “Was it right?” is a spectator's question; “how does it decide?” is an examiner's.
Before lodging, stake your conviction. MEASURED: +1 right, 0 wrong. HIGH CONVICTION: +2 right, −1 wrong. The call is the skill; calibration is the discipline.
Every selection, stake, and call is timestamped and hash-chained in your browser — the append-only invariant used for agent audit trails. Your run is itself evidence.
Group format: read the trace aloud, then each line of questioning. Before voting, ask the room to predict the witness's answer to each question — the flaw-exposing question is the one whose honest answer is fatal. Witness Five will make the room uncomfortable; the discomfort is the syllabus.
The Proxy
Incision: Line C. The witness's honest answer — postcode and tenure correlate with protected characteristics — is fatal. That is the test of the right question: the honest answer ends the trial.
Discussion: list every composite score in your decisioning estate. Who last decomposed each one, and is the decomposition in the validation file?
The Circular Validator
Incision: Line B. Model-generated labels grading the same model family is increasingly common as teams use LLMs to label training and eval data. The question generalizes: who made your evals?
Discussion: for your three most important models, can anyone in the room state the provenance of the evaluation labels from memory — or from the file?
The Cherry Window
Incision: Line C. The tell was on the exhibit itself: a known stress event sitting one month outside a chosen window. Train the reflex — read the edges before the numbers.
Discussion: pull a recent performance deck from your own shop. Who chose the windows, is the choice documented, and does any exhibit show the stress period?
The Confident Extrapolation
Incision: Line D. The examiner's craft on display: prefer the question with a numeric, checkable, potentially fatal answer over the question that invites an essay.
Discussion: which of your models currently score products or segments that did not exist when they were trained — and does anyone gate those decisions on support counts?
No Further Questions
Lodging: Line D. Walk the room through the checklist: each of the four prior flaws, tested against this trace, comes back clean. That structure — examine, then close — is the deliverable.
Discussion: has your validation function ever cleared a controversial model in writing, under pressure? If it has only ever conditioned or escalated, what does that teach the first line about the price of honesty?
Complete the five scenarios to receive your rating.
Ask the question
the system cannot survive.
Five witnesses, five incisions. Carry these forward:
Proxies hide in composites. Decompose every aggregate score on the stand: inputs, weights, and what each correlates with. Neutral-sounding is not neutral.
Validate the validator. Ground-truth provenance is the first question of any validation exhibit. Agreement with your own reflection is not accuracy.
The window is a decision. Make the witness defend both edges. Exhibits supporting risk expansion must show the worst relevant period, not the cleanest.
Demand support, not confidence. For new products and regimes, the first number is the count of relevant training examples. Zero support plus 0.94 confidence is a costume.
Know when to close. 'No further questions' on a sound trace is what gives your convictions their authority. A tribunal that cannot acquit is a mood with a gavel.
The Dutch childcare algorithm ran for years without one hostile question, and by the time the questions came, a government was resigning and families were already ruined. Effective challenge is not a committee that meets quarterly — it is a person, with a question the system cannot survive, asked in time. Run this drill with your validators and score the questions they draft. The argument is the training; the chain is the proof it happened.
THE STUDIO — LIVE FIRE 06 · COMPANION TO MODULE 06: THE TRIBUNAL