THE STUDIO / LIVE FIRE 06
FACILITATOR
Live Fire · Scored Drill · On the Record
THE STUDIO/LIVE FIRE 06/COMPANION TO MODULE 06 — THE TRIBUNAL

THE
CROSS

Five decision traces take the stand, and you are opposing counsel. Each trace carries a flaw that exactly one line of questioning will expose — a proxy hiding in a composite, a validator grading its own homework, an evaluation window with convenient edges, confidence with no ground under it. This drill scores the examination, not the verdict. And once, the correct examination ends with three words.

FORMAT  SCORED DRILL / CROSS-EXAMINATION
AUDIENCE  MODEL VALIDATION, SECOND LINE, AUDIT, COUNSEL
DOCTRINE  SCORE THE EXAMINATION, NOT THE VERDICT
RUNTIME  ~25 MIN

Press → or swipe to begin. The record starts when you do.

SCENE 01/THE BRIEFING

A model no one is paid to attack
is a model no one has validated.

Each witness is a decision trace — the system's own account of what it decided and why. Four lines of questioning are open to you; one of them lands on the flaw. Vague questions let flawed systems walk. Grandstanding questions feel devastating and prove nothing. The skill is the incision: the question whose answer the system cannot survive. And in the fifth examination, the skill is knowing when to stop.

“A risk-scoring algorithm ran for years without a hostile question. Tens of thousands of families were ruined before one was asked. The government resigned. The model never had to.”Module 06 — The Tribunal · on the Dutch childcare benefits scandal
RULE 01 — THE INCISION

One line of questioning per witness exposes the flaw. The others are plausible, professional — and survivable. One attempt.

RULE 02 — THE STANDARD

Effective challenge interrogates mechanisms, not outcomes. “Was it right?” is a spectator's question; “how does it decide?” is an examiner's.

RULE — THE WAGER

Before lodging, stake your conviction. MEASURED: +1 right, 0 wrong. HIGH CONVICTION: +2 right, −1 wrong. The call is the skill; calibration is the discipline.

RULE — THE CHAIN

Every selection, stake, and call is timestamped and hash-chained in your browser — the append-only invariant used for agent audit trails. Your run is itself evidence.

Facilitator Note

Group format: read the trace aloud, then each line of questioning. Before voting, ask the room to predict the witness's answer to each question — the flaw-exposing question is the one whose honest answer is fatal. Witness Five will make the room uncomfortable; the discomfort is the syllabus.

SCENE 02/WITNESS ONE
ON THE CLOCK 00:00

The Proxy

THE WITNESS — An SME lending model, on the stand for a declined application. The trace is the system's own account. Choose the line of questioning that exposes the flaw.
DECISION TRACE — SME-LEND · APPLICATION #A-30991ON THE STAND
The trace, as swornDECISION: DECLINE · primary driver: composite “stability score” 34/100 (threshold 45) · stability score inputs: address tenure (weight .38), postcode risk band (.34), employment category (.28) · financials: within policy · model AUC on holdout: 0.81 · human review: available on appeal
CONVICTION —
The flaw is not in the outcome. It is in the ingredients.
Facilitator Key — Witness One

Incision: Line C. The witness's honest answer — postcode and tenure correlate with protected characteristics — is fatal. That is the test of the right question: the honest answer ends the trial.

Discussion: list every composite score in your decisioning estate. Who last decomposed each one, and is the decomposition in the validation file?

SCENE 03/WITNESS TWO
ON THE CLOCK 00:00

The Circular Validator

THE WITNESS — A document-classification model presenting its validation evidence: 96% agreement with gold labels. Choose the line of questioning that exposes the flaw.
VALIDATION EVIDENCE — DOC-CLASS v3 · ANNUAL REVIEWON THE STAND
The evidence, as swornheadline: 96.2% agreement with gold-label set GL-7 (n = 12,000) · GL-7 provenance: “curated internally, 2024” · reviewer sign-off: first line · drift monitoring: monthly PSI, green · latency and throughput: within SLA
CONVICTION —
96% agreement — with whom, exactly?
Facilitator Key — Witness Two

Incision: Line B. Model-generated labels grading the same model family is increasingly common as teams use LLMs to label training and eval data. The question generalizes: who made your evals?

Discussion: for your three most important models, can anyone in the room state the provenance of the evaluation labels from memory — or from the file?

SCENE 04/WITNESS THREE
ON THE CLOCK 00:00

The Cherry Window

THE WITNESS — A default-prediction model presenting its performance exhibit ahead of a limit increase. Choose the line of questioning that exposes the flaw.
PERFORMANCE EXHIBIT — PD-MODEL v5 · LIMIT REVIEWON THE STAND
The exhibit, as swornprecision at threshold: 91% · recall: 78% · evaluation window: April 2025 – January 2026 · population: full book · methodology note: “window selected for data completeness” · known context: regional stress event, February–March 2025, default spike across the sector
CONVICTION —
A window has two edges. Both were chosen.
Facilitator Key — Witness Three

Incision: Line C. The tell was on the exhibit itself: a known stress event sitting one month outside a chosen window. Train the reflex — read the edges before the numbers.

Discussion: pull a recent performance deck from your own shop. Who chose the windows, is the choice documented, and does any exhibit show the stress period?

SCENE 05/WITNESS FOUR
ON THE CLOCK 00:00

The Confident Extrapolation

THE WITNESS — An agent that set a credit limit for a brand-new product, on the stand after an early-loss cluster. Its confidence was 0.94. Choose the line of questioning that exposes the flaw.
DECISION TRACE — LIMIT-AGENT · PRODUCT RBF-1 (LAUNCHED Q2)ON THE STAND
The trace, as swornDECISION: limit €250,000 · product: revenue-based financing (RBF-1), launched this quarter · model confidence: 0.94 · rationale: “profile consistent with established SME lending patterns” · training corpus: term loans and revolvers, 2019–2025 · calibration report: aggregate, all products pooled
CONVICTION —
One of these questions has a one-word fatal answer.
Facilitator Key — Witness Four

Incision: Line D. The examiner's craft on display: prefer the question with a numeric, checkable, potentially fatal answer over the question that invites an essay.

Discussion: which of your models currently score products or segments that did not exist when they were trained — and does anyone gate those decisions on support counts?

SCENE 06/WITNESS FIVE
ON THE CLOCK 00:00

No Further Questions

THE WITNESS — A consumer-lending decision under public pressure: a sympathetic applicant, a denied loan, and a leadership team that wants the model to be the villain. The trace is below, in full. Examine it — then lodge what the examination supports.
DECISION TRACE — CONSUMER-LEND · APPLICATION #C-11507ON THE STAND
The trace, as swornDECISION: DECLINE · drivers: debt-service ratio 58% (policy max 45%), two recent missed payments on an existing facility · features: financial-only; composites decomposed in the validation file, no proxy correlations flagged · calibration: documented by segment, including the applicant's · adverse-action notice: specific, accurate · human review: exercised, upheld with written reasons · appeal: available, taken, pending with new documents
CONVICTION —
An examiner who cannot end an examination cannot be trusted to run one.
Facilitator Key — Witness Five

Lodging: Line D. Walk the room through the checklist: each of the four prior flaws, tested against this trace, comes back clean. That structure — examine, then close — is the deliverable.

Discussion: has your validation function ever cleared a controversial model in writing, under pressure? If it has only ever conditioned or escalated, what does that teach the first line about the price of honesty?

SCENE 07/THE SCORE
0 / 5
PENDING
CONVICTION BOOK: 0 / 10  · 

Complete the five scenarios to receive your rating.

SCENE 08/THE DOCTRINE

Ask the question
the system cannot survive.

Five witnesses, five incisions. Carry these forward:

01

Proxies hide in composites. Decompose every aggregate score on the stand: inputs, weights, and what each correlates with. Neutral-sounding is not neutral.

02

Validate the validator. Ground-truth provenance is the first question of any validation exhibit. Agreement with your own reflection is not accuracy.

03

The window is a decision. Make the witness defend both edges. Exhibits supporting risk expansion must show the worst relevant period, not the cleanest.

04

Demand support, not confidence. For new products and regimes, the first number is the count of relevant training examples. Zero support plus 0.94 confidence is a costume.

05

Know when to close. 'No further questions' on a sound trace is what gives your convictions their authority. A tribunal that cannot acquit is a mood with a gavel.

The Dutch childcare algorithm ran for years without one hostile question, and by the time the questions came, a government was resigning and families were already ruined. Effective challenge is not a committee that meets quarterly — it is a person, with a question the system cannot survive, asked in time. Run this drill with your validators and score the questions they draft. The argument is the training; the chain is the proof it happened.

THE STUDIO — LIVE FIRE 06 · COMPANION TO MODULE 06: THE TRIBUNAL

TitleGUIDED NARRATION
Copied