THE
OVERRIDE
Five live scenarios. One planted flaw in each. You stake your conviction on every call, and every action you take is hash-chained into an evidence pack — because in this Studio, even the training leaves a trail.
Press → or swipe to begin. The record starts when you do.
Agreement is not review.
Confidence is not calibration.
Every scenario is a real artifact type from a regulated institution — a credit memo, a policy answer, a compliance briefing, a transaction alert — drafted by an AI system that sounds certain. In each one, exactly one thing is wrong. Find it, decide how sure you are, and lodge the challenge. Both calls go on the record.
"Effective challenge — critical analysis by objective, informed parties who can identify model limitations and produce appropriate changes." Federal Reserve SR 11-7 / OCC 2011-12 — Supervisory Guidance on Model Risk Management
One flaw per scenario, never labeled, hidden inside fluent, verifiable material. Everything else is true. One attempt.
Before lodging, stake your conviction. MEASURED: +1 if right, 0 if wrong. HIGH CONVICTION: +2 if right, −1 if wrong. Detection is the skill; calibration is the discipline.
Every selection, stake, and lodge is timestamped and hash-chained in your browser — the same append-only invariant used for agent audit trails. Your run is itself evidence.
Five detections and a clean wager book earns EFFECTIVE CHALLENGER. Zero earns THE RUBBER STAMP. The chain head goes on your card either way.
Run scenes 02–06 as group exercises: display the exhibit, take a floor vote on the flawed claim, then a second vote on the stake before lodging. The stake vote is where the best arguments surface — people who agree on the flaw rarely agree on the confidence. Target 4–6 minutes of debate per scenario.
The Citation
Flaw: Claim 3. SR 15-19 is real — but it addresses capital planning expectations for large firms. It says nothing about AI models or quarterly revalidation. A real citation carrying invented content is the hardest hallucination class to catch.
Discussion: Who in the room would have looked the letter up? What is your institution's rule for verifying citations in AI-drafted regulatory text — and is it written down anywhere?
The Number
Flaw: Claim 4. Policy defines DSCR as NOI over debt service: $4.2M ÷ $2.8M = 1.5x, not 1.9x. The model reached 1.9x by silently substituting EBITDA ($5.3M) as the numerator. Every input was real; the definition was swapped.
Discussion: Ask the room to recompute before revealing. Then ask — whose covenant definition does your AI use, yours or the one most common in its training data?
The Retrieval
Flaw: Answer Part 2. The answer contradicts its own displayed source. §4.7 requires EDD for all PEPs irrespective of jurisdiction; the bot narrowed it to FATF grey-list domiciles — a reading the cited text does not permit.
Discussion: A visible citation raises trust and lowers scrutiny. Who actually read the snippet before the reveal? What would this failure cost in a regulatory exam?
The Rule
Flaw: Claim 3. Article 4 requires providers and deployers to ensure a "sufficient level of AI literacy" — it specifies no certified course, no 40 hours, no renewal cycle, no filing with authorities. The model invented precision because precision reads as competence. Fabricated specificity is a signature failure mode.
Discussion: Would anyone have challenged the claim if it had said 20 hours? Why do invented numbers survive review better than invented concepts?
The Verdict Trap
T-1 16:47 WIRE OUT $9,500.00 → Bremen Industrial Fittings GmbH, Hamburg, DE
T-2 10:03 WIRE OUT $9,500.00 → Bremen Industrial Fittings GmbH, Hamburg, DE
CTR THRESHOLD: $10,000 · COUNTERPARTY JURISDICTION: GERMANY (NOT LISTED HIGH-RISK)
Correct: Option C. The conclusion is right and the reasoning is broken. Option A embeds a defective control — the agent's stated logic becomes training signal, audit precedent, and the documented basis of the SAR. Option B throws away a genuine structuring pattern because the paperwork was wrong.
Discussion: this is the hardest habit in the drill — challenging an answer you agree with. In your shop, is there any workflow field for "right outcome, wrong reasoning"? If not, where does that signal go?
Complete the five scenarios to receive your rating.
The override is the control.
Five scenarios, five failure classes. None of them announced themselves. All of them survived a fluent first read. Carry these forward:
Real citations can carry invented content. Verify what the source says, not that the source exists.
Correct inputs do not guarantee correct arithmetic. Recompute the ratio that carries the decision.
A displayed citation is not a faithful reading. When a system shows its source, read the source against the answer.
Fabricated precision reads as competence. The more specific the invented number, the fewer people challenge it.
Agreement is not validation. A right answer on wrong reasoning is a broken control wearing a correct outcome.
Effective challenge is not a temperament. It is a regulated skill — named in SR 11-7, presumed by the EU AI Act's human-oversight provisions, and now measured, staked, and sealed on a chain. Run this drill with your team. Argue about every scenario. The argument is the training; the chain is the proof it happened.
THE STUDIO — LIVE FIRE 02 · COMPANION TO MODULE 02: THE COUNTERSIGN