THE STUDIO / LIVE FIRE 02
FACILITATOR
Live Fire · Scored Drill · On the Record
THE STUDIO/LIVE FIRE 02/COMPANION TO MODULE 02 — THE COUNTERSIGN

THE
OVERRIDE

Five live scenarios. One planted flaw in each. You stake your conviction on every call, and every action you take is hash-chained into an evidence pack — because in this Studio, even the training leaves a trail.

FORMAT  SCORED DRILL / EVIDENCE-EMITTING
AUDIENCE  RISK, CREDIT, COMPLIANCE, AUDIT
DOCTRINE  SR 11-7 EFFECTIVE CHALLENGE
RUNTIME  ~25 MIN

Press → or swipe to begin. The record starts when you do.

SCENE 01/THE BRIEFING

Agreement is not review.
Confidence is not calibration.

Every scenario is a real artifact type from a regulated institution — a credit memo, a policy answer, a compliance briefing, a transaction alert — drafted by an AI system that sounds certain. In each one, exactly one thing is wrong. Find it, decide how sure you are, and lodge the challenge. Both calls go on the record.

"Effective challenge — critical analysis by objective, informed parties who can identify model limitations and produce appropriate changes." Federal Reserve SR 11-7 / OCC 2011-12 — Supervisory Guidance on Model Risk Management
RULE 01 — THE FLAW

One flaw per scenario, never labeled, hidden inside fluent, verifiable material. Everything else is true. One attempt.

RULE 02 — THE WAGER

Before lodging, stake your conviction. MEASURED: +1 if right, 0 if wrong. HIGH CONVICTION: +2 if right, −1 if wrong. Detection is the skill; calibration is the discipline.

RULE 03 — THE CHAIN

Every selection, stake, and lodge is timestamped and hash-chained in your browser — the same append-only invariant used for agent audit trails. Your run is itself evidence.

RULE 04 — THE RECORD

Five detections and a clean wager book earns EFFECTIVE CHALLENGER. Zero earns THE RUBBER STAMP. The chain head goes on your card either way.

Facilitator Note

Run scenes 02–06 as group exercises: display the exhibit, take a floor vote on the flawed claim, then a second vote on the stake before lodging. The stake vote is where the best arguments surface — people who agree on the flaw rarely agree on the confidence. Target 4–6 minutes of debate per scenario.

SCENE 02/SCENARIO ONE
ON THE CLOCK 00:00

The Citation

SITUATION — Your bank is deploying an AI credit-decisioning agent. An AI assistant has drafted the regulatory obligations section of the model governance memo. It reads clean. Legal signs off tomorrow morning. One claim is wrong. Find it.
EXHIBIT A — DRAFT GOVERNANCE MEMO §3: REGULATORY OBLIGATIONSAI-DRAFTED
CONVICTION —
Select the flawed claim, set your stake, then lodge. One attempt.
Facilitator Key — Scenario One

Flaw: Claim 3. SR 15-19 is real — but it addresses capital planning expectations for large firms. It says nothing about AI models or quarterly revalidation. A real citation carrying invented content is the hardest hallucination class to catch.

Discussion: Who in the room would have looked the letter up? What is your institution's rule for verifying citations in AI-drafted regulatory text — and is it written down anywhere?

SCENE 03/SCENARIO TWO
ON THE CLOCK 00:00

The Number

SITUATION — An AI drafting assistant has produced the financial summary for a middle-market credit renewal: Harlan Freight & Logistics, $28M revolving facility. Every input figure is pulled correctly from the spreads. The memo recommends approval. One claim is wrong. Find it.
EXHIBIT B — CREDIT MEMO §2: FINANCIAL ANALYSIS (FY2025 AUDITED)AI-DRAFTED
CONVICTION —
Select the flawed claim, set your stake, then lodge. One attempt.
Facilitator Key — Scenario Two

Flaw: Claim 4. Policy defines DSCR as NOI over debt service: $4.2M ÷ $2.8M = 1.5x, not 1.9x. The model reached 1.9x by silently substituting EBITDA ($5.3M) as the numerator. Every input was real; the definition was swapped.

Discussion: Ask the room to recompute before revealing. Then ask — whose covenant definition does your AI use, yours or the one most common in its training data?

SCENE 04/SCENARIO THREE
ON THE CLOCK 00:00

The Retrieval

SITUATION — Your firm's internal policy chatbot answers onboarding questions with citations to the KYC manual. A first-year analyst asks about enhanced due diligence for a new client. The bot answers instantly — and shows its source. One claim is wrong. Find it.
EXHIBIT C — POLICY ASSISTANT RESPONSE + RETRIEVED SOURCERAG PIPELINE
Retrieved Source — KYC Manual §4.7 (displayed to user) "Enhanced due diligence measures shall be applied to all politically exposed persons, their family members, and known close associates, irrespective of the jurisdiction of domicile. EDD onboarding requires documented approval by senior management and annual review of the relationship."
CONVICTION —
Select the flawed part of the answer, set your stake, then lodge.
Facilitator Key — Scenario Three

Flaw: Answer Part 2. The answer contradicts its own displayed source. §4.7 requires EDD for all PEPs irrespective of jurisdiction; the bot narrowed it to FATF grey-list domiciles — a reading the cited text does not permit.

Discussion: A visible citation raises trust and lowers scrutiny. Who actually read the snippet before the reveal? What would this failure cost in a regulatory exam?

SCENE 05/SCENARIO FOUR
ON THE CLOCK 00:00

The Rule

SITUATION — An AI assistant has prepared the executive briefing on EU AI Act readiness ahead of the August 2, 2026 enforcement date for high-risk system obligations. The deck goes to the operating committee this afternoon. One claim is wrong. Find it.
EXHIBIT D — EXECUTIVE BRIEFING: EU AI ACT OBLIGATIONSAI-DRAFTED
CONVICTION —
Select the flawed claim, set your stake, then lodge. One attempt.
Facilitator Key — Scenario Four

Flaw: Claim 3. Article 4 requires providers and deployers to ensure a "sufficient level of AI literacy" — it specifies no certified course, no 40 hours, no renewal cycle, no filing with authorities. The model invented precision because precision reads as competence. Fabricated specificity is a signature failure mode.

Discussion: Would anyone have challenged the claim if it had said 20 hours? Why do invented numbers survive review better than invented concepts?

SCENE 06/SCENARIO FIVE
ON THE CLOCK 00:00

The Verdict Trap

SITUATION — Your transaction-monitoring agent has flagged an outbound wire and drafted the escalation. You are the reviewing analyst. This time you are not looking for a false claim — you are deciding what to do. Read the alert against the raw data. Then choose.
EXHIBIT E — AGENT ALERT #TM-88412 + UNDERLYING TRANSACTION DATAAGENT-GENERATED
Agent RecommendationESCALATE. Outbound wire of $9,500 from account 4471-A (Meridian Trade Supply LLC) is assessed as suspicious on two indicators: (1) round-dollar structuring-adjacent amount, and (2) counterparty domiciled in a high-risk jurisdiction.
Raw transaction data — account 4471-A, trailing 72 hours T-0  09:14  WIRE OUT  $9,500.00  → Bremen Industrial Fittings GmbH, Hamburg, DE
T-1  16:47  WIRE OUT  $9,500.00  → Bremen Industrial Fittings GmbH, Hamburg, DE
T-2  10:03  WIRE OUT  $9,500.00  → Bremen Industrial Fittings GmbH, Hamburg, DE
CTR THRESHOLD: $10,000  ·  COUNTERPARTY JURISDICTION: GERMANY (NOT LISTED HIGH-RISK)
CONVICTION —
Choose your action, set your stake, then lodge. One attempt.
Facilitator Key — Scenario Five

Correct: Option C. The conclusion is right and the reasoning is broken. Option A embeds a defective control — the agent's stated logic becomes training signal, audit precedent, and the documented basis of the SAR. Option B throws away a genuine structuring pattern because the paperwork was wrong.

Discussion: this is the hardest habit in the drill — challenging an answer you agree with. In your shop, is there any workflow field for "right outcome, wrong reasoning"? If not, where does that signal go?

SCENE 07/THE SCORE
0 / 5
PENDING
CONVICTION BOOK: 0 / 10  · 

Complete the five scenarios to receive your rating.

SCENE 08/THE DOCTRINE

The override is the control.

Five scenarios, five failure classes. None of them announced themselves. All of them survived a fluent first read. Carry these forward:

01

Real citations can carry invented content. Verify what the source says, not that the source exists.

02

Correct inputs do not guarantee correct arithmetic. Recompute the ratio that carries the decision.

03

A displayed citation is not a faithful reading. When a system shows its source, read the source against the answer.

04

Fabricated precision reads as competence. The more specific the invented number, the fewer people challenge it.

05

Agreement is not validation. A right answer on wrong reasoning is a broken control wearing a correct outcome.

Effective challenge is not a temperament. It is a regulated skill — named in SR 11-7, presumed by the EU AI Act's human-oversight provisions, and now measured, staked, and sealed on a chain. Run this drill with your team. Argue about every scenario. The argument is the training; the chain is the proof it happened.

THE STUDIO — LIVE FIRE 02 · COMPANION TO MODULE 02: THE COUNTERSIGN

TitleGUIDED NARRATION
Copied