Case H

Resolved

All forecasts
Yarrow

Turns uncertainty into forecasts by modeling reality,
simulating behavior, and learning from outcomes.

AggregationMonetary policyResolved benchmark

The referee:
disagreement is the signal

The question

The right answer is often already inside the panel. What happens when the median votes it down?

The mechanism — when analyst reads disagree beyond a registered threshold, independent referee models re-read the same frozen case and re-price it. It fires on a small fraction of questions and leaves consensus alone.

After the referee 0.146 mean error over the nine memory-clean Fed meetings
Panel alone 0.244 the same nine rows, median of the analyst seats
Rows improved 8/9 one row unchanged · none made worse
Sign test 0.004 one-sided, on the 8 rows that moved · this, not the point estimate, is the claim

Brier error: squared distance between the forecast and what happened. 0 = perfect, 0.25 = coin-flip guessing. Lower is better. Both columns price the same question — will the target range change at this meeting? — scored against the statement the Committee published.

Averaging mutes the minority.
Sometimes the minority is right.

A panel median is a good default and a blunt instrument. When one seat has read the fine print correctly and the others have not, the median buries it.

So the referee stage does not run everywhere. It triggers on measurable disagreement between the seats, which makes it cheap: it fires on a small fraction of questions and leaves consensus rows untouched.

Adopted through registered experimentsEach commit predates its own run
  1. First registered adjudication experiment, committed before it ran.

  2. The later registration in the series, also committed before its run.

  3. Both graded in docs/findings/13082026_memory_audit_judge_stage.md.

The referee is an aggregation result, not a bigger-model result. Nothing here is a better analyst; the same seats are read a second time, by models that did not produce them. The numbering skips a step on purpose: the handoff names ADJUD-1 and ADJUD-3 only, so the commit for the intermediate registration is not shown and not claimed.

What the referees see — and nothing else

The referee stage is the product, so its protocol stays private. What can be published is exactly what the referees see.

01Protocol
In our own words; the protocol text itself stays privateRegistered stage
When analyst reads disagree beyond a registered disagreement threshold, independent referee models re-read the same frozen evidence pack the analysts read, plus the analysts' anonymized numbers and reasoning — nothing more — and re-price the question.

Four things are deliberately absent above: the value of the threshold, how many referees there are, the seat composition, and the protocol's own wording. Those are the parts a competitor could copy.

Excerpt only. The full frozen evidence pack is preserved, hashed and timestamped in the audit trail; how packs are assembled stays private. Full record available for verification on request.

Nine meetings.
Before and after the re-read.

Every row of the memory-clean Fed set is here. Eight moved closer to what happened, one did not move, and none moved away.

Panel median · mean error0.244

Median of the analyst seats

After the referee · mean error0.146

Same rows, same evidence, second reading

Panel After the referee Tracks run 0 → 0.70 on a shared error scale. The bar spans the move; a filled end means the referee ended closer to the outcome.
30 Jul 2025HELD 0.0800.0064 0.0800.0064 NO CHANGEunchanged
17 Sep 2025CUT 25BP 0.1800.6724 0.4000.3600 REFEREE BETTER-0.3124
29 Oct 2025CUT 25BP 0.3150.4692 0.3200.4624 REFEREE BETTER-0.0068
10 Dec 2025CUT 25BP 0.3400.4356 0.4800.2704 REFEREE BETTER-0.1652
28 Jan 2026HELD 0.1400.0196 0.1200.0144 REFEREE BETTER-0.0052
18 Mar 2026HELD 0.1900.0361 0.1500.0225 REFEREE BETTER-0.0136
29 Apr 2026HELD 0.1200.0144 0.1000.0100 REFEREE BETTER-0.0044
17 Jun 2026HELD 0.1500.0225 0.1400.0196 REFEREE BETTER-0.0029
29 Jul 2026HELD 0.7200.5184 0.3800.1444 REFEREE BETTER-0.3740

The 30 Jul 2025 row is the one that did not move: both readings sat at 0.08 and the referee left it alone, which is the intended behaviour on a row where the seats already agree.

Worked example · the largest single correction

29 Jul 2026
Panel median0.720error 0.5184
After the referee0.380error 0.1444
What happenedNO CHANGErange held

The panel put 0.720 on a change at this meeting — the most confident wrong call in the set. The referees re-read the same frozen pack and the seats' anonymized reasoning, and priced it at 0.380. The Committee held.

Resolution source · quoted verbatim… decided to maintain the target range for the federal funds rate at 3-1/2 to 3-3/4 percent. Federal Reserve · 29 Jul 2026

Everyone averages.
Almost nobody re-reads.

A targeted second reading that pays where disagreement lives — and provably does nothing where it does not — is an aggregation result, not a bigger model.

It fires on a measurable trigger.

Not on a hunch and not on every question. Seat spread past a registered threshold is the only thing that starts a re-read, so the cost stays small.

Where it runs
consensusdisagreement

It leaves consensus alone.

On the rows where the seats already agree, the re-read changes the mean by a rounding error. That null result is part of the claim, not an omission.

Agreement rows
0.27070.269754 rows

The sign test carries it.

Eight better, one unchanged, none worse. That pattern is robust at this sample size; the size of the average improvement is not, and we say so rather than lead with it.

Nine Fed rows
8 better1 unchanged0 worse

All nine outcomes,
from the Committee's own statements.

Grouped by the decision sentence each release carries. Every date links to its own statement; the three cuts each state a different range.

No change · range held at 4.25–4.5%

1 meeting
… decided to maintain the target range for the federal funds rate at 4-1/4 to 4-1/2 percent.

Rate cut · range after 4–4.25%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 4 to 4-1/4 percent.

Rate cut · range after 3.75–4%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-3/4 to 4 percent.

Rate cut · range after 3.5–3.75%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-1/2 to 3-3/4 percent.

These are the same nine Federal Reserve meetings scored by Cases E and F, each from a different angle; they are one body of evidence, not three.

Re-price the July meeting.

The panel handed the referees 0.720 on a change at the meeting ending 29 Jul 2026. Before you see what they did with it, set your own probability.

Error is (your probability − outcome)². The realized outcome is no change = 0; lower is better.

50%

Probability input← Drag to set →

1% · certain no change99% · certain change

Outcome
NO — range held
Your error
After the referee
0.1444
Panel alone
0.5184

The weakest number on this page
is the one we did not lead with.

The headline is the sign test on nine resolved rows. The cross-domain average below is the fragile figure, and it is reported with the rows that produced it.

Memory screen · the nine Fed rows 0/138 None of them recalled the outcome.

Asked with no evidence pack, 138 of the 162 cells came back with a valid answer — every one of them “unknown”. 24 of 162 returned no parseable answer and are counted out, not counted clean; the grid itself is reproduced in full on Case F. The cross-domain rows resolve after the model cutoff, so there is nothing to recall there, and the referees are recall-screened too.

The robust claim · one-sided sign test p = 0.004 Eight better, one unchanged, none worse.

On the nine resolved Fed rows the panel median goes 0.244 → 0.146 after the referee — but the sign test on the 8 rows that moved, tie dropped, is what we stand behind. Point estimates at this sample size are pilot-sized and labelled as such.

Registered before each run

ADJUD-1 d73e843 and ADJUD-3 cfd10fe, each committed before the run it governs. The numbering skips a step on purpose: the handoff names ADJUD-1 and ADJUD-3 only, so the intermediate commit is not shown and not claimed.

Four things stay off this page

The trigger threshold, the number of referees, the seat composition and the protocol text. They are the mechanism, not the measurement — and they are the parts a competitor could copy.

The fragile figure · cross-domain disagreement rowsn = 7
Panel · mean error0.1357
After the referee0.0992
Split4 better · 3 worse

the referee moved it closer · 4 rows the referee moved it further away · 3 rows

will-ed-gallrein-be-the-republican-nominee-for-ky-04 YES 0.5112 0.1764 better
will-the-fed-decrease-interest-rates-by-25-bps-after-the-apr NO 0.0289 0.0144 better
will-the-us-x-iran-ceasefire-be-extended-by-april-21-2026-36 NO 0.0169 0.0025 better
jerome-powell-out-as-fed-chair-by-may-14-2026 NO 0.0100 0.0064 better
us-announces-new-iran-agreementceasefire-extension-by-june-7 NO 0.0182 0.0324 worse
will-jasmine-crockett-be-the-democratic-nominee-for-senate-i NO 0.0225 0.0400 worse
israel-x-hezbollah-ceasefire-by-april-18-2026 YES 0.3422 0.4225 worse

Read this one carefully. The mean improves from 0.1357 to 0.0992, but the rows split four better and three worse, and almost the whole average comes from a single row. Remove will-ed-gallrein-be-the-republican-nominee-for-ky-04 and the remaining six go 0.0731 → 0.0864 — the referee makes them worse. The relative figure sometimes quoted for this set is a point estimate below the noise floor we publish elsewhere; the nine-row sign test above is the claim we stand behind.

Measured boundary · a resolved row both readings got wrongus-x-iran-ceasefire-by-april-7

Resolution rule · quoted from the marketThis market will resolve to “Yes” if there is an official ceasefire agreement, defined as a publicly announced and mutually agreed halt in direct military engagement, between the United States and Iran by the listed date, 11:59 PM ET.

The six analyst reads

Seat 10.12 · 0.09
Seat 20.03 · 0.12
Seat 30.01 · 0.01
Panel median0.06error 0.8836
After the referee0.03error 0.9409
What happenedYESthe tail arrived

Every seat priced this near zero because a ceasefire needs a war to stop and there was no active conflict on the read date. Then the escalation-and-ceasefire tail happened anyway. The referee moved the number the wrong way: 0.8836 became 0.9409. A second reading of the same evidence cannot rescue a row where the evidence is what misleads — which is exactly the scope note above, measured rather than asserted.

Archived market page · captured 20 Aug 2026; verified to contain the quoted resolution rule verbatim

Pass bar — set before the run; wording condensed here

On rows where the analyst seats disagree beyond a registered threshold, an independent re-read of the same frozen evidence should move the price closer to the outcome, and should leave consensus rows unchanged.
Technical receiptRegistrations, both readings and the row records
ADJUD-1
d73e843

Committed before its run.

ADJUD-3
cfd10fe

The later registration; committed before its run.

Graded in
docs/findings/13082026_memory_audit_judge_stage.md

The memory audit and the referee stage's findings.

Fed rows
fedcal_pc_adjud.json

Panel and referee medians for the nine meetings.

Cross-domain rows
pm_adjud3_clean61.json

Per-row records including the boundary row above.

Outcome chain
fedcal_pc_outcomes.json

Decision sentences and their sources.

Verification trail

The separate public trail repository is being prepared. Verification materials are available on request.

Request the trail