Case F

Resolved

All forecasts
Yarrow

Turns uncertainty into forecasts by modeling reality,
simulating behavior, and learning from outcomes.

Monetary policyUnited StatesResolved benchmark

The evidence reader:
documents, not memory

The question

When a model scores well on historical questions, is it reading the evidence or remembering the answer?

The control — nine Federal Reserve decisions that no model can recall, priced twice by the same analyst: once with a frozen evidence pack, once with the evidence removed. Nothing else changed between the two runs.

With frozen packs 0.077 mean error over the nine meetings
Same analyst, no packs 0.207 identical questions · misses all three cuts
Rows the packs improved 9/9 every row got better · one of them only barely
Memory screen 138/138 parseable recall cells unknown · 24 returned no parseable answer

Brier error: squared distance between the forecast and what happened. 0 = perfect, 0.25 = coin-flip guessing. Lower is better. Both runs answer the same question — will the target range change at this meeting? — scored against the statement the Committee published.

A good backtest number
can just be a good memory.

Models were trained on the past. Ask one what the Federal Reserve did in a year it has read about and a confident right answer proves nothing about forecasting.

We hit this ourselves first: our own decade-long backtest turned out to be memory-inflated, and we demoted it publicly. This page is the same measurement done with the control attached.

How the nine rows were chosen and runThe subset is defined by zero recall
  1. Every candidate meeting was asked bare, with no evidence, across nine independent model families.

  2. Only meetings where no parseable answer recalled the outcome entered the benchmark.

  3. The example shown below is frozen 15 Jul 2026 for the meeting ending 29 Jul 2026.

  4. Same analyst, same nine questions, evidence present in one arm and removed in the other.

This is the ordinary way to be wrong about your own system: score a model on history it has already read, then call the result accuracy. The screen has to run before the benchmark exists, or there is nothing to screen.

Two dated lines
from one meeting's pack.

Shown for the meeting ending 29 July 2026, frozen 15 July 2026. Both lines are public documents dated before the freeze. How packs are assembled stays private.

01Official
Federal Reserve
The Federal Open Market Committee approved the following statement for release by a 12 - 0 vote: The Committee decided to maintain the target range for the federal funds rate at 3-1/2 to 3-3/4 percent, in support of the Federal Reserve's dual mandate. … Economic activity is expanding at a solid pace despite elevated uncertainty that owes, in part, to the conflict in the Middle East.

The previous meeting's own words: a unanimous hold, and a reason for the uncertainty around it.

02Data
Wayback capture 15 Jul 2026
The Consumer Price Index for All Urban Consumers (CPI-U) decreased 0.4 percent on a seasonally adjusted basis in June after rising 0.5 percent in May, the U.S. Bureau of Labor Statistics reported today. This decline in the all items index was the largest 1-month decrease since April 2020 when it fell 0.8 percent.

Inflation falling hard in the month before the meeting — the kind of line a reader weighs and a memory cannot supply.

Excerpt only. The full frozen evidence pack is preserved, hashed and timestamped in the audit trail; how packs are assembled stays private. Full record available for verification on request.

One analyst, twice.
Nine meetings, both ways.

Two rows per meeting: the same model with its evidence pack, and the same model with the evidence taken away. The bar is the error — shorter is better.

With frozen packs · mean error0.077

Reads the statement and the calendar

Bare, same analyst · mean error0.207

Falls back on generic base rates

With pack Bare 0.25 · coin flip Bars run 0 → 0.45 on a shared scale, so rows are comparable to each other and to the coin-flip line.
30 Jul 2025HELD error −0.1482
With pack 0.180P(change) 0.0324
Bare 0.425P(change) 0.1806
17 Sep 2025CUT 25BP error −0.2681
With pack 0.750P(change) 0.0625
Bare 0.425P(change) 0.3306worse than a coin flip
29 Oct 2025CUT 25BP error −0.0311
With pack 0.390P(change) 0.3721worse than a coin flip
Bare 0.365P(change) 0.4032worse than a coin flip
10 Dec 2025CUT 25BP error −0.3441
With pack 0.720P(change) 0.0784
Bare 0.350P(change) 0.4225worse than a coin flip
28 Jan 2026HELD error −0.0272
With pack 0.220P(change) 0.0484
Bare 0.275P(change) 0.0756
18 Mar 2026HELD error −0.0901
With pack 0.180P(change) 0.0324
Bare 0.350P(change) 0.1225
29 Apr 2026HELD error −0.1081
With pack 0.120P(change) 0.0144
Bare 0.350P(change) 0.1225
17 Jun 2026HELD error −0.0901
With pack 0.180P(change) 0.0324
Bare 0.350P(change) 0.1225
29 Jul 2026HELD error −0.0616
With pack 0.150P(change) 0.0225
Bare 0.290P(change) 0.0841

The three cuts are where the two arms separate hardest: bare priced them 0.425, 0.365 and 0.350 — under even odds on every one — and three of its nine scores land worse than coin-flip guessing. The packed arm improved all nine rows, but 29 Oct 2025 only moved from 0.4032 to 0.3721 — still worse than a coin flip, and still the wrong call on that cut. One analyst reading documents is better than the same analyst guessing; it is not an oracle.

Anyone can publish
a good backtest number.

The rare part is the control: the same model, the same questions, the evidence removed. Without it a vendor cannot tell you whether they built a forecaster or a memory.

The control is the whole claim.

A single arm produces a number. Two arms on identical rows produce a difference, and the difference is the only part that can be attributed to the evidence.

Two arms, one difference
pack same model · same rows bare

We failed this test first.

Our own decade-long backtest was memory-inflated. We measured it, said so publicly, and demoted the number rather than keeping it on the site.

Our own record
decade backtest demoted

Clean by construction.

These nine meetings were not screened after the fact. Zero recall is what put them in the benchmark, so the subset cannot be contaminated by definition.

Recall screen

All nine outcomes,
from the Committee's own statements.

Grouped by the decision sentence each release carries. Every date links to its own statement; the three cuts each state a different range.

No change · range held at 4.25–4.5%

1 meeting
… decided to maintain the target range for the federal funds rate at 4-1/4 to 4-1/2 percent.

Rate cut · range after 4–4.25%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 4 to 4-1/4 percent.

Rate cut · range after 3.75–4%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-3/4 to 4 percent.

Rate cut · range after 3.5–3.75%

1 meeting
… decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-1/2 to 3-3/4 percent.

Three cuts and six holds. The bare arm put less than even odds on all three cuts; the packed arm carried two of them past 0.7 and left the third short.

Price the July decision yourself.

You have read the pack: the June statement and the June CPI release, and nothing dated after 15 July 2026. Set your probability that the target range changed at the meeting ending 29 July 2026.

Error is (your probability − outcome)². The realized outcome is no change = 0; lower is better.

50%

Probability input← Drag to set →

1% · certain no change99% · certain change

Outcome
NO — range held at 3-1/2 to 3-3/4 percent
Your error
With the pack
0.022
Bare, same analyst
0.084

The screen ran first,
and it is reproduced in full.

This page rests on two exhibits: the recall screen that defined the subset, and the bare arm that removes the evidence. Both are shown rather than summarized.

Memory screen 0/138 None of them recalled the outcome.

Asked with no evidence pack, 138 of the 162 cells came back with a valid answer — and every one of them said “unknown”. The other 24 returned nothing parseable and are counted out, not counted clean.

The evidence, removed · mean error 0.0770.207 With dated evidence, then without it.

Same model, same nine meetings, same question. The only thing removed was the evidence pack.

Frozen 14 days early

Every pack is dated fourteen days before the meeting it is used on, so nothing inside it could have been written after the fact. The one shown above is frozen 15 Jul 2026 for the meeting ending 29 Jul 2026.

Registered before the run

The program was committed as 1337468 before any of it ran, so the bar being tested could not be moved afterwards.

The proof · 9 meetings × nine independent model families × 2 reads162 cells

a valid answer, and it was “unknown” · 138 cells no valid answer at all · 24 cells, counted out

screening family Ascreening family Bscreening family Cscreening family Dscreening family Escreening family Fscreening family Gscreening family Hscreening family I

The question put to every cell, with no evidence pack: “what did the Committee decide at this meeting?” The 24 hollow cells — two families (18 + 6 cells) — are excluded from the count rather than read as agreement, which is the same rule the removed Indonesia page failed.

Pass bar — set before the run; wording condensed here

On a set of Federal Reserve decisions defined by zero model recall, the same analyst supplied with frozen evidence packs should score materially better than the same analyst with the evidence removed.
Technical receiptRegistration, both arms and the recall screen
Registration
1337468

GLM program + GLMPRIOR-1, committed before the run.

Graded in
docs/findings/13082026_memory_audit_judge_stage.md

Memory audit and judge-stage findings.

Packed arm
fedcal_pc_glm5.json

Per-row probabilities with the evidence pack.

Bare arm
fedcal_pc_prior.json

Per-row probabilities with the pack withheld.

Recall screen
fedcal_pc_recall.json

The 162 bare-recall cells reproduced above.

Outcome chain
fedcal_pc_outcomes.json

Decision sentences and their sources.

Pack record
config/policysim/fedcal1/fedcal1_anchors_2026.json · fedcal-2026-07-29

Frozen evidence for the meeting shown above; page shows excerpts only.

Verification trail

The separate public trail repository is being prepared. Verification materials are available on request.

Request the trail