“In support of its goals, the Committee decided to maintain the target range for the federal funds rate at 4-1/4 to 4-1/2 percent. In considering the extent and timing of additional adjustments to the target range for the federal funds rate, the Committee will carefully assess incoming data, the evolving outlook, and the balance of risks. …”
Policy sizingUnited StatesResolved case
The sizing call:
the Fed's 2025 cuts
Not only whether the Federal Reserve would move, but by how much.
Registered format — each analyst spreads probability across HOLD, ±25bp and ±50bp. The answer is computed from those stored masses.
The two score families shown on this page retain their original labels. The Clean-9 score above is not the RPS comparison shown in the panel below.
01 / The setup
Calling the move
is not sizing the move.
Binary questions hide direction and size. A model can call a cut and still put no probability on the number of basis points that actually matters.
The action-space format was adopted through registered experiments on a decade of meetings, then graded on the recall-clean 2025–2026 subset.
The five-bucket format was registered before adoption.
The memory-screen program was committed before the clean subset was defined.
The size-grading rule was committed before the benchmark was graded.
Only the nine rows that passed the memory screen counted in the public result.
02 / The frozen evidence
Three dated lines.
No hidden recipe.
This public excerpt shows the evidence available for the first 2025 cut, frozen on 3 September 2025. How packs are assembled stays private.
“The Consumer Price Index for All Urban Consumers (CPI-U) increased 0.2 percent on a seasonally adjusted basis in July, after rising 0.3 percent in June, the U.S. Bureau of Labor Statistics reported today. Over the last 12 months, the all items index increased 2.7 percent before seasonal adjustment.”
“The unemployment rate, at 4.2 percent, also changed little in July. Employment continued to trend up in health care and in social assistance. Federal government continued to lose jobs.”
Excerpt only. The full frozen evidence pack is preserved, hashed and timestamped in the audit trail; how packs are assembled stays private. Full record available for verification on request.
03 / The panel
One forecast.
Five places to put the mass.
Six valid seats supplied five-bucket masses; two seats returned no valid table. The September meeting is shown seat by seat, followed by the three-cut aggregate.
Mean of the three 2025 cuts
Pooled five-bucket masses
Benchmark baseline
FAILED · NO VALID MASS TABLE
FAILED · NO VALID MASS TABLE
Method: add the six valid seats' five buckets with equal weight, then divide the CUT 25 bucket by all change mass (the four non-HOLD buckets). This page scores the same nine Fed meetings as Cases F and H; they are one body of evidence, not independent samples.
04 / Why this was rare
YES can be correct
and still be useless.
Size is invisible to binary scoring. A forecast can look perfect on move/no-move while assigning zero mass to the policy size that happened.
Binary scoring hides the miss.
Calling “a cut” does not distinguish 25 basis points from a double-size move.
Stored masses preserve the shape.
HOLD, ±25bp and ±50bp remain separate, so the forecast can be graded at the resolution anyone can hedge.
Basis points carry the value.
Direction is only the first decision. Position sizing depends on where the probability sits inside that direction.
05 / The resolution
The mass landed
on the right size.
Each 2025 cut is resolved against its archived FOMC statement. The three decision sentences and official links are quoted below.
Resolved · exact size
93.3%mean share of change probability on the true size across the three 2025 cuts.
This is a sizing result, not a binary “the Fed moved” claim. HOLD is excluded from each meeting's change-mass denominator.
Lower is better. Comparisons remain inside their own score family.
decided to lower the target range for the federal funds rate by 1/4 percentage point to 4 to 4-1/4 percentSOURCE ↗
decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-3/4 to 4 percentSOURCE ↗
decided to lower the target range for the federal funds rate by 1/4 percentage point to 3-1/2 to 3-3/4 percentSOURCE ↗
Calibration lab
Size the September move yourself.
Assume the Fed will move. Spread 100 points across CUT 50, CUT 25, HOLD, HIKE 25 and HIKE 50.
The controls always sum to 100. On reveal, your CUT 25 share is also normalized over change mass so it can be compared with the panel's 93.3% three-cut mean.
How we kept it honest
Clean rows only.
The blind spot stays visible.
Meetings any analyst could recite were excluded before scoring. The September 2024 double-size cut remains visible as a memory-contaminated mechanism exhibit, not a graded row.
Every parseable cell was unknown; 24 additional cells returned no parseable answer.
SIZEGRADE-1 was committed before the result set was graded; action space and recall screen were also pre-registered.
Dated official lines may be shown; how packs are assembled stays private.
No aggregate accuracy number is computed across the curated public cases.
Mechanism exhibit only: this meeting sits in the memory-contaminated backtest span and is not part of the graded clean nine. The other four slots are shown only to locate the true size.
Registered format
Spread probability over the actual action space: HOLD, ±25bp and ±50bp. Grade the stored masses at the realized size.
Technical receiptRegistration, score files and recall screen
- Registration
016bafc · 30ea729 · 8736742- Evidence example
fedcal-2025-09-17 · freeze 2025-09-03- Finding
docs/findings/13082026_memory_audit_judge_stage.md- Frozen anchor
config/policysim/fedcal1/fedcal1_anchors_2025.json- Action-space masses
case_e_seat_masses.json- Memory screen
fed_recall_grid.csv
Action space, recall screen and SIZEGRADE-1; each committed before its corresponding adoption or grading step.
The first 2025 cut, frozen 14 days before the meeting ended.
Action-space and recall-screen results.
Source named; three dated lines shown from a larger pack.
Nine meetings plus the separate September 2024 mechanism exhibit.
138 parseable cells unknown; 24 cells unparseable.
The separate public trail repository is being prepared. Verification materials are available on request.