Executive Signal Summary
▶ The Radar grades motion, not the space. Scenario exercises define what could happen. The Radar produces dated, falsifiable calls on whether specific scenario elements are materializing, and scores those calls against measured evidence at each sweep. The wargame maps the space; the Radar grades the motion.
▶ Every threshold is set before tracking begins. A threshold is derived from a named historical reference class, published together with the call and its falsifier, and frozen for the tracking window. Nothing is graded after the fact, and nothing is re-thresholded mid-window.
▶ The dotted ring is the tripwire. On the board, distance from center is proximity to the scenario as written. The dotted ring marks the historical threshold: a marker outside it is a held call, a marker drifting inward is the thing to watch, and a marker crossing it is a tripped falsifier that forces a public revision.
▶ Three desks, one editorial standard. Vera sources the reference class and signs off on every threshold value. Manticus red-teams against premature calls and maps signals to decisions. Darśan tests whether the historical analogue actually fits. A threshold no desk can defend does not ship.
▶ The method itself carries falsification triggers. Section VIII names four conditions under which the Radar's methodology, not merely an individual call, is judged to have failed and is restructured in public.
I. What the Radar Is and Is Not
The Radar produces signals: dated, confidence-tiered, falsifiable calls on whether specific elements of a scenario or transition dynamic are materializing, tracked against measured evidence on a stated cadence. Three distinctions are worth making explicit.
Not a forecast. The Radar does not assign probabilities to futures. It states what would have to be observed for a call to fail, then reports whether that observation occurred. The epistemics are closer to a pre-registered experiment than to a prediction market: the value produced is the discipline of the grading, not the cleverness of the guess.
Not a newsfeed. Events enter the Radar only through a signal they bear on. An event that moves no marker is commentary and belongs in the Briefings. The board changes only at sweeps, on evidence that Vera has sourced, so a reader can distinguish the instrument's motion from the news cycle's.
Not the wargame. The scenario being graded (in SA-001, RAND's Infinite Potential) defines the possibility space and deserves its own credit for doing so. The Radar's claim is narrower: given that space, here is which branch the evidence is selecting, and here is the receipt.
II. The Anatomy of a Signal
Every signal on the board is a record with six required fields. If any field is missing, the signal is not published.
| Field | What it contains | SA-001 example (Substrate) |
|---|---|---|
| Dimension | The axis of the scenario being graded | Substrate: power, compute, and physical deployment capacity |
| Call | The dated, scoreable claim | Energization in the top US data-center markets stays beyond three years through 2027, foreclosing the scenario's 18-month full-economy deployment |
| Conviction tier | High · Medium · Watch, set at publication | High |
| Historical threshold | The reference-class boundary the evidence is tracked against (Section III) | Energization > 3 years, held to 2027 |
| Falsifier | The threshold restated as an observable event | Interconnection or transformer lead times compress materially inside two quarters |
| Sources & desk | Primary sources graded by Vera; owning desks named | Sightline / Bloomberg · PwC · HSBC · Vera · Manticus |
The falsifier and the threshold are the same object seen twice. The threshold is the boundary in the historical record; the falsifier is what crossing that boundary would look like in the news. Publishing both, in the same document as the call, is what makes the later grading auditable.
III. How a Historical Threshold Is Set
"Historical threshold" is a specific claim: the line on the board sits where the historical record, not our judgment of the moment, separates ordinary variation from regime change. The derivation has four steps, and each leaves an artifact a reader can inspect.
Name the reference class
Every dimension is assigned a historical population of comparable episodes before any threshold value is discussed. For the SA-001 Substrate call, the reference class is large-load grid interconnection and transformer procurement timelines in US markets. For Labor, it is technology-driven displacement episodes read at the cohort level rather than in aggregate. The reference class is stated in the signal record; a threshold with no named reference class is an opinion wearing a costume.
Set the value where the record breaks
The threshold value is the boundary at which the reference class historically stopped behaving like noise and started behaving like a regime change. It is always a value plus a duration ("compress materially inside two quarters"), never a bare number, because the historical record distinguishes a blip from a break by persistence. Manticus runs the red-team against premature placement: would this threshold have false-alarmed on past episodes that resolved as noise?
Test the analogue
Darśan audits whether the historical analogue actually maps: same causal structure, not merely surface resemblance. Where the analogue is weak, the signal is either re-classed or explicitly downgraded. Some dimensions have thin histories by nature; Section IX states how those are handled.
Freeze and publish
Reference class, threshold value, and falsifier are published together with the call, and the threshold is frozen for the tracking window. Re-derivation happens only at window close, in the changelog, with the old and new values shown side by side. A threshold moved mid-window voids the signal (Section VIII, Trigger 1).
What "historical" rules out is as important as what it requires. A threshold may not be placed by intuition about the present, by what would make the board dramatic, or by where the consensus currently sits. If the reference class does not support a line, the dimension ships as Watch, with no threshold, and says so.
IV. The SA-001 Thresholds, Worked
The five calls of SA-001, published 28 May 2026, each with the falsifier as stated on that date and the status after the first monthly sweep (29 June 2026). This table is the receipt the method promises; the full sweep record is on the Delta board.
| Dimension | Falsifier as published | Sweep 1 status |
|---|---|---|
| Substrate | Interconnection or transformer lead times compress materially inside two quarters | Not tripped; moved the opposite way. Held · strengthened |
| Agentic reliability | A credible public demonstration at the Capricorn bar | Not tripped; the one axis drifting toward the scenario. Held · watch |
| Labor | A sustained AI-attributable spike in headline unemployment | Not tripped; mechanism sharpened to reduced junior hiring. Held · strengthened |
| Crisis vector | A credibly autonomous, unattributable infrastructure attack | Not tripped; attribution intact, autonomy rising. Closest-watched |
| Governance | A containment crisis arriving before a distributional one | Not tripped. Held |
A sixth signal (Substrate Sovereignty, on the US–China seam) opened at Sweep 1 under the Hé desk. New desks open new signals; they do not rescore existing ones. Its falsifier, full indigenous substitution or full embargo, was stated at opening, in keeping with Step 4.
V. Scoring: Mass, Velocity, Proximity
Between publication and falsification, a signal is scored on three components at each sweep.
Mass. How much evidence has converged: independent primary sources, not citations of citations. Vera grades source independence explicitly, because three articles quoting one report are one source. Higher mass makes a reading harder to dismiss; it never substitutes for the threshold.
Velocity. Rate and direction of change: accelerating, plateauing, or reverting. A fast move on thin mass is logged as a question, not a signal. Velocity is what separates "held · strengthened" from "held · watch" in the sweep verdicts.
Proximity. Distance to the historical threshold, shown on the board as ring position. Proximity is ordinal: it ranks how close the evidence sits to the line and which way it moved since the last sweep. It is not a probability, and the board never presents it as one.
The structure the scoring borrows
The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. Decision latitude — which options remain open — is a function of that position, exactly as action availability is conditioned on belief in the formal model. This is why the board is worth watching between verdicts: the motion is the belief update, and the belief update is what changes the choice set.
What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section IX is unambiguous.
VI. Reading the Board
The board's geometry encodes the method. The center is the scenario as written. Each marker sits at the distance the measured world is from that point on its dimension. The dotted ring is the historical threshold from Section III. Outward motion between sweeps means the gap to the scenario widened and the call strengthened; inward motion means the world drifted toward the scenario and the signal is watched harder. A marker crossing the dotted ring is a tripped falsifier: the call is scored as failed, publicly, at the sweep where it happened.
Each sweep is dated and appended to the same timeline, so the board can be scrubbed through its own history. The Delta board for a window shows every stop; nothing shown at a later sweep is permitted to alter an earlier one.
VII. Grading Discipline
Four rules govern the grading, and they are the substance of the claim "nothing was graded after the fact."
Pre-registration. Calls, tiers, thresholds, and falsifiers are published before the tracking window opens. The publication date is on the board.
Sweep cadence. Grading happens at stated intervals (monthly for SA-001), not when the news is convenient. An event between sweeps enters the record at the next sweep, dated to the event.
Verdict vocabulary. Sweep verdicts come from a closed set: held · strengthened, held, held · watch, tripped, retired. A closed vocabulary prevents the grading from softening its own failures with prose.
Desk separation. The desk that sourced a threshold (Vera) is not the desk that red-teams it (Manticus) or the desk that audits the analogue (Darśan). Disagreement between desks is recorded, not smoothed.
VIII. Falsification of the Method
Individual calls fail by their falsifiers. The method fails by the following triggers, mirroring the standard set for FNC-1 in NCB-003: stated up front, so the reader knows what would force a restructuring.
Threshold churn
Any threshold re-set inside its tracking window voids the pre-registration claim for that signal. The signal is retired, the change is documented in the changelog, and the retirement is shown on the board rather than deleted from it.
Surprise trip
If a falsifier trips without the board having shown inward drift on that dimension across the preceding sweeps, the proximity model failed for that dimension class: the instrument was not measuring the approach it claims to measure. The dimension's scoring is rebuilt before any new signal ships on it.
Reference-class failure
If consecutive scored windows show calls held while the world moves materially in ways the board's dimensions cannot represent (the Hé seam in SA-001 is the near-miss case: a real dynamic invisible to the original three-desk read), the dimension architecture, not just a threshold, is revised, and the revision is documented.
Grading drift
If the desks, grading the same sweep evidence under the stated standard, produce materially different verdicts that editorial cannot resolve on the record, the editorial standard itself is judged underspecified and is restructured in public before the next sweep.
IX. Stated Limitations
Reference-class selection is judgment. Two honest analysts can assign the same dimension to different historical populations. The method makes that judgment inspectable, not infallible: the class is named, so the reader can dispute it.
Some histories are thin. Agentic reliability has no long historical record; its SA-001 threshold leans on a benchmark-defined bar (the Capricorn bar) rather than a base rate, and the signal record says so. Where a dimension's threshold is benchmark-anchored rather than history-anchored, it is labeled as such and carries at most Medium conviction.
Ring position is ordinal. The board shows ranking and direction, not calibrated probability. A marker twice as far from the ring is not twice as safe, and the Radar makes no such claim.
The Radar is not a POMDP. Section V borrows the structure of a partially observable Markov decision process; it does not implement the formalism. The Radar computes no belief distribution over states, runs no expected-free-energy calculation, and derives Proximity from a historical reference class rather than from a probability model. The correspondence is conceptual and is used to keep the scoring honest about what it represents: an evidence-updated position relative to a boundary, not a solved decision process. Any reading of the board as a calibrated POMDP would overstate what the instrument does.
The grader grades itself. FP1 both places the thresholds and scores them. The mitigations are structural (desk separation, closed verdict vocabulary, public receipts), not a substitute for external audit. Readers who find a graded window they believe was scored generously are asked to say so, on the record.
Changelog. v0.2 — added the POMDP lineage for the scoring model (Section V) and the corresponding limitation stating the Radar does not implement the formalism (Section IX). No change to thresholds, calls, or grading.
Reference
1. Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.