Executive Signal Summary
▶ The Radar grades motion, not the space. Scenario exercises define what could happen. The Radar produces dated, falsifiable calls on whether specific scenario elements are materializing, and scores those calls against measured evidence at each sweep. The wargame maps the space; the Radar grades the motion.
▶ Every threshold is set before tracking begins. A threshold is derived from a named historical reference class, published together with the call and its falsifier, and frozen for the tracking window. Nothing is graded after the fact, and nothing is re-thresholded mid-window.
▶ The dotted ring is the tripwire. On a scenario-grading board, distance from center is proximity to the scenario as written. Threshold-crossing readings use the opposite convention and a different graphic; Section VI states the rule. The dotted ring marks the historical threshold: a marker outside it is a held call, a marker drifting inward is the thing to watch, and a marker crossing it is a tripped falsifier that forces a public revision.
▶ Monitors gather; lenses interpret. Domain monitors log evidence to a shared ledger without interpreting it. Vera sources the reference class and signs off on every threshold value. Manticus red-teams against premature calls and maps signals to decisions. Darśan tests whether the historical analogue actually fits. A threshold no lens can defend does not ship, and where the lenses cannot agree the brief says whether the disagreement resolved by convergence, stands as documented dissent, or is waiting on evidence that does not yet exist.
▶ The method itself carries falsification triggers. Section IX names four conditions under which the Radar's methodology, not merely an individual call, is judged to have failed and is restructured in public.
I. What the Radar Is and Is Not
The Radar produces signals: dated, confidence-tiered, falsifiable calls on whether specific elements of a scenario or transition dynamic are materializing, tracked against measured evidence on a stated cadence. Three distinctions are worth making explicit.
Not a forecast. The Radar does not assign probabilities to futures. It states what would have to be observed for a call to fail, then reports whether that observation occurred. The epistemics are closer to a pre-registered experiment than to a prediction market: the value produced is the discipline of the grading, not the cleverness of the guess.
Not a newsfeed. Events enter the Radar only through a signal they bear on. An event that moves no marker is commentary and belongs in the Briefings. The board changes only at sweeps, on evidence that Vera has sourced, so a reader can distinguish the instrument's motion from the news cycle's.
Not the wargame. The scenario being graded defines the possibility space and deserves its own credit for doing so; where FP1 grades an externally authored scenario, that work is cited by name in the assessment itself. The Radar's claim is narrower: given that space, here is which branch the evidence is selecting, and here is the receipt.
II. The Anatomy of a Signal
Every signal on the board is a record with six required fields. If any field is missing, the signal is not published.
| Field | What it contains | SA-001 example (Substrate) |
|---|---|---|
| Dimension | The axis of the scenario being graded | Substrate: power, compute, and physical deployment capacity |
| Call | The dated, scoreable claim | Energization in the top US data-center markets stays beyond three years through 2027, foreclosing the scenario's 18-month full-economy deployment |
| Conviction tier | High · Medium · Watch, set at publication | High |
| Historical threshold | The reference-class boundary the evidence is tracked against (Section III) | Energization > 3 years, held to 2027 |
| Falsifier | The threshold restated as an observable event | Interconnection or transformer lead times compress materially inside two quarters |
| Sources & desk | Primary sources graded by Vera; owning desks named | Sightline / Bloomberg · PwC · HSBC · Vera · Manticus |
The falsifier and the threshold are the same object seen twice. The threshold is the boundary in the historical record; the falsifier is what crossing that boundary would look like in the news. Publishing both, in the same document as the call, is what makes the later grading auditable.
III. How a Historical Threshold Is Set
"Historical threshold" is a specific claim: the line on the board sits where the historical record, not our judgment of the moment, separates ordinary variation from regime change. The derivation has four steps, and each leaves an artifact a reader can inspect.
Name the reference class
Every dimension is assigned a historical population of comparable episodes before any threshold value is discussed. For the SA-001 Substrate call, the reference class is large-load grid interconnection and transformer procurement timelines in US markets. For Labor, it is technology-driven displacement episodes read at the cohort level rather than in aggregate. The reference class is stated in the signal record; a threshold with no named reference class is an opinion wearing a costume.
Set the value where the record breaks
The threshold value is the boundary at which the reference class historically stopped behaving like noise and started behaving like a regime change. It is always a value plus a duration ("compress materially inside two quarters"), never a bare number, because the historical record distinguishes a blip from a break by persistence. Manticus runs the red-team against premature placement: would this threshold have false-alarmed on past episodes that resolved as noise?
Test the analogue
Darśan audits whether the historical analogue actually maps: same causal structure, not merely surface resemblance. Where the analogue is weak, the signal is either re-classed or explicitly downgraded. Some dimensions have thin histories by nature; Section IX states how those are handled.
Freeze and publish
Reference class, threshold value, and falsifier are published together with the call, and the threshold is frozen for the tracking window. Re-derivation happens only at window close, in the changelog, with the old and new values shown side by side. A threshold moved mid-window voids the signal (Section IX, Trigger 1).
What "historical" rules out is as important as what it requires. A threshold may not be placed by intuition about the present, by what would make the board dramatic, or by where the consensus currently sits. If the reference class does not support a line, the dimension ships as Watch, with no threshold, and says so.
IV. The SA-001 Thresholds, Worked
The five calls of SA-001, published 28 May 2026, each with the falsifier as stated on that date and the status after the first monthly sweep (29 June 2026). This table is the receipt the method promises; the full sweep record is on the Delta board.
| Dimension | Falsifier as published | Sweep 1 status |
|---|---|---|
| Substrate | Interconnection or transformer lead times compress materially inside two quarters | Not tripped; moved the opposite way. Held · strengthened |
| Agentic reliability | A credible public demonstration at the Capricorn bar | Not tripped; the one axis drifting toward the scenario. Held · watch |
| Labor | A sustained AI-attributable spike in headline unemployment | Not tripped; mechanism sharpened to reduced junior hiring. Held · strengthened |
| Crisis vector | A credibly autonomous, unattributable infrastructure attack | Not tripped; attribution intact, autonomy rising. Closest-watched |
| Governance | A containment crisis arriving before a distributional one | Not tripped. Held |
A sixth signal (Substrate Sovereignty, on the US–China seam) opened at Sweep 1 under the Hé desk. New desks open new signals; they do not rescore existing ones. Its falsifier, full indigenous substitution or full embargo, was stated at opening, in keeping with Step 4.
V. Scoring: Mass, Velocity, Proximity
Between publication and falsification, a signal is scored on three components at each sweep.
Mass. How much evidence has converged: independent primary sources, not citations of citations. Vera grades source independence explicitly, because three articles quoting one report are one source. Higher mass makes a reading harder to dismiss; it never substitutes for the threshold.
Velocity. Rate and direction of change: accelerating, plateauing, or reverting. A fast move on thin mass is logged as a question, not a signal. Velocity is what separates "held · strengthened" from "held · watch" in the sweep verdicts.
Proximity. Distance to the historical threshold, shown on the board as ring position. Proximity is ordinal: it ranks how close the evidence sits to the line and which way it moved since the last sweep. It is not a probability, and the board never presents it as one.
The structure the scoring borrows
The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. In the formal model, action availability is conditioned on belief. The Radar does not inherit that step. Position on the board is an observed signal; which options remain open is a separate claim that has to be argued from the evidence, in writing, and it is made in the decision layer of a briefing rather than read off a marker. This is why the board is worth watching between verdicts, and also why the board alone is never the deliverable: the motion is the belief update, and the reasoning that carries an update into a changed choice set is the analysis.
What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section X is unambiguous.
VI. Reading the Board
The board's geometry encodes the method. The center is the scenario as written. Each marker sits at the distance the measured world is from that point on its dimension. The dotted ring is the historical threshold from Section III. Outward motion between sweeps means the gap to the scenario widened and the call strengthened; inward motion means the world drifted toward the scenario and the signal is watched harder. A marker crossing the dotted ring is a tripped falsifier: the call is scored as failed, publicly, at the sweep where it happened.
Each sweep is dated and appended to the same timeline, so the board can be scrubbed through its own history. The Delta board for a window shows every stop; nothing shown at a later sweep is permitted to alter an earlier one.
Two instruments, two geometries, never one graphic
FP1 publishes two kinds of measurement, and they read in opposite directions. Conflating them is the most damaging error available to this method, so the rule is stated here rather than left to the legend.
Scenario-grading board
Used when the object is a named scenario, as in SA-001. The center is the scenario. Inward is approach: the world moving toward the scenario. The ring is the falsifier, and crossing it fails the standing call. This is the board described above.
Threshold-crossing reading
Used when the object is a measured dimension against a historical line, as in the published Hyperscaling reading. There is no scenario at the center. Outward is crossing: a dimension moving past the level that preceded prior unwinds. Crossing confirms the measurement rather than falsifying a call.
The rule. A single visual grammar may not carry both meanings. Every published board states its type, and a Type B reading does not use the Type A ring-and-spoke graphic. Where a Type B reading has previously shown both a threshold chart and a scope, the two are the same numbers twice and the scope is the redundant one.
And neither type shows decision impact. Position on either board is an observed signal. Section VII states the steps that have to be completed before a position becomes a claim about anyone's options.
VII. The Evidence-to-Decision Pipeline
Eight stages sit between an observation and a decision brief. Each is labelled by how it is produced, because a reader deciding whether to trust an output needs to know which parts a machine did and which parts a named person is accountable for.
| Stage | What it produces | How | What a person is accountable for |
|---|---|---|---|
| 1 · Domain evidence | Raw items from a monitored beat, timestamped | Automated | Choosing the beat and the source set |
| 2 · Claims and observations | Discrete claims separated from commentary | Automated | Spot-checking extraction against the source |
| 3 · Source and independence grading | Quality grade, and whether a claim rests on one source | Rule-based | Setting the rules; any override, on the record |
| 4 · Relationship mapping | How each claim bears on the others | Rule-based Human | Every relationship asserted, and its confidence grade |
| 5 · Pivotal uncertainty | The question the reading turns on | Human judgment | Naming it, and defending it against alternatives |
| 6 · Options opened or closed | Effect on the decision holder's option set | Human judgment | The reasoning chain from evidence to option, in writing |
| 7 · Decision brief | The seven-part briefing, same order every time | Rule-based | Editorial sign-off before publication |
| 8 · Human review and feedback | Corrections that re-grade earlier ledger entries | Human judgment | Recording what was wrong and what changed |
Stages 5 and 6 are the ones that cannot be automated away without losing the thing being sold. Everything above them is collection and grading, which is labour. Everything below them is format. The judgment in the middle is the product, and it is the reason a position on a board is not itself a decision brief.
Where the current method sits on this pipeline
Today, stages 1 through 8 are executed by desks working to the rules in this paper, with software assisting collection. Stage 3 is rule-based in the sense that the grading rules are published and applied consistently, not in the sense that a program applies them. A computational engine to automate stages 1 through 4 — evidence collection, claim extraction, source grading, and relationship mapping, with probabilistic updating over a graphical representation of the domain — is in development. It is not running behind anything currently published, and the distinction is restated in Section X.
VIII. Grading Discipline
Four rules govern the grading, and they are the substance of the claim "nothing was graded after the fact."
Pre-registration. Calls, tiers, thresholds, and falsifiers are published before the tracking window opens. The publication date is on the board.
Sweep cadence. Grading happens at stated intervals (monthly for SA-001), not when the news is convenient. An event between sweeps enters the record at the next sweep, dated to the event.
Verdict vocabulary. Sweep verdicts come from a closed set: held · strengthened, held, held · watch, tripped, retired. A closed vocabulary prevents the grading from softening its own failures with prose.
Desk separation. The lens that sourced a threshold (Vera) is not the lens that red-teams it (Manticus) or the lens that audits the analogue (Darśan). Gathering and interpretation are also separated: domain monitors log evidence to the ledger and do not interpret it, and the lenses interpret without owning a beat. Disagreement between lenses is recorded, not smoothed, and resolves in one of three states — convergence, documented dissent, or insufficient evidence — each named in the brief.
IX. Falsification of the Method
Individual calls fail by their falsifiers. The method fails by the following triggers, mirroring the standard set for FNC-1 in NCB-003: stated up front, so the reader knows what would force a restructuring.
Threshold churn
Any threshold re-set inside its tracking window voids the pre-registration claim for that signal. The signal is retired, the change is documented in the changelog, and the retirement is shown on the board rather than deleted from it.
Surprise trip
If a falsifier trips without the board having shown drift toward the line on that dimension across the preceding sweeps (inward on a Type A board, outward on Type B), the proximity model failed for that dimension class: the instrument was not measuring the approach it claims to measure. The dimension's scoring is rebuilt before any new signal ships on it.
Reference-class failure
If consecutive scored windows show calls held while the world moves materially in ways the board's dimensions cannot represent (the Hé seam in SA-001 is the near-miss case: a real dynamic invisible to the original three-desk read), the dimension architecture, not just a threshold, is revised, and the revision is documented.
Grading drift
If the desks, grading the same sweep evidence under the stated standard, produce materially different verdicts that editorial cannot resolve on the record, the editorial standard itself is judged underspecified and is restructured in public before the next sweep.
X. Stated Limitations
Reference-class selection is judgment. Two honest analysts can assign the same dimension to different historical populations. The method makes that judgment inspectable, not infallible: the class is named, so the reader can dispute it.
Some histories are thin. Agentic reliability has no long historical record; its SA-001 threshold leans on a benchmark-defined bar (the Capricorn bar) rather than a base rate, and the signal record says so. Where a dimension's threshold is benchmark-anchored rather than history-anchored, it is labeled as such and carries at most Medium conviction.
Ring position is ordinal. The board shows ranking and direction, not calibrated probability. A marker twice as far from the ring is not twice as safe, and the Radar makes no such claim.
The Radar is not a POMDP. Section V borrows the structure of a partially observable Markov decision process; it does not implement the formalism. The Radar computes no belief distribution over states, runs no expected-free-energy calculation, and derives Proximity from a historical reference class rather than from a probability model. The correspondence is conceptual and is used to keep the scoring honest about what it represents: an evidence-updated position relative to a boundary, not a solved decision process. Any reading of the board as a calibrated POMDP would overstate what the instrument does.
What is published is not the engine. The instrument documented here is an editorially governed measurement and grading method: thresholds set by hand and defended in public, sources graded against published rules, relationships argued by named lenses. The computational engine described at the end of Section VII — automated collection and grading, relationship mapping, and probabilistic updating over a graphical model — is in development and is a separate object. The two can coexist, and the engine is intended to occupy stages 1 through 4 of the same pipeline. Nothing currently published should be read as its output, and this paper labels proposed architecture as proposed wherever it appears.
The grader grades itself. FP1 both places the thresholds and scores them. The mitigations are structural (desk separation, closed verdict vocabulary, public receipts), not a substitute for external audit. Readers who find a graded window they believe was scored generously are asked to say so, on the record.
Changelog. v0.3 — separated the two board geometries into Type A and Type B with an explicit prohibition on sharing one graphic (Section VI); removed the inference from board position to decision latitude and moved that claim into the briefing decision layer (Sections V, VI); added the evidence-to-decision pipeline with per-stage accountability (Section VII); separated domain monitors from analytical lenses and named the three disagreement-resolution states (Section VIII); added the limitation distinguishing the published instrument from the engine in development (Section X). Sections renumbered from VII onward. No change to thresholds, calls, or grading.
Changelog. v0.2 — added the POMDP lineage for the scoring model (Section V) and the corresponding limitation stating the Radar does not implement the formalism (Section IX of v0.2). No change to thresholds, calls, or grading.
Reference
1. Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.