The Radar is a grading instrument. A scenario or wargame maps a possibility space; the Radar grades which branch the measured world is selecting, on a stated cadence, against thresholds set before tracking begins. This paper documents the method as applied in Scenario Assessment SA-001: what a signal is, how its historical threshold is derived, how motion is scored, how evidence becomes a decision, and how the method itself can fail. The scenario graded in SA-001 is an externally published exercise, credited in full in SA-001 itself; this paper describes only the instrument. Every rule here is written so a skeptical reader can check whether we followed it.

Executive Signal Summary

The Radar grades motion, not the space. Scenario exercises define what could happen. The Radar produces dated, falsifiable calls on whether specific scenario elements are materializing, and scores those calls against measured evidence at each sweep. The wargame maps the space; the Radar grades the motion.

Every threshold is set before tracking begins. A threshold is derived from a named historical reference class, published together with the call and its falsifier, and frozen for the tracking window. Nothing is graded after the fact, and nothing is re-thresholded mid-window.

The dotted ring is the tripwire. On a scenario-grading board, distance from center is proximity to the scenario as written. Threshold-crossing readings use the opposite convention and a different graphic; Section VI states the rule. The dotted ring marks the historical threshold: a marker outside it is a held call, a marker drifting inward is the thing to watch, and a marker crossing it is a tripped falsifier that forces a public revision.

Monitors gather; lenses interpret. Domain monitors log evidence to a shared ledger without interpreting it. Vera sources the reference class and signs off on every threshold value. Manticus red-teams against premature calls and maps signals to decisions. Darśan tests whether the historical analogue actually fits. A threshold no lens can defend does not ship, and where the lenses cannot agree the brief says whether the disagreement resolved by convergence, stands as documented dissent, or is waiting on evidence that does not yet exist.

The method itself carries falsification triggers. Section IX names four conditions under which the Radar's methodology, not merely an individual call, is judged to have failed and is restructured in public.

I. What the Radar Is and Is Not

The Radar produces signals: dated, confidence-tiered, falsifiable calls on whether specific elements of a scenario or transition dynamic are materializing, tracked against measured evidence on a stated cadence. Three distinctions are worth making explicit.

Not a forecast. The Radar does not assign probabilities to futures. It states what would have to be observed for a call to fail, then reports whether that observation occurred. The epistemics are closer to a pre-registered experiment than to a prediction market: the value produced is the discipline of the grading, not the cleverness of the guess.

Not a newsfeed. Events enter the Radar only through a signal they bear on. An event that moves no marker is commentary and belongs in the Briefings. The board changes only at sweeps, on evidence that Vera has sourced, so a reader can distinguish the instrument's motion from the news cycle's.

Not the wargame. The scenario being graded defines the possibility space and deserves its own credit for doing so; where FP1 grades an externally authored scenario, that work is cited by name in the assessment itself. The Radar's claim is narrower: given that space, here is which branch the evidence is selecting, and here is the receipt.

II. The Anatomy of a Signal

Every signal on the board is a record with six required fields. If any field is missing, the signal is not published.

FieldWhat it containsSA-001 example (Substrate)
DimensionThe axis of the scenario being gradedSubstrate: power, compute, and physical deployment capacity
CallThe dated, scoreable claimEnergization in the top US data-center markets stays beyond three years through 2027, foreclosing the scenario's 18-month full-economy deployment
Conviction tierHigh · Medium · Watch, set at publicationHigh
Historical thresholdThe reference-class boundary the evidence is tracked against (Section III)Energization > 3 years, held to 2027
FalsifierThe threshold restated as an observable eventInterconnection or transformer lead times compress materially inside two quarters
Sources & deskPrimary sources graded by Vera; owning desks namedSightline / Bloomberg · PwC · HSBC · Vera · Manticus

The falsifier and the threshold are the same object seen twice. The threshold is the boundary in the historical record; the falsifier is what crossing that boundary would look like in the news. Publishing both, in the same document as the call, is what makes the later grading auditable.

III. How a Historical Threshold Is Set

"Historical threshold" is a specific claim: the line on the board sits where the historical record, not our judgment of the moment, separates ordinary variation from regime change. The derivation has four steps, and each leaves an artifact a reader can inspect.

Step 1 · Vera

Name the reference class

Every dimension is assigned a historical population of comparable episodes before any threshold value is discussed. For the SA-001 Substrate call, the reference class is large-load grid interconnection and transformer procurement timelines in US markets. For Labor, it is technology-driven displacement episodes read at the cohort level rather than in aggregate. The reference class is stated in the signal record; a threshold with no named reference class is an opinion wearing a costume.

Step 2 · Vera, red-teamed by Manticus

Set the value where the record breaks

The threshold value is the boundary at which the reference class historically stopped behaving like noise and started behaving like a regime change. It is always a value plus a duration ("compress materially inside two quarters"), never a bare number, because the historical record distinguishes a blip from a break by persistence. Manticus runs the red-team against premature placement: would this threshold have false-alarmed on past episodes that resolved as noise?

Step 3 · Darśan

Test the analogue

Darśan audits whether the historical analogue actually maps: same causal structure, not merely surface resemblance. Where the analogue is weak, the signal is either re-classed or explicitly downgraded. Some dimensions have thin histories by nature; Section IX states how those are handled.

Step 4 · Editorial

Freeze and publish

Reference class, threshold value, and falsifier are published together with the call, and the threshold is frozen for the tracking window. Re-derivation happens only at window close, in the changelog, with the old and new values shown side by side. A threshold moved mid-window voids the signal (Section IX, Trigger 1).

What "historical" rules out is as important as what it requires. A threshold may not be placed by intuition about the present, by what would make the board dramatic, or by where the consensus currently sits. If the reference class does not support a line, the dimension ships as Watch, with no threshold, and says so.

IV. The SA-001 Thresholds, Worked

The five calls of SA-001, published 28 May 2026, each with the falsifier as stated on that date and the status after the first monthly sweep (29 June 2026). This table is the receipt the method promises; the full sweep record is on the Delta board.

DimensionFalsifier as publishedSweep 1 status
SubstrateInterconnection or transformer lead times compress materially inside two quartersNot tripped; moved the opposite way. Held · strengthened
Agentic reliabilityA credible public demonstration at the Capricorn barNot tripped; the one axis drifting toward the scenario. Held · watch
LaborA sustained AI-attributable spike in headline unemploymentNot tripped; mechanism sharpened to reduced junior hiring. Held · strengthened
Crisis vectorA credibly autonomous, unattributable infrastructure attackNot tripped; attribution intact, autonomy rising. Closest-watched
GovernanceA containment crisis arriving before a distributional oneNot tripped. Held

A sixth signal (Substrate Sovereignty, on the US–China seam) opened at Sweep 1 under the Hé desk. New desks open new signals; they do not rescore existing ones. Its falsifier, full indigenous substitution or full embargo, was stated at opening, in keeping with Step 4.

V. Scoring: Mass, Velocity, Proximity

Between publication and falsification, a signal is scored on three components at each sweep.

Mass. How much evidence has converged: independent primary sources, not citations of citations. Vera grades source independence explicitly, because three articles quoting one report are one source. Higher mass makes a reading harder to dismiss; it never substitutes for the threshold.

Velocity. Rate and direction of change: accelerating, plateauing, or reverting. A fast move on thin mass is logged as a question, not a signal. Velocity is what separates "held · strengthened" from "held · watch" in the sweep verdicts.

Proximity. Distance to the historical threshold, shown on the board as ring position. Proximity is ordinal: it ranks how close the evidence sits to the line and which way it moved since the last sweep. It is not a probability, and the board never presents it as one.

The structure the scoring borrows

The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. In the formal model, action availability is conditioned on belief. The Radar does not inherit that step. Position on the board is an observed signal; which options remain open is a separate claim that has to be argued from the evidence, in writing, and it is made in the decision layer of a briefing rather than read off a marker. This is why the board is worth watching between verdicts, and also why the board alone is never the deliverable: the motion is the belief update, and the reasoning that carries an update into a changed choice set is the analysis.

What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section X is unambiguous.

VI. Reading the Board

The board's geometry encodes the method. The center is the scenario as written. Each marker sits at the distance the measured world is from that point on its dimension. The dotted ring is the historical threshold from Section III. Outward motion between sweeps means the gap to the scenario widened and the call strengthened; inward motion means the world drifted toward the scenario and the signal is watched harder. A marker crossing the dotted ring is a tripped falsifier: the call is scored as failed, publicly, at the sweep where it happened.

Each sweep is dated and appended to the same timeline, so the board can be scrubbed through its own history. The Delta board for a window shows every stop; nothing shown at a later sweep is permitted to alter an earlier one.

Two instruments, two geometries, never one graphic

FP1 publishes two kinds of measurement, and they read in opposite directions. Conflating them is the most damaging error available to this method, so the rule is stated here rather than left to the legend.

Type A

Scenario-grading board

Used when the object is a named scenario, as in SA-001. The center is the scenario. Inward is approach: the world moving toward the scenario. The ring is the falsifier, and crossing it fails the standing call. This is the board described above.

Type B

Threshold-crossing reading

Used when the object is a measured dimension against a historical line, as in the published Hyperscaling reading. There is no scenario at the center. Outward is crossing: a dimension moving past the level that preceded prior unwinds. Crossing confirms the measurement rather than falsifying a call.

The rule. A single visual grammar may not carry both meanings. Every published board states its type, and a Type B reading does not use the Type A ring-and-spoke graphic. Where a Type B reading has previously shown both a threshold chart and a scope, the two are the same numbers twice and the scope is the redundant one.

And neither type shows decision impact. Position on either board is an observed signal. Section VII states the steps that have to be completed before a position becomes a claim about anyone's options.

VII. The Evidence-to-Decision Pipeline

Eight stages sit between an observation and a decision brief. Each is labelled by how it is produced, because a reader deciding whether to trust an output needs to know which parts a machine did and which parts a named person is accountable for.

Domain evidence Monitors sweep their assigned beats AUTOMATED 1 Claims and observations Extraction into discrete, dated claims AUTOMATED 2 Source and independence grading Quality and independence scored (Mass) RULE-BASED 3 Relationship mapping How claims bear on one another RULE-BASED 4 Pivotal uncertainty The question the reading turns on HUMAN JUDGMENT 5 Options opened or closed Effect on the decision holder's set HUMAN JUDGMENT 6 Decision brief Fixed seven-part format RULE-BASED 7 Human review and feedback Editorial sign-off; ledger re-grading HUMAN JUDGMENT 8
Automated — runs without a person in the loop Rule-based — deterministic, published rules a reader can check Human judgment — a named desk is accountable
StageWhat it producesHowWhat a person is accountable for
1 · Domain evidenceRaw items from a monitored beat, timestampedAutomatedChoosing the beat and the source set
2 · Claims and observationsDiscrete claims separated from commentaryAutomatedSpot-checking extraction against the source
3 · Source and independence gradingQuality grade, and whether a claim rests on one sourceRule-basedSetting the rules; any override, on the record
4 · Relationship mappingHow each claim bears on the othersRule-based HumanEvery relationship asserted, and its confidence grade
5 · Pivotal uncertaintyThe question the reading turns onHuman judgmentNaming it, and defending it against alternatives
6 · Options opened or closedEffect on the decision holder's option setHuman judgmentThe reasoning chain from evidence to option, in writing
7 · Decision briefThe seven-part briefing, same order every timeRule-basedEditorial sign-off before publication
8 · Human review and feedbackCorrections that re-grade earlier ledger entriesHuman judgmentRecording what was wrong and what changed

Stages 5 and 6 are the ones that cannot be automated away without losing the thing being sold. Everything above them is collection and grading, which is labour. Everything below them is format. The judgment in the middle is the product, and it is the reason a position on a board is not itself a decision brief.

Where the current method sits on this pipeline

Today, stages 1 through 8 are executed by desks working to the rules in this paper, with software assisting collection. Stage 3 is rule-based in the sense that the grading rules are published and applied consistently, not in the sense that a program applies them. A computational engine to automate stages 1 through 4 — evidence collection, claim extraction, source grading, and relationship mapping, with probabilistic updating over a graphical representation of the domain — is in development. It is not running behind anything currently published, and the distinction is restated in Section X.

VIII. Grading Discipline

Four rules govern the grading, and they are the substance of the claim "nothing was graded after the fact."

Pre-registration. Calls, tiers, thresholds, and falsifiers are published before the tracking window opens. The publication date is on the board.

Sweep cadence. Grading happens at stated intervals (monthly for SA-001), not when the news is convenient. An event between sweeps enters the record at the next sweep, dated to the event.

Verdict vocabulary. Sweep verdicts come from a closed set: held · strengthened, held, held · watch, tripped, retired. A closed vocabulary prevents the grading from softening its own failures with prose.

Desk separation. The lens that sourced a threshold (Vera) is not the lens that red-teams it (Manticus) or the lens that audits the analogue (Darśan). Gathering and interpretation are also separated: domain monitors log evidence to the ledger and do not interpret it, and the lenses interpret without owning a beat. Disagreement between lenses is recorded, not smoothed, and resolves in one of three states — convergence, documented dissent, or insufficient evidence — each named in the brief.

IX. Falsification of the Method

Individual calls fail by their falsifiers. The method fails by the following triggers, mirroring the standard set for FNC-1 in NCB-003: stated up front, so the reader knows what would force a restructuring.

Trigger 1

Threshold churn

Any threshold re-set inside its tracking window voids the pre-registration claim for that signal. The signal is retired, the change is documented in the changelog, and the retirement is shown on the board rather than deleted from it.

Trigger 2

Surprise trip

If a falsifier trips without the board having shown drift toward the line on that dimension across the preceding sweeps (inward on a Type A board, outward on Type B), the proximity model failed for that dimension class: the instrument was not measuring the approach it claims to measure. The dimension's scoring is rebuilt before any new signal ships on it.

Trigger 3

Reference-class failure

If consecutive scored windows show calls held while the world moves materially in ways the board's dimensions cannot represent (the Hé seam in SA-001 is the near-miss case: a real dynamic invisible to the original three-desk read), the dimension architecture, not just a threshold, is revised, and the revision is documented.

Trigger 4

Grading drift

If the desks, grading the same sweep evidence under the stated standard, produce materially different verdicts that editorial cannot resolve on the record, the editorial standard itself is judged underspecified and is restructured in public before the next sweep.

X. Stated Limitations

Reference-class selection is judgment. Two honest analysts can assign the same dimension to different historical populations. The method makes that judgment inspectable, not infallible: the class is named, so the reader can dispute it.

Some histories are thin. Agentic reliability has no long historical record; its SA-001 threshold leans on a benchmark-defined bar (the Capricorn bar) rather than a base rate, and the signal record says so. Where a dimension's threshold is benchmark-anchored rather than history-anchored, it is labeled as such and carries at most Medium conviction.

Ring position is ordinal. The board shows ranking and direction, not calibrated probability. A marker twice as far from the ring is not twice as safe, and the Radar makes no such claim.

The Radar is not a POMDP. Section V borrows the structure of a partially observable Markov decision process; it does not implement the formalism. The Radar computes no belief distribution over states, runs no expected-free-energy calculation, and derives Proximity from a historical reference class rather than from a probability model. The correspondence is conceptual and is used to keep the scoring honest about what it represents: an evidence-updated position relative to a boundary, not a solved decision process. Any reading of the board as a calibrated POMDP would overstate what the instrument does.

What is published is not the engine. The instrument documented here is an editorially governed measurement and grading method: thresholds set by hand and defended in public, sources graded against published rules, relationships argued by named lenses. The computational engine described at the end of Section VII — automated collection and grading, relationship mapping, and probabilistic updating over a graphical model — is in development and is a separate object. The two can coexist, and the engine is intended to occupy stages 1 through 4 of the same pipeline. Nothing currently published should be read as its output, and this paper labels proposed architecture as proposed wherever it appears.

The grader grades itself. FP1 both places the thresholds and scores them. The mitigations are structural (desk separation, closed verdict vocabulary, public receipts), not a substitute for external audit. Readers who find a graded window they believe was scored generously are asked to say so, on the record.

Status. v0.3, August 2026. Documents the method as applied in SA-001 (published 28 May 2026; first sweep 29 June 2026). The Radar and scenario-assessment pages reference this paper as the method of record; the FNC-1 instrument remains documented separately in NCB-003. Reference-class citations for each SA-001 threshold are being assembled into an appendix before v1.0.

Changelog. v0.3 — separated the two board geometries into Type A and Type B with an explicit prohibition on sharing one graphic (Section VI); removed the inference from board position to decision latitude and moved that claim into the briefing decision layer (Sections V, VI); added the evidence-to-decision pipeline with per-stage accountability (Section VII); separated domain monitors from analytical lenses and named the three disagreement-resolution states (Section VIII); added the limitation distinguishing the published instrument from the engine in development (Section X). Sections renumbered from VII onward. No change to thresholds, calls, or grading.

Changelog. v0.2 — added the POMDP lineage for the scoring model (Section V) and the corresponding limitation stating the Radar does not implement the formalism (Section IX of v0.2). No change to thresholds, calls, or grading.

Reference

1.  Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.