The Radar is a grading instrument. A scenario or wargame maps a possibility space; the Radar grades which branch the measured world is selecting, on a stated cadence, against thresholds set before tracking begins. This paper documents the method as applied across FP1 assessments and readings: what a signal is, how its historical threshold is derived, how motion is scored, how evidence becomes a decision, and how the method itself can fail. This paper describes only the instrument. Scenario assessments issued to third parties are cited in the assessments themselves and are not reproduced here. Every rule here is written so a skeptical reader can check whether we followed it.

Executive Signal Summary

The Radar grades motion, not the space. Scenario exercises define what could happen. The Radar produces dated, falsifiable calls on whether specific scenario elements are materializing, and scores those calls against measured evidence at each sweep. The wargame maps the space; the Radar grades the motion.

Every threshold is set before tracking begins. A threshold is derived from a named historical reference class, published together with the call and its falsifier, and frozen for the tracking window. Nothing is graded after the fact, and nothing is re-thresholded mid-window.

The dotted ring is the tripwire. On a scenario-grading board, distance from center is proximity to the scenario as written. Threshold-crossing readings use the opposite convention and a different graphic; Section VI states the rule. The dotted ring marks the historical threshold: a marker outside it is a held call, a marker drifting inward is the thing to watch, and a marker crossing it is a tripped falsifier that forces a public revision.

Monitors gather; lenses interpret. Domain monitors log evidence to a shared ledger without interpreting it. Vera sources the reference class and signs off on every threshold value. Manticus red-teams against premature calls and maps signals to decisions. Darśan tests whether the historical analogue actually fits. A threshold no lens can defend does not ship, and where the lenses cannot agree the brief says whether the disagreement resolved by convergence, stands as documented dissent, or is waiting on evidence that does not yet exist.

The method itself carries falsification triggers. Section IX names four conditions under which the Radar's methodology, not merely an individual call, is judged to have failed and is restructured in public.

I. What the Radar Is and Is Not

The Radar produces signals: dated, confidence-tiered, falsifiable calls on whether specific elements of a scenario or transition dynamic are materializing, tracked against measured evidence on a stated cadence. Three distinctions are worth making explicit.

Not a forecast. The Radar does not assign probabilities to futures. It states what would have to be observed for a call to fail, then reports whether that observation occurred. The epistemics are closer to a pre-registered experiment than to a prediction market: the value produced is the discipline of the grading, not the cleverness of the guess.

Not a newsfeed. Events enter the Radar only through a signal they bear on. An event that moves no marker is commentary and belongs in the Briefings. The board changes only at sweeps, on evidence that Vera has sourced, so a reader can distinguish the instrument's motion from the news cycle's.

Not the wargame. The scenario being graded defines the possibility space and deserves its own credit for doing so; where FP1 grades an externally authored scenario, that work is cited by name in the assessment itself. The Radar's claim is narrower: given that space, here is which branch the evidence is selecting, and here is the receipt.

II. The Anatomy of a Signal

Every signal on the board is a record with six required fields. If any field is missing, the signal is not published.

FieldWhat it containsWorked example
DimensionThe axis of the scenario being gradedSubstrate: power, compute, and physical deployment capacity
CallThe dated, scoreable claimEnergization in the top US data-center markets stays beyond three years through 2027, foreclosing the scenario's 18-month full-economy deployment
Conviction tierHigh · Medium · Watch, set at publicationHigh
Historical thresholdThe reference-class boundary the evidence is tracked against (Section III)Energization > 3 years, held to 2027
FalsifierThe threshold restated as an observable eventInterconnection or transformer lead times compress materially inside two quarters
Sources & deskPrimary sources graded by Vera; owning desks namedSightline / Bloomberg · PwC · HSBC · Vera · Manticus

The falsifier and the threshold are the same object seen twice. The threshold is the boundary in the historical record; the falsifier is what crossing that boundary would look like in the news. Publishing both, in the same document as the call, is what makes the later grading auditable.

III. How a Historical Threshold Is Set

"Historical threshold" is a specific claim: the line on the board sits where the historical record, not our judgment of the moment, separates ordinary variation from regime change. The derivation has four steps, and each leaves an artifact a reader can inspect.

Step 1 · Vera

Name the reference class

Every dimension is assigned a historical population of comparable episodes before any threshold value is discussed. For a physical-substrate call, the reference class is large-load grid interconnection and transformer procurement timelines in US markets. For Labor, it is technology-driven displacement episodes read at the cohort level rather than in aggregate. The reference class is stated in the signal record; a threshold with no named reference class is an opinion wearing a costume.

Step 2 · Vera, red-teamed by Manticus

Set the value where the record breaks

The threshold value is the boundary at which the reference class historically stopped behaving like noise and started behaving like a regime change. It is always a value plus a duration ("compress materially inside two quarters"), never a bare number, because the historical record distinguishes a blip from a break by persistence. Manticus runs the red-team against premature placement: would this threshold have false-alarmed on past episodes that resolved as noise?

Step 3 · Darśan

Test the analogue

Darśan audits whether the historical analogue actually maps: same causal structure, not merely surface resemblance. Where the analogue is weak, the signal is either re-classed or explicitly downgraded. Some dimensions have thin histories by nature; Section IX states how those are handled.

Step 4 · Editorial

Freeze and publish

Reference class, threshold value, and falsifier are published together with the call, and the threshold is frozen for the tracking window. Re-derivation happens only at window close, in the changelog, with the old and new values shown side by side. A threshold moved mid-window voids the signal (Section IX, Trigger 1).

What "historical" rules out is as important as what it requires. A threshold may not be placed by intuition about the present, by what would make the board dramatic, or by where the consensus currently sits. If the reference class does not support a line, the dimension ships as Watch, with no threshold, and says so.

IV. A Threshold, Worked

The method's claim is that a reader can trace any published decision sentence back to an observation and to the transformation applied in between. This section discharges that claim on one reading rather than asserting it. It shows how a single threshold is derived and applied; it is not a list of what FP1 currently tracks. Every live threshold, with its falsifier, provable direction and grade date, is in the Register, which is the only place a current threshold should ever be looked up. Reading R‑013 (27 July 2026) is used because it contains the case the method most needs to survive: a source that was graded, published, and then downgraded after the fact.

The decision question. Whether a utility should preserve interconnection capacity for a data-centre commitment running past 2030. The holder is illustrative. What makes it a useful test is that the evidence arrives on clocks that differ by more than an order of magnitude: listed equities reprice in days, fab and transformer decisions resolve over years.

ObservationEvidence typeThreshold appliedContribution to the reading
Chip complex gives back more than a fifth of its value over three weeksMarket-derived measure; the data wrapper is not the sourcePrice movement alone may not rerate physical demand, at any magnitudeInsufficient. Logged, not decision-bearing
Hyperscaler capital guidance does not reverse with the marketPrimary record, self-reportedStated intention is not realized load; counts against collapse, never for confirmationContradicts a complete-collapse reading
Micron Q3 FY2026 operating resultsPrimary record, self-reportedCompany-reported operating evidence is not independent corroboration of a company-reported planSupports the unchanged assessment; does not pool with the row above
Federal Reserve projections, 17 June: 9 of 18 participants above the June midpointPrimary official recordFinancing cost moves independently of installed capacityNarrows financing latitude. The one dimension that moved
Reported HBM capacity deferralIndependent status not established; single apparent reporting chainA claim with one chain does not carry decision weightPublished at low confidence, then downgraded at Sweep 2

The downgrade is the receipt. On 7 August a primary company disclosure announced a large investment plan including future HBM capacity. That does not falsify the earlier report, and the method does not pretend it does. It weakens the broader retrenchment inference the reading leaned on, which is enough to remove the item from the decision chain. The row stays in the ledger, marked revised, with the date. The reading is not re‑dated and the conclusion that rested partly on it is restated without it. An instrument that quietly deletes a weak row, or silently improves a reading after the fact, has no record worth grading — which is why Section IX treats post‑hoc threshold movement as a voiding event rather than a correction.

What the four states produced. Unchanged on the demand case, because no admissible evidence bore on realized long-horizon load. Narrowed on financing latitude, on the strength of one official record. Revised on near-term memory capacity. Preserved on reversible interconnection and permitting optionality, subject to its carrying cost. Four dimensions, one movement, and the movement is the smallest claim the evidence supports rather than the most interesting one available.

The full ledger, including the rows that contributed nothing, is published with the reading at the worked example. Scenario‑grading applications of the same method, where the object is a named third‑party scenario rather than a decision question, follow the identical rules under a different board geometry; those assessments are issued to the parties concerned and are not published here.

V. Scoring: Mass, Velocity, Proximity

Between publication and falsification, a signal is scored on three components at each sweep.

Mass. How much evidence has converged: independent primary sources, not citations of citations. Vera grades source independence explicitly, because three articles quoting one report are one source. Higher mass makes a reading harder to dismiss; it never substitutes for the threshold.

Velocity. Rate and direction of change: accelerating, plateauing, or reverting. A fast move on thin mass is logged as a question, not a signal. Velocity is what separates "held · strengthened" from "held · watch" in the sweep verdicts.

Proximity. Distance to the historical threshold, shown on the board as ring position. Proximity is ordinal: it ranks how close the evidence sits to the line and which way it moved since the last sweep. It is not a probability, and the board never presents it as one.

The structure the scoring borrows

The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. In the formal model, action availability is conditioned on belief. The Radar does not inherit that step. Position on the board is an observed signal; which options remain open is a separate claim that has to be argued from the evidence, in writing, and it is made in the decision layer of a briefing rather than read off a marker. This is why the board is worth watching between verdicts, and also why the board alone is never the deliverable: the motion is the belief update, and the reasoning that carries an update into a changed choice set is the analysis.

What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section X is unambiguous.

VI. Reading the Board

The board's geometry encodes the method. The center is the scenario as written. Each marker sits at the distance the measured world is from that point on its dimension. The dotted ring is the historical threshold from Section III. Outward motion between sweeps means the gap to the scenario widened and the call strengthened; inward motion means the world drifted toward the scenario and the signal is watched harder. A marker crossing the dotted ring is a tripped falsifier: the call is scored as failed, publicly, at the sweep where it happened.

Each sweep is dated and appended to the same timeline, so the board can be scrubbed through its own history. The Delta board for a window shows every stop; nothing shown at a later sweep is permitted to alter an earlier one.

Two instruments, two geometries, never one graphic

FP1 publishes two kinds of measurement, and they read in opposite directions. Conflating them is the most damaging error available to this method, so the rule is stated here rather than left to the legend.

Type A

Scenario-grading board

Used when the object is a named scenario. The center is the scenario. Inward is approach: the world moving toward the scenario. The ring is the falsifier, and crossing it fails the standing call. This is the board described above.

Type B

Threshold-crossing reading

Used when the object is a measured dimension against a historical line, as in the published Hyperscaling reading. There is no scenario at the center. Outward is crossing: a dimension moving past the level that preceded prior unwinds. Crossing confirms the measurement rather than falsifying a call.

The rule. A single visual grammar may not carry both meanings. Every published board states its type, and a Type B reading does not use the Type A ring-and-spoke graphic. Where a Type B reading has previously shown both a threshold chart and a scope, the two are the same numbers twice and the scope is the redundant one.

And neither type shows decision impact. Position on either board is an observed signal. Section VII states the steps that have to be completed before a position becomes a claim about anyone's options.

VII. The Evidence-to-Decision Pipeline

Eight stages sit between an observation and a decision brief. Each is labelled by how it is produced, because a reader deciding whether to trust an output needs to know which parts a machine did and which parts a named person is accountable for.

Domain evidence Monitors sweep their assigned beats AUTOMATED 1 Claims and observations Extraction into discrete, dated claims AUTOMATED 2 Source and independence grading Quality and independence scored (Mass) RULE-BASED 3 Relationship mapping How claims bear on one another RULE-BASED 4 Pivotal uncertainty The question the reading turns on HUMAN JUDGMENT 5 Options opened or closed Effect on the decision holder's set HUMAN JUDGMENT 6 Decision brief Fixed seven-part format RULE-BASED 7 Human review and feedback Editorial sign-off; ledger re-grading HUMAN JUDGMENT 8
Automated — runs without a person in the loop Rule-based — deterministic, published rules a reader can check Human judgment — a named desk is accountable
StageWhat it producesHowWhat a person is accountable for
1 · Domain evidenceRaw items from a monitored beat, timestampedAutomatedChoosing the beat and the source set
2 · Claims and observationsDiscrete claims separated from commentaryAutomatedSpot-checking extraction against the source
3 · Source and independence gradingQuality grade, and whether a claim rests on one sourceRule-basedSetting the rules; any override, on the record
4 · Relationship mappingHow each claim bears on the othersRule-based HumanEvery relationship asserted, and its confidence grade
5 · Pivotal uncertaintyThe question the reading turns onHuman judgmentNaming it, and defending it against alternatives
6 · Options opened or closedEffect on the decision holder's option setHuman judgmentThe reasoning chain from evidence to option, in writing
7 · Decision briefThe seven-part briefing, same order every timeRule-basedEditorial sign-off before publication
8 · Human review and feedbackCorrections that re-grade earlier ledger entriesHuman judgmentRecording what was wrong and what changed

Stages 5 and 6 are the ones that cannot be automated away without losing the thing being sold. Everything above them is collection and grading, which is labour. Everything below them is format. The judgment in the middle is the product, and it is the reason a position on a board is not itself a decision brief.

Where the current method sits on this pipeline

Today, stages 1 through 8 are executed by desks working to the rules in this paper, with software assisting collection. Stage 3 is rule-based in the sense that the grading rules are published and applied consistently, not in the sense that a program applies them. A computational engine to automate stages 1 through 4 — evidence collection, claim extraction, source grading, and relationship mapping, with probabilistic updating over a graphical representation of the domain — is in development. It is not running behind anything currently published, and the distinction is restated in Section X.

VIII. Grading Discipline

Four rules govern the grading, and they are the substance of the claim "nothing was graded after the fact."

Pre-registration. Calls, tiers, thresholds, and falsifiers are published before the tracking window opens. The publication date is on the board.

Sweep cadence. Grading happens at stated intervals (monthly for scenario assessments), not when the news is convenient. An event between sweeps enters the record at the next sweep, dated to the event.

Verdict vocabulary. Sweep verdicts come from a closed set: held · strengthened, held, held · watch, tripped, retired. A closed vocabulary prevents the grading from softening its own failures with prose.

Desk separation. The lens that sourced a threshold (Vera) is not the lens that red-teams it (Manticus) or the lens that audits the analogue (Darśan). Gathering and interpretation are also separated: domain monitors log evidence to the ledger and do not interpret it, and the lenses interpret without owning a beat. Disagreement between lenses is recorded, not smoothed, and resolves in one of three states — convergence, documented dissent, or insufficient evidence — each named in the brief.

IX. Falsification of the Method

Individual calls fail by their falsifiers. The method fails by the following triggers, mirroring the standard set for FNC-1 in NCB-003: stated up front, so the reader knows what would force a restructuring.

Trigger 1

Threshold churn

Any threshold re-set inside its tracking window voids the pre-registration claim for that signal. The signal is retired, the change is documented in the changelog, and the retirement is shown on the board rather than deleted from it.

Trigger 2

Surprise trip

If a falsifier trips without the board having shown drift toward the line on that dimension across the preceding sweeps (inward on a Type A board, outward on Type B), the proximity model failed for that dimension class: the instrument was not measuring the approach it claims to measure. The dimension's scoring is rebuilt before any new signal ships on it.

Trigger 3

Reference-class failure

If consecutive scored windows show calls held while the world moves materially in ways the board's dimensions cannot represent (the Hé seam is the near-miss case: a real dynamic invisible to the original three-desk read), the dimension architecture, not just a threshold, is revised, and the revision is documented.

Trigger 4

Grading drift

If the desks, grading the same sweep evidence under the stated standard, produce materially different verdicts that editorial cannot resolve on the record, the editorial standard itself is judged underspecified and is restructured in public before the next sweep.

X. Stated Limitations

Reference-class selection is judgment. Two honest analysts can assign the same dimension to different historical populations. The method makes that judgment inspectable, not infallible: the class is named, so the reader can dispute it.

Some histories are thin. Agentic reliability has no long historical record; its threshold leans on a benchmark-defined bar (the Capricorn bar) rather than a base rate, and the signal record says so. Where a dimension's threshold is benchmark-anchored rather than history-anchored, it is labeled as such and carries at most Medium conviction.

Ring position is ordinal. The board shows ranking and direction, not calibrated probability. A marker twice as far from the ring is not twice as safe, and the Radar makes no such claim.

The Radar is not a POMDP. Section V borrows the structure of a partially observable Markov decision process; it does not implement the formalism. The Radar computes no belief distribution over states, runs no expected-free-energy calculation, and derives Proximity from a historical reference class rather than from a probability model. The correspondence is conceptual and is used to keep the scoring honest about what it represents: an evidence-updated position relative to a boundary, not a solved decision process. Any reading of the board as a calibrated POMDP would overstate what the instrument does.

What is published is not the engine. The instrument documented here is an editorially governed measurement and grading method: thresholds set by hand and defended in public, sources graded against published rules, relationships argued by named lenses. The computational engine described at the end of Section VII — automated collection and grading, relationship mapping, and probabilistic updating over a graphical model — is in development and is a separate object. The two can coexist, and the engine is intended to occupy stages 1 through 4 of the same pipeline. Nothing currently published should be read as its output, and this paper labels proposed architecture as proposed wherever it appears.

The grader grades itself. FP1 both places the thresholds and scores them. The mitigations are structural (desk separation, closed verdict vocabulary, public receipts), not a substitute for external audit. Readers who find a graded window they believe was scored generously are asked to say so, on the record.

Status. v0.3, August 2026. Documents the method of record for FP1 readings and scenario assessments. The Radar and scenario assessments reference this paper as the method of record; the FNC-1 instrument remains documented separately in NCB-003. Reference-class citations for each published threshold are being assembled into an appendix before v1.0.

Changelog. v0.3 — separated the two board geometries into Type A and Type B with an explicit prohibition on sharing one graphic (Section VI); removed the inference from board position to decision latitude and moved that claim into the briefing decision layer (Sections V, VI); added the evidence-to-decision pipeline with per-stage accountability (Section VII); separated domain monitors from analytical lenses and named the three disagreement-resolution states (Section VIII); added the limitation distinguishing the published instrument from the engine in development (Section X). Sections renumbered from VII onward. No change to thresholds, calls, or grading.

Changelog. v0.2 — added the POMDP lineage for the scoring model (Section V) and the corresponding limitation stating the Radar does not implement the formalism (Section IX of v0.2). No change to thresholds, calls, or grading.

Reference

1.  Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.