Executive Signal Summary
▶ The Radar grades motion, not the space. Scenario exercises define what could happen. The Radar produces dated, falsifiable calls on whether specific scenario elements are materializing, and scores those calls against measured evidence at each sweep. The wargame maps the space; the Radar grades the motion.
▶ Every threshold is set before tracking begins. A threshold is derived from a named historical reference class, published together with the call and its falsifier, and frozen for the tracking window. Nothing is graded after the fact, and nothing is re-thresholded mid-window.
▶ The dotted ring is the tripwire. On a scenario-grading board, distance from center is proximity to the scenario as written. Threshold-crossing readings use the opposite convention and a different graphic; Section VI states the rule. The dotted ring marks the historical threshold: a marker outside it is a held call, a marker drifting inward is the thing to watch, and a marker crossing it is a tripped falsifier that forces a public revision.
▶ Monitors gather; lenses interpret. Domain monitors log evidence to a shared ledger without interpreting it. Vera sources the reference class and signs off on every threshold value. Manticus red-teams against premature calls and maps signals to decisions. Darśan tests whether the historical analogue actually fits. A threshold no lens can defend does not ship, and where the lenses cannot agree the brief says whether the disagreement resolved by convergence, stands as documented dissent, or is waiting on evidence that does not yet exist.
▶ The method itself carries falsification triggers. Section IX names four conditions under which the Radar's methodology, not merely an individual call, is judged to have failed and is restructured in public.
I. What the Radar Is and Is Not
The Radar produces signals: dated, confidence-tiered, falsifiable calls on whether specific elements of a scenario or transition dynamic are materializing, tracked against measured evidence on a stated cadence. Three distinctions are worth making explicit.
Not a forecast. The Radar does not assign probabilities to futures. It states what would have to be observed for a call to fail, then reports whether that observation occurred. The epistemics are closer to a pre-registered experiment than to a prediction market: the value produced is the discipline of the grading, not the cleverness of the guess.
Not a newsfeed. Events enter the Radar only through a signal they bear on. An event that moves no marker is commentary and belongs in the Briefings. The board changes only at sweeps, on evidence that Vera has sourced, so a reader can distinguish the instrument's motion from the news cycle's.
Not the wargame. The scenario being graded defines the possibility space and deserves its own credit for doing so; where FP1 grades an externally authored scenario, that work is cited by name in the assessment itself. The Radar's claim is narrower: given that space, here is which branch the evidence is selecting, and here is the receipt.
II. The Anatomy of a Signal
Every signal on the board is a record with six required fields. If any field is missing, the signal is not published.
| Field | What it contains | Worked example |
|---|---|---|
| Dimension | The axis of the scenario being graded | Substrate: power, compute, and physical deployment capacity |
| Call | The dated, scoreable claim | Energization in the top US data-center markets stays beyond three years through 2027, foreclosing the scenario's 18-month full-economy deployment |
| Conviction tier | High · Medium · Watch, set at publication | High |
| Historical threshold | The reference-class boundary the evidence is tracked against (Section III) | Energization > 3 years, held to 2027 |
| Falsifier | The threshold restated as an observable event | Interconnection or transformer lead times compress materially inside two quarters |
| Sources & desk | Primary sources graded by Vera; owning desks named | Sightline / Bloomberg · PwC · HSBC · Vera · Manticus |
The falsifier and the threshold are the same object seen twice. The threshold is the boundary in the historical record; the falsifier is what crossing that boundary would look like in the news. Publishing both, in the same document as the call, is what makes the later grading auditable.
III. How a Historical Threshold Is Set
"Historical threshold" is a specific claim: the line on the board sits where the historical record, not our judgment of the moment, separates ordinary variation from regime change. The derivation has four steps, and each leaves an artifact a reader can inspect.
Name the reference class
Every dimension is assigned a historical population of comparable episodes before any threshold value is discussed. For a physical-substrate call, the reference class is large-load grid interconnection and transformer procurement timelines in US markets. For Labor, it is technology-driven displacement episodes read at the cohort level rather than in aggregate. The reference class is stated in the signal record; a threshold with no named reference class is an opinion wearing a costume.
Set the value where the record breaks
The threshold value is the boundary at which the reference class historically stopped behaving like noise and started behaving like a regime change. It is always a value plus a duration ("compress materially inside two quarters"), never a bare number, because the historical record distinguishes a blip from a break by persistence. Manticus runs the red-team against premature placement: would this threshold have false-alarmed on past episodes that resolved as noise?
Test the analogue
Darśan audits whether the historical analogue actually maps: same causal structure, not merely surface resemblance. Where the analogue is weak, the signal is either re-classed or explicitly downgraded. Some dimensions have thin histories by nature; Section IX states how those are handled.
Freeze and publish
Reference class, threshold value, and falsifier are published together with the call, and the threshold is frozen for the tracking window. Re-derivation happens only at window close, in the changelog, with the old and new values shown side by side. A threshold moved mid-window voids the signal (Section IX, Trigger 1).
What "historical" rules out is as important as what it requires. A threshold may not be placed by intuition about the present, by what would make the board dramatic, or by where the consensus currently sits. If the reference class does not support a line, the dimension ships as Watch, with no threshold, and says so.
IV. A Threshold, Worked
The method's claim is that a reader can trace any published decision sentence back to an observation and to the transformation applied in between. This section discharges that claim on one reading rather than asserting it. It shows how a single threshold is derived and applied; it is not a list of what FP1 currently tracks. Every live threshold, with its falsifier, provable direction and grade date, is in the Register, which is the only place a current threshold should ever be looked up. Reading R‑013 (27 July 2026) is used because it contains the case the method most needs to survive: a source that was graded, published, and then downgraded after the fact.
The decision question. Whether a utility should preserve interconnection capacity for a data-centre commitment running past 2030. The holder is illustrative. What makes it a useful test is that the evidence arrives on clocks that differ by more than an order of magnitude: listed equities reprice in days, fab and transformer decisions resolve over years.
| Observation | Evidence type | Threshold applied | Contribution to the reading |
|---|---|---|---|
| Chip complex gives back more than a fifth of its value over three weeks | Market-derived measure; the data wrapper is not the source | Price movement alone may not rerate physical demand, at any magnitude | Insufficient. Logged, not decision-bearing |
| Hyperscaler capital guidance does not reverse with the market | Primary record, self-reported | Stated intention is not realized load; counts against collapse, never for confirmation | Contradicts a complete-collapse reading |
| Micron Q3 FY2026 operating results | Primary record, self-reported | Company-reported operating evidence is not independent corroboration of a company-reported plan | Supports the unchanged assessment; does not pool with the row above |
| Federal Reserve projections, 17 June: 9 of 18 participants above the June midpoint | Primary official record | Financing cost moves independently of installed capacity | Narrows financing latitude. The one dimension that moved |
| Reported HBM capacity deferral | Independent status not established; single apparent reporting chain | A claim with one chain does not carry decision weight | Published at low confidence, then downgraded at Sweep 2 |
The downgrade is the receipt. On 7 August a primary company disclosure announced a large investment plan including future HBM capacity. That does not falsify the earlier report, and the method does not pretend it does. It weakens the broader retrenchment inference the reading leaned on, which is enough to remove the item from the decision chain. The row stays in the ledger, marked revised, with the date. The reading is not re‑dated and the conclusion that rested partly on it is restated without it. An instrument that quietly deletes a weak row, or silently improves a reading after the fact, has no record worth grading — which is why Section IX treats post‑hoc threshold movement as a voiding event rather than a correction.
What the four states produced. Unchanged on the demand case, because no admissible evidence bore on realized long-horizon load. Narrowed on financing latitude, on the strength of one official record. Revised on near-term memory capacity. Preserved on reversible interconnection and permitting optionality, subject to its carrying cost. Four dimensions, one movement, and the movement is the smallest claim the evidence supports rather than the most interesting one available.
The full ledger, including the rows that contributed nothing, is published with the reading at the worked example. Scenario‑grading applications of the same method, where the object is a named third‑party scenario rather than a decision question, follow the identical rules under a different board geometry; those assessments are issued to the parties concerned and are not published here.
V. Scoring: Mass, Velocity, Proximity
Between publication and falsification, a signal is scored on three components at each sweep.
Mass. How much evidence has converged: independent primary sources, not citations of citations. Vera grades source independence explicitly, because three articles quoting one report are one source. Higher mass makes a reading harder to dismiss; it never substitutes for the threshold.
Velocity. Rate and direction of change: accelerating, plateauing, or reverting. A fast move on thin mass is logged as a question, not a signal. Velocity is what separates "held · strengthened" from "held · watch" in the sweep verdicts.
Proximity. Distance to the historical threshold, shown on the board as ring position. Proximity is ordinal: it ranks how close the evidence sits to the line and which way it moved since the last sweep. It is not a probability, and the board never presents it as one.
The structure the scoring borrows
The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. In the formal model, action availability is conditioned on belief. The Radar does not inherit that step. Position on the board is an observed signal; which options remain open is a separate claim that has to be argued from the evidence, in writing, and it is made in the decision layer of a briefing rather than read off a marker. This is why the board is worth watching between verdicts, and also why the board alone is never the deliverable: the motion is the belief update, and the reasoning that carries an update into a changed choice set is the analysis.
What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section X is unambiguous.
VI. Reading the Board
The board's geometry encodes the method. The center is the scenario as written. Each marker sits at the distance the measured world is from that point on its dimension. The dotted ring is the historical threshold from Section III. Outward motion between sweeps means the gap to the scenario widened and the call strengthened; inward motion means the world drifted toward the scenario and the signal is watched harder. A marker crossing the dotted ring is a tripped falsifier: the call is scored as failed, publicly, at the sweep where it happened.
Each sweep is dated and appended to the same timeline, so the board can be scrubbed through its own history. The Delta board for a window shows every stop; nothing shown at a later sweep is permitted to alter an earlier one.
Two instruments, two geometries, never one graphic
FP1 publishes two kinds of measurement, and they read in opposite directions. Conflating them is the most damaging error available to this method, so the rule is stated here rather than left to the legend.
Scenario-grading board
Used when the object is a named scenario. The center is the scenario. Inward is approach: the world moving toward the scenario. The ring is the falsifier, and crossing it fails the standing call. This is the board described above.
Threshold-crossing reading
Used when the object is a measured dimension against a historical line, as in the published Hyperscaling reading. There is no scenario at the center. Outward is crossing: a dimension moving past the level that preceded prior unwinds. Crossing confirms the measurement rather than falsifying a call.
The rule. A single visual grammar may not carry both meanings. Every published board states its type, and a Type B reading does not use the Type A ring-and-spoke graphic. Where a Type B reading has previously shown both a threshold chart and a scope, the two are the same numbers twice and the scope is the redundant one.
And neither type shows decision impact. Position on either board is an observed signal. Section VII states the steps that have to be completed before a position becomes a claim about anyone's options.
VII. The Evidence-to-Decision Pipeline
Eight stages sit between an observation and a decision brief. Each is labelled by how it is produced, because a reader deciding whether to trust an output needs to know which parts a machine did and which parts a named person is accountable for.
| Stage | What it produces | How | What a person is accountable for |
|---|---|---|---|
| 1 · Domain evidence | Raw items from a monitored beat, timestamped | Automated | Choosing the beat and the source set |
| 2 · Claims and observations | Discrete claims separated from commentary | Automated | Spot-checking extraction against the source |
| 3 · Source and independence grading | Quality grade, and whether a claim rests on one source | Rule-based | Setting the rules; any override, on the record |
| 4 · Relationship mapping | How each claim bears on the others | Rule-based Human | Every relationship asserted, and its confidence grade |
| 5 · Pivotal uncertainty | The question the reading turns on | Human judgment | Naming it, and defending it against alternatives |
| 6 · Options opened or closed | Effect on the decision holder's option set | Human judgment | The reasoning chain from evidence to option, in writing |
| 7 · Decision brief | The seven-part briefing, same order every time | Rule-based | Editorial sign-off before publication |
| 8 · Human review and feedback | Corrections that re-grade earlier ledger entries | Human judgment | Recording what was wrong and what changed |
Stages 5 and 6 are the ones that cannot be automated away without losing the thing being sold. Everything above them is collection and grading, which is labour. Everything below them is format. The judgment in the middle is the product, and it is the reason a position on a board is not itself a decision brief.
Where the current method sits on this pipeline
Today, stages 1 through 8 are executed by desks working to the rules in this paper, with software assisting collection. Stage 3 is rule-based in the sense that the grading rules are published and applied consistently, not in the sense that a program applies them. A computational engine to automate stages 1 through 4 — evidence collection, claim extraction, source grading, and relationship mapping, with probabilistic updating over a graphical representation of the domain — is in development. It is not running behind anything currently published, and the distinction is restated in Section X.
VIII. Grading Discipline
Four rules govern the grading, and they are the substance of the claim "nothing was graded after the fact."
Pre-registration. Calls, tiers, thresholds, and falsifiers are published before the tracking window opens. The publication date is on the board.
Sweep cadence. Grading happens at stated intervals (monthly for scenario assessments), not when the news is convenient. An event between sweeps enters the record at the next sweep, dated to the event.
Verdict vocabulary. Sweep verdicts come from a closed set: held · strengthened, held, held · watch, tripped, retired. A closed vocabulary prevents the grading from softening its own failures with prose.
Desk separation. The lens that sourced a threshold (Vera) is not the lens that red-teams it (Manticus) or the lens that audits the analogue (Darśan). Gathering and interpretation are also separated: domain monitors log evidence to the ledger and do not interpret it, and the lenses interpret without owning a beat. Disagreement between lenses is recorded, not smoothed, and resolves in one of three states — convergence, documented dissent, or insufficient evidence — each named in the brief.
IX. Falsification of the Method
Individual calls fail by their falsifiers. The method fails by the following triggers, mirroring the standard set for FNC-1 in NCB-003: stated up front, so the reader knows what would force a restructuring.
Threshold churn
Any threshold re-set inside its tracking window voids the pre-registration claim for that signal. The signal is retired, the change is documented in the changelog, and the retirement is shown on the board rather than deleted from it.
Surprise trip
If a falsifier trips without the board having shown drift toward the line on that dimension across the preceding sweeps (inward on a Type A board, outward on Type B), the proximity model failed for that dimension class: the instrument was not measuring the approach it claims to measure. The dimension's scoring is rebuilt before any new signal ships on it.
Reference-class failure
If consecutive scored windows show calls held while the world moves materially in ways the board's dimensions cannot represent (the Hé seam is the near-miss case: a real dynamic invisible to the original three-desk read), the dimension architecture, not just a threshold, is revised, and the revision is documented.
Grading drift
If the desks, grading the same sweep evidence under the stated standard, produce materially different verdicts that editorial cannot resolve on the record, the editorial standard itself is judged underspecified and is restructured in public before the next sweep.
X. Stated Limitations
Reference-class selection is judgment. Two honest analysts can assign the same dimension to different historical populations. The method makes that judgment inspectable, not infallible: the class is named, so the reader can dispute it.
Some histories are thin. Agentic reliability has no long historical record; its threshold leans on a benchmark-defined bar (the Capricorn bar) rather than a base rate, and the signal record says so. Where a dimension's threshold is benchmark-anchored rather than history-anchored, it is labeled as such and carries at most Medium conviction.
Ring position is ordinal. The board shows ranking and direction, not calibrated probability. A marker twice as far from the ring is not twice as safe, and the Radar makes no such claim.
The Radar is not a POMDP. Section V borrows the structure of a partially observable Markov decision process; it does not implement the formalism. The Radar computes no belief distribution over states, runs no expected-free-energy calculation, and derives Proximity from a historical reference class rather than from a probability model. The correspondence is conceptual and is used to keep the scoring honest about what it represents: an evidence-updated position relative to a boundary, not a solved decision process. Any reading of the board as a calibrated POMDP would overstate what the instrument does.
What is published is not the engine. The instrument documented here is an editorially governed measurement and grading method: thresholds set by hand and defended in public, sources graded against published rules, relationships argued by named lenses. The computational engine described at the end of Section VII — automated collection and grading, relationship mapping, and probabilistic updating over a graphical model — is in development and is a separate object. The two can coexist, and the engine is intended to occupy stages 1 through 4 of the same pipeline. Nothing currently published should be read as its output, and this paper labels proposed architecture as proposed wherever it appears.
The grader grades itself. FP1 both places the thresholds and scores them. The mitigations are structural (desk separation, closed verdict vocabulary, public receipts), not a substitute for external audit. Readers who find a graded window they believe was scored generously are asked to say so, on the record.
Changelog. v0.3 — separated the two board geometries into Type A and Type B with an explicit prohibition on sharing one graphic (Section VI); removed the inference from board position to decision latitude and moved that claim into the briefing decision layer (Sections V, VI); added the evidence-to-decision pipeline with per-stage accountability (Section VII); separated domain monitors from analytical lenses and named the three disagreement-resolution states (Section VIII); added the limitation distinguishing the published instrument from the engine in development (Section X). Sections renumbered from VII onward. No change to thresholds, calls, or grading.
Changelog. v0.2 — added the POMDP lineage for the scoring model (Section V) and the corresponding limitation stating the Radar does not implement the formalism (Section IX of v0.2). No change to thresholds, calls, or grading.
Reference
1. Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.