← Research & methods

Methodology paper / NCB-004 · v0.3

The Radar: Signal scoring and historical thresholds

How FP1 checks an assessment against new evidence, records mistakes and decides when to change its view.

In brief

Before we start

Write down the assessment, what we will measure, and what result would prove it wrong.

As evidence arrives

Check how reliable the evidence is, what has changed, and whether the agreed limit has been reached.

When we review it

Publish the result, explain any disagreement, and name the person responsible for the assessment.

About this paper

The Radar checks published assessments against new evidence. This paper explains what is recorded, how the checks work, and when the method needs to change. This reader edition, dated 17 September 2026, simplifies the v0.3 paper and replaces exercise-specific examples with an example from FP1’s public Register. It does not change the Register’s rules or results.

The method in brief

Publish an assessment and the evidence behind it. Say in advance what would prove it wrong. Review it on the agreed date, using the same rule. Keep the result and any correction visible.

People run this process. Software helps collect and organize information. A named human reviewer is responsible for what FP1 publishes.

What the Radar does

The Radar follows specific questions over time. For each one, FP1 records its assessment, the evidence, a review date and the result that would require a change of view.

It does not calculate the probability of a future event. It checks whether the evidence has met a stated condition. A news event matters to the board only if it changes one of the questions being tracked.

Where a review starts from someone else’s scenario, the original work must be credited and available for readers to inspect. Private exercise details are not part of this public edition.

What each assessment must include

A chart marker is the short version of a fuller record. That record needs six things:

FieldWhat a reader should find
QuestionThe specific change we are tracking.
AssessmentWhat FP1 thinks, and the date it said so.
ConfidenceHow strongly the evidence supports that view, with its limits.
ThresholdThe level or condition that would change the assessment, including how long it must last.
What would prove it wrongAn observable result, expressed plainly enough for someone else to check.
Sources and responsibilityWhere the evidence came from and who reviewed the conclusion.

A threshold is a boundary. The test explains what crossing that boundary would look like in practice. Both must be published before the result is known.

How a threshold is chosen

A historical threshold needs a relevant historical comparison. It cannot be chosen because it makes today’s chart look dramatic.

1. Choose the comparison

Vera identifies comparable past cases and their sources. A study of power constraints might examine earlier grid-connection delays; a study of job losses needs comparable workers and industries.

2. Set the level and duration

Specify the change that would matter and how long it must last. Manticus checks whether the same rule would have produced false alarms in the earlier cases.

3. Test whether the history fits

Darśan examines the causes behind the comparison. Similar-looking charts are not enough. If the comparison is weak, lower the confidence or find a better one.

4. Publish the rule and keep it fixed

Publish the comparison, the threshold and the test together. Keep them unchanged for the period being assessed. Any later revision must show the old and new rules, with the reason for the change.

If the evidence cannot support a threshold, label the question “Watch” and explain what is missing.

An example you can check

The Register shows two rules and their recorded August 2026 results. These public examples replace the exercise-specific table in the earlier web edition.

MeasureRule set before the resultRecorded result
FNC-1 / SOX spreadCloses to −40, or widens beyond −120−68.745. Neither condition was met.
Belief IndexTwo consecutive weekly readings above 55, or any one below 4057.490, with one week above 55. The two-week condition was not met.

The result was published five days late. That delay is part of the record. The calculation correction in NCB-003 is also relevant when interpreting the FNC-1 results.

Evidence, change and distance to the limit

The paper uses three terms for what reviewers assess:

Mass: how much independent evidence?

Three articles quoting one report are still one source. Vera checks the quality and independence of the underlying evidence. More supporting evidence does not replace the agreed test.

Velocity: how quickly is it changing?

Is the trend speeding up, slowing down or reversing? A sharp move supported by little evidence is a question to investigate, not a firm conclusion.

Proximity: how close to the threshold?

This records how near the evidence is to the agreed limit, and whether it has moved closer since the last review. It is a relative position, not a calculated probability.

Background: the connection to active inference

The structure the scoring borrows

The three components are a deliberate, informal borrowing from the structure of a partially observable Markov decision process (POMDP), the formulation of active inference set out by Smith, Friston, and Whyte.1 In that structure an agent cannot see the underlying state of the world directly; it maintains a belief about that state, updates the belief as observations arrive, and — the part that matters here — finds that which actions are available to it depends on which belief state it is in. The Radar maps onto this one-to-one at the level of structure: the scenario is the state that cannot be observed directly, evidence accumulating at each sweep is the stream of observations, Mass and Velocity are how the reading is updated, and Proximity is where the updated belief sits relative to the boundary. In the formal model, action availability is conditioned on belief. The Radar does not inherit that step. Position on the board is an observed signal; which options remain open is a separate claim that has to be argued from the evidence, in writing, and it is made in the decision layer of a briefing rather than read off a marker. This is why the board is worth watching between verdicts, and also why the board alone is never the deliverable: the motion is the belief update, and the reasoning that carries an update into a changed choice set is the analysis.

What the Radar borrows is the shape of the problem, not the mathematics. It is stated here so the scoring is not mistaken for intuition dressed as instrumentation, and so the departure in Section X is unambiguous.

How to read the charts

FP1 has used two kinds of chart. Their directions mean different things, so each chart must explain its own scale.

Comparing the world with a scenario

The centre represents the scenario being examined. Movement inward means the evidence is moving closer to it. A marked boundary shows where the original assessment would need to change.

Comparing a measure with a threshold

The chart shows a measure against a stated limit. Movement beyond the limit means it has crossed that threshold. There is no imagined scenario at the centre.

Dates, units and a clear legend are essential. Neither chart tells a reader what decision to make: that requires a separate explanation of the evidence and the available choices.

The earlier SA-001 comparison depended on exercise-specific criteria. Its numeric display and pass/fail grades are not reproduced in the current public overview, because those criteria are not available for readers to check.

From evidence to a reviewed brief

Eight steps connect a source to a published brief. The labels show where software helps, where written rules guide the work, and where a person must make and defend a judgment.

Evidence → judgment → review

Eight steps, with responsibility at each one.

  1. 01

    Gather the evidence

    Choose the subject and collect the sources.

    Software-assisted
  2. 02

    Separate the claims

    Record each claim and when it was made.

    Software-assisted
  3. 03

    Check the sources

    Assess quality and whether sources are independent.

    Rule-guided
  4. 04

    Connect the evidence

    Show which claims support or challenge one another.

    Rule-guided
  5. 05

    Find the key uncertainty

    Identify the question that could change the answer.

    Accountable judgment
  6. 06

    Explain the choices

    Say what the evidence means for the decision.

    Accountable judgment
  7. 07

    Write the brief

    Use the same seven-part format.

    Rule-guided
  8. 08

    Review and correct

    Name the reviewer and record any corrections.

    Accountable judgment

Review feeds back into source grading. Corrections update the earlier record.

Software-assisted — current retrieval, extraction, or formatting support Rule-guided — declared rules govern the step Accountable judgment — a named reviewer owns the conclusion
Detailed stage outputs and accountability
StageWhat it producesCurrent boundaryWhat remains accountable
1 · Domain evidenceRaw items from a monitored beat, timestampedSoftware-assistedChoosing the beat and accepting the source set
2 · Claims and observationsDiscrete claims separated from commentarySoftware-assistedChecking extraction against the source
3 · Source and independence gradingQuality grade, and whether a claim rests on one sourceRule-guidedSetting the rules; any override, on the record
4 · Relationship mappingHow each claim bears on the othersRule-guidedEvery relationship asserted, and its confidence grade
5 · Pivotal uncertaintyThe question the reading turns onAccountable judgmentNaming it, and defending it against alternatives
6 · Options opened or closedEffect on the decision holder's option setAccountable judgmentThe reasoning chain from evidence to option, in writing
7 · Decision briefThe seven-part briefing, same order every timeRule-guidedSign-off before publication
8 · Review, correction, and regradingCorrections that re-grade earlier ledger entriesAccountable judgmentRecording what was wrong and what changed

The key uncertainty and its effect on a decision require human judgment. The brief must explain that reasoning and name who is responsible for it.

Where the current method sits on this pipeline

People currently run the full process, with software support for collecting information. An engine to assist the first four stages is in development. It is not the engine behind the assessments already published.

How reviews are recorded

Four rules make a later review possible:

  • Set the rule first. Publish the assessment, confidence, threshold and test before the period being reviewed.
  • Keep the review date. Review at the stated interval. Record events between reviews with their actual dates.
  • Use the same result labels. The original labels are “held · strengthened,” “held,” “held · watch,” “tripped” and “retired.” They mean stronger support, unchanged, closer to failing, failed, or withdrawn. A failed assessment must not be softened into a success through wording.
  • Separate the checks. Vera checks sources; Manticus challenges the reasoning; Darśan tests historical comparisons. Hé examines competing explanations. Rāwī explains the findings. If a disagreement remains, publish it.

When the method needs to change

An assessment can be wrong. The process for checking it can also fail. These are the warning signs:

The rule changes after the start

A threshold changed during the review period invalidates that assessment’s claim to have been set in advance. Withdraw the assessment and record why.

The chart misses the approach

If a condition is met without the chart having shown movement toward it, review how that question was scored before using the same approach again.

The questions miss important changes

If major changes repeatedly fall outside the categories being tracked, revise the categories. Changing one threshold will not solve that problem.

Reviewers cannot apply the rules consistently

If reviewers reach materially different grades from the same evidence and the rules cannot settle the difference, make the rules clearer before the next review.

What this method cannot establish

Historical comparisons involve judgment. Reasonable reviewers may choose different past cases. Naming those cases lets readers challenge the choice.

Some subjects have little history. Reliable AI agents are one example. A benchmark may help, but it must be named and available for readers to inspect. A weak historical record limits the confidence of the assessment.

Chart position is not probability. A marker twice as far from a line is not “twice as safe.” The chart shows relative position and direction.

The published method is not a mathematical inference engine. It draws ideas from active inference, but it does not calculate beliefs about hidden states or expected free energy. People review the evidence and write the reasoning.

FP1 is checking its own work. Separate review roles and visible corrections help, but they do not replace independent scrutiny.

Edition note · 17 September 2026.

Revised for clarity. Exercise-specific criteria and grades are omitted; the worked example now uses the public Register. Earlier editions are retained in FP1’s archive.

Reference

1.  Smith, R., Friston, K. J., & Whyte, C. J. (2022). A step-by-step tutorial on active inference and its application to empirical data. Journal of Mathematical Psychology, 107, 102632.