Source to Claim Record

See the evidence behind the conclusion.

FP1 Evidence turns source material into discrete claims, checks them against cited sources, preserves uncertainty, and records what would overturn each assessment. The current Chrome side-panel beta works with captioned YouTube videos and pasted transcripts. The same Claim Record structure supports FP1's Radar, Register, and decision briefs.

Working beta · v0.7.0 Local Chrome loading · not a store release
Inspectable evidence object

How a Claim Record is graded.

Everything in the system reduces to this object. A claim, the sources behind it, whether those sources are actually independent of each other, an assessment drawn from a closed vocabulary, and the evidence that would change the assessment. Every downstream instrument renders records of this shape.

Claim record · FP1-EV-0001
Source class: widely repeated public claim
“A ChatGPT query uses about ten times the electricity of a Google search.”
In circulation since late 2023. Checked 4 September 2026.
Where the number comes from

The ratio is not a measurement. It is one 2023 estimate divided by one 2009 figure, and almost every later appearance restates that division rather than repeating it.

2009
Denominator. Google publishes roughly 0.3 Wh for a search, attributed to Urs Hölzle. Pre-dates the smartphone era and Google’s ownership of YouTube. Never updated.
Oct 2023
Numerator. Alex de Vries, Joule 7(10), estimates roughly 3 Wh per query, derived in part from a third-party estimate of ChatGPT’s daily request volume and server fleet.
2023–24
Restatement. The ratio is carried into an EPRI figure, a Goldman Sachs report, IEA commentary and general press. Each cites an earlier restatement, not an independent measurement.
Feb 2025
First independent recalculation. Epoch AI, using the same method with updated hardware and volume assumptions, puts a typical GPT-4o text query near 0.3 Wh, roughly ten times below the 2023 estimate.
Jun 2025
Operator statement. OpenAI states an average query near 0.34 Wh. Self-reported, no published methodology.
Aug 2025
First production measurement. Google publishes a technical paper (arXiv:2508.15734) reporting a median Gemini text prompt at 0.24 Wh, measured from in-production telemetry including host CPU, idle capacity and cooling overhead. Not peer reviewed. It measures a chatbot, not Search.
Evidence
  • Google, Measuring the environmental impact of delivering AI at Google Scale (arXiv:2508.15734)Primary · operator telemetryB · not peer reviewed
  • de Vries, The growing energy footprint of artificial intelligence, Joule 7(10), 2191–2194Primary · peer-reviewed commentaryA · estimate, not measurement
  • Epoch AI, energy-per-query recalculation, February 2025Independent of both operatorsB
  • OpenAI, statement on average query energy, June 2025Self-reportedC · no methodology
  • Google, 2009 blog figure for a searchPrimary but supersededWithdrawn from use
  • EPRI figure, Goldman Sachs report, IEA commentary, general pressNot independent · restate the two aboveExcluded

Dozens of citations. Three independent measurements or estimates, none of which measures both sides of the ratio under one method.

Assessment
Unsupported as stated

Not false, and not merely outdated. The ratio cannot currently be computed. The numerator has moved by roughly a factor of ten since 2023 and is now measured by two operators using different methods. The denominator has not been published by anyone since 2009, so no current figure for a Google search exists to divide into it. A claim whose denominator is unmeasured is unsupported regardless of how many outlets carry it.

The two operator figures that do exist (0.24 Wh and 0.34 Wh) are for chatbot prompts, not for search, and are not comparable to each other: only one publishes its methodology, and only one includes idle capacity and cooling overhead.

Falsifier

A current per-query figure for Google Search, measured under the same boundary as the Gemini paper (accelerator, host CPU, provisioned idle capacity and data-centre overhead), published alongside a comparable figure for an LLM query. That would make the ratio computable and would move this record to supported or false depending on the result.

What this record does not say

Nothing here bears on whether AI energy demand in aggregate is large or growing. That is a different claim, resting on different evidence, and it is tracked separately. Retiring a bad ratio is not the same as answering the question the ratio was being used to settle.

Real record, checked 4 September 2026 against the cited sources. Grades apply to the source as evidence for this specific claim, not to the source in general. Machine-readable: claims.json.

Running beta architecture

Claims, Context, Integrity, and Analysis.

Chrome side-panel beta · v0.7.0

The panel opens beside whatever you are reading or watching. Claims is the evidence. Context and Integrity describe the information environment around it. Analysis applies the same lenses used in FP1's published work. Nothing here filters or blocks anything.

FP1 Evidence · working beta · Chrome side panel

Checkable claims, extracted and timestamped. Each one can be checked on demand against live sources, and every check returns its citations. Claims that cannot be checked are marked as such rather than assessed anyway.

00:14:22
Capacity claim
Assessed: context required. Two independent primary sources, conflicting capacity definitions.
00:22:07
Cost-per-unit claim
Assessed: supported. Three independent sources agree within the stated tolerance.
00:41:55
Attribution of a policy decision
Assessed: contested. Primary sources disagree, and the disagreement is preserved rather than resolved.
00:58:10
Forward projection
Not checkable. A statement about the future carries no evidence yet, so it is logged and left ungraded.

Faithful web replica of the v0.7.0 extension source. The extension itself must still be loaded locally for functional testing. Analysis is model-generated and only as reliable as the transcript and sources, which is why timestamps and sources remain inspectable.

Processing order

Verification remains separate from interpretation.

Context and Integrity are annotations on the evidence environment. They operate before the analytical lenses, not alongside them, and they are kept separate from the claim assessment so that neither can quietly contaminate the other.

Layer 1
Extract

Source text and a stable source identity. Claims separated from the commentary around them, each one timestamped against the original.

Layer 2
Verify

Live checking against sources. Quality graded, independence graded separately, an assessment from the closed set, and a stated falsifier.

Layer 3 · two passes, one layer
Context

Framing, one-sided sourcing, unanswered counterarguments, and the moments that were fair.

Integrity

Inspectable patterns in presentation, omission, and source use. The beta makes no longitudinal claim and never scores a person or their intent.

Layer 4
Analyse

Strategy, historical orientation and synthesis, applied to the completed record rather than to the headline. Translation runs automatically over the output.

Context and Integrity occupy the same layer and run as two separate passes. Neither writes to the claim assessment, and the assessment does not feed either of them, so a source that reads badly cannot quietly downgrade a well-sourced claim and a well-sourced claim cannot launder a manipulative source.

Integrity does not rate people or sources

The obvious product here is a rating that tells people which sources to believe. FP1 does not build that, and will not.

A rating would move judgment from the reader to the tool without establishing whether a claim is true. The Integrity pass instead shows the passage, describes the communication pattern, and explains why it may affect interpretation.

Honest limits

  • Captions or text are required. No transcript, no analysis.
  • Automatic captions mis-hear names and numbers, and assessments inherit those errors. Auto-caption sources are flagged.
  • Integrity operates on one transcript. It does not establish recurring behaviour or motive.
  • Context and Integrity findings are model-generated observations, not truth, bias, or credibility scores.
  • Live fact-checks require the tester's Anthropic key and network access.
Downstream use

Claim Records support the Radar, Register, and decision brief.

One video is the smallest case. The same claim records, graded the same way, are what the Radar reads and what the Register grades. The evidence layer is not a separate product from the instruments; it is the layer they stand on.

Evidence → threshold → change → decision. The same four moves at every scale, from one sentence in a transcript to a monthly reading. Nothing enters a brief that cannot be traced back along this line to a graded claim record.