← All insights
Published · August 17, 2026

Same number, three levels of reliability: evidence provenance in carbon data

Two factories report the same figure for the same period: 1,240 tonnes CO₂e.

In the first, that number is the sum of hourly readings from a meter fitted to the gas line. In the second, an engineer reviewed the invoices and typed the total into a spreadsheet.

The two numbers are identical. Their reliability is not. And most reporting today shows no trace of that difference.

”Verified” does not say what you think it says

There are two separate questions in carbon data, and they are routinely confused:

First question: who checked this number? Nobody, the company’s own team, or an independent verifier? The industry has established practice on this axis; the assurance statements on reports speak to it.

Second question: where did this number come from? Did an instrument measure it, was it read off a source document, or did someone type it?

These two are independent — and that is the whole point. A number that has passed independent verification may rest entirely on hand-entered data. The verifier examines process, checks methodological conformance, samples. It does not descend to every line and confirm that each figure came off a device. Nor is it expected to, because the trail usually does not exist.

The result: a buyer sees “verified” and assumes the data was measured. Those are not the same thing.

Three grades

We write the second question onto every record. Three grades:

🥇 Gold — signed instrument reading. The number came from a device that cryptographically signs its own readings. The meter has an identity; a message whose value has been altered but whose signature no longer matches is refused, not accepted. This is the firmest grade because nothing is typed between the reading and the record. The human hand is not fully out of it — someone still installs the meter, chooses where it sits and configures it; the honest-limit section below is about exactly that.

🥈 Silver — source document. The number was read from an invoice, a waybill or a weighbridge ticket. The document’s bytes are stored under their own hash, so an auditor can go back to exactly that document and compare what was read against what was written. Not as tight as a device, but it leaves a trail.

🥉 Bronze — declaration. Someone entered the number. It may be correct; it usually is. But its basis is a statement, not a source record.

Seeing bronze is not a bad thing — it is an honest thing. What is bad is presenting every line as equally certain. Saying “I measured it” about something you did not measure is the politest form of getting it wrong.

If a calculation rests on both a signed meter reading and a hand-entered value, its grade is silver or bronze — not gold.

The reason is simple: a number is only as reliable as its weakest input. If you read the gas off a meter but estimated the production volume, the resulting intensity figure is as solid as that estimate. Averaging destroys the very information that matters here.

A proportional statement — “94% of this data comes from signed measurement” — is meaningful, but at inventory level, across hundreds of records. On a single calculation line there is no weighting with which to compute a proportion; the honest thing to state there is the name of the weakest link.

Why we do not produce a single trust score

It is tempting to combine all of this into one number like 87/100. We don’t, for two reasons.

First, it destroys information. Seeing 87 tells you nothing about whether this is “well verified but hand-entered” or “unverified but instrument-measured”. Those are very different situations demanding different responses: the first calls for measurement infrastructure, the second for an audit engagement. A single score makes them look the same.

Second, it is not reproducible. Our entire claim is that you can recompute every number yourself. A score produced by a formula whose weights you cannot see is precisely what we argue against: a black box placed on top of a system built to eliminate black boxes.

If you want a score, one condition makes it legitimate: its formula must be versioned, deterministic, and written into the receipt — so that you can arrive at the same result. Otherwise it is an opinion with a number attached.

What changes in practice

For the buyer: you no longer have to demand “everything verified”. You can require gold on the lines that matter most and accept silver elsewhere. The demand becomes measurable.

For the producer: the return on a meter becomes visible. Today a careful operator and a careless one look identical to the buyer; the grade puts that difference on the record.

For the auditor: where to spend sampling effort becomes obvious. Bronze lines get examined before gold ones.

The honest limit

The grade tells you where the data came from. It does not tell you that the world was correctly represented.

A meter fitted to the wrong pipe, working properly and signing correctly, produces a flawlessly gold-grade wrong number. Provenance narrows that gap; it does not close it.

What narrows it further is comparing a record against its neighbours: whether the quantity you purchased and the energy your supplier declared are consistent with each other. We will take that up in a separate article.

For now the point is this: knowing how reliable a number is, is worth as much as the number itself. And there is no good reason to withhold that.

Make your carbon data impossible to doubt.

Request a demo