Does the assessment predict what happens next?
This page compares the402's earlier assessments with later observed outcomes. It is where the system's accuracy, mistakes, and blind spots should become measurable.
There is not enough outcome data to report accuracy yet.
the402 will publish results when a meaningful number of assessments can be compared with later purchases, failures, disputes, and recoveries.
Rather than wait for a number worth showing, the page ships with its measurements named and its shape drawn, so that what the402 intends to be held to is on the record before the first result is in. A metric chosen after the data is in is not a metric.
There is also no accuracy API route, and nothing on this page asks for one, which is why the accuracy record below never says "could not load": there is a difference between a door that is shut and work that has not been done, and only one of them is true here. The one request this page does make is for the counters in who is watching, and that route exists.
What will be reported
- Assessments evaluated
- Correctly predicted outcomes
- False positives
- False negatives
- Unknown at time of purchase
- Time period
- Methodology version
- Verdict against later outcome
- For every verdict standing at a moment in time, what happened next: did purchases through endpoints the402 called verified succeed, and did the ones it called failed go on failing? Reported over a 30-day and a 90-day window, because a verdict that is right for a week and wrong for a quarter is a different failure from one that was never right.
- Calibration by confidence
- the402 states a confidence with every verdict. Calibration asks whether that number means anything: among statements issued at a given confidence, how often did the outcome match? A well-calibrated 0.6 is right about three times in five. An over-confident engine is wrong in a more expensive way than an uncertain one, and only this measurement can tell them apart.
- Corrections
- Every verdict the402 took back, and why: disputes upheld after review, and verdicts an operator pinned by hand. A correction is the clearest evidence a rule was wrong that the402 can have, and counting them is the cheapest honesty available.
- Sample sizes
- The count behind every figure above, published beside it, never as a footnote. A rate computed from four outcomes is not a small result, it is not a result; the reader is owed the denominator in the same glance as the number.
Verdict against later outcome
One row per verdict value, over a fixed window. The question it answers
is the only one that matters to somebody deciding whether to trust a
badge: when the402 said this, how often was it borne out? A
verdict of unknown is included and is not a prediction,
what it reports is how often an endpoint the402 had said nothing about
turned out to work, which is the fairest available measure of what a
silence from the402 is worth.
| Verdict at the time | Endpoints | Outcomes observed | Outcomes matching the verdict | Match rate | Window |
|---|---|---|---|---|---|
| No accuracy record has been published yet. | |||||
Calibration by confidence
Confidence is published with every verdict, so it can be checked like anything else. Grouped into bands, the observed match rate is compared against the confidence the402 stated, and the gap between them is the number, not the match rate on its own. An engine that is right 90% of the time while claiming 99% is worse than one that is right 70% of the time and says so.
| Confidence band | Statements | Outcomes observed | Observed match rate | Gap to stated confidence |
|---|---|---|---|---|
| No accuracy record has been published yet. | ||||
Corrections
Every time the402 took a statement back. Two sources feed it: a dispute upheld after a person reviewed it, and a verdict an operator pinned by hand pending review. Both are recorded append-only inside the trust engine today, and the two surface differently: an upheld dispute will be listed individually in the public dispute register once it opens, while a pin shows on the endpoint itself as a risk flag on its statement rather than in that register. This page is where both are counted, which is a different claim from where either is listed.
| When | Endpoint | What was corrected | How it was raised | Outcome |
|---|---|---|---|---|
| No accuracy record has been published yet. | ||||
Paused The dispute routes are switched off with the rest of the public read API, so no dispute has been filed, and no verdict has been pinned or corrected by hand.
Small samples are published, and labelled
Every figure carries the count it was computed from. Below a stated minimum the page still publishes the figure and marks it as too small to draw a conclusion from, it does not hide the row, and it does not quietly round the cohort up by widening the window until the number looks respectable.
Suppressing a thin cohort is the more flattering choice and the less honest one: the rows that vanish are the ones where the402 knows least, which is precisely where a reader most needs to know that it knows little. A row that says "three outcomes, take nothing from this" tells the truth. A missing row tells the reader nothing and lets them assume the rest is complete.
What would make the402 stop
This page exists to make one commitment checkable from outside. the402 set itself a falsification test at day 90 with two conditions: that independent third parties are actually querying its verdicts each week, and that those queries are followed by purchases, a query and a payment to the same endpoint, measured from the chain, within ten minutes of each other.
If both are near zero at day 90, the402 stops building this. A trust network nobody consults and nobody buys through is not an early-stage trust network; it is an expensive opinion. The day-90 metrics in full are in the402's trust-layer plan (not public yet) and this page is where they will be reported against.
Paused the402's free verdict lookup is live
(GET /v1/reputation/:wallet on api.the402.ai), and it
answers unknown for every wallet today, not because the
engine is in shadow, but because no endpoint has been linked to an
account yet: the verdict a wallet reads is the weakest verdict over the
endpoints its operator has claimed, and claiming is switched off.
The first half of that test, whether anybody is querying at all, is now counted, and the numbers are in who is watching below. The second half is not: a query followed by a payment to the same endpoint needs paid outcomes, which need the checkout, and no purchase has been made through it.
Who is watching
Live
the402 counts three things: requests to its free verdict
routes (GET /v1/reputation/:wallet, the batch
route beside it, the paid on-chain lookup counted separately, and the
in-worker MCP's trust tool) today, and the three explorer
interactions once the explorer is live, beside
pageviews from Cloudflare Web Analytics, which is
cookieless and adaptively sampled. One nightly snapshot, one closed UTC
day at a time.
What is not counted is the longer half of the sentence. No IP
address and no user agent is stored: a caller is a salted hash of the
address and of the first word of the client's user agent (its
name, without the version), so one client sending a hundred version
strings is one caller, and distinct client names behind one
address (curl, an SDK, an agent runtime) are counted apart. Every
browser announces itself as Mozilla, so browsers behind one
address collapse to a single caller, and a client that rotates its own
name mints one per name: the count is bounded by the402's rate limits
and is not verified per caller. The salt is random, it is replaced every 28 days, and each salt is
thrown away sixty days after it is minted, so no hash can be traced back
to an address once its salt is gone and the same visitor is unlinkable
across two windows. No wallet is recorded, the address in a verdict
lookup is the subject of the question, not the asker, and neither is
kept. No page content, no URL, no referrer and no subject id: an
explorer event is one word, so a counter cannot become a record of who
is interested in whom.
Do Not Track and Global Privacy Control are honoured,
in the browser and again at the route, and a request that carries either
is answered and not counted.
| Day (UTC) | Trust queries | Distinct callers | Endpoint views | Probe I/O expanded | Signature checks | Pageviews |
|---|---|---|---|---|---|---|
| 2026-09-21 | 3 | 3 | 3 | 0 | 0 | – |
| 2026-09-20 | 5 | 6 | 0 | 0 | 0 | 10 |
| 2026-09-19 | 3 | 3 | 0 | 0 | 0 | 10 |
| 2026-09-18 | 4 | 3 | 0 | 0 | 0 | 10 |
| 2026-09-17 | 4 | 4 | 0 | 0 | 0 | 10 |
| 2026-09-16 | 7 | 5 | 1 | 0 | 0 | 10 |
| 2026-09-15 | 4 | 4 | 0 | 0 | 0 | 40 |
The seven most recent closed days; today is partial and is left out rather than shown as a slow morning. Snapshot written 2026-09-22T02:00:48.240Z.
Day 90, metric 1
≥ 25 distinct third-party consumers issuing trust queries weekly, ≥ 5 recurring
- Distinct callers, last 7 closed days
- 13
- Recurring callers, last 28 days
- 3
the402 cannot yet separate third-party callers from its own; every caller is counted A recurring caller is one seen on at least three distinct days inside the same 28-day window, three days, because the salt cannot follow anybody past the end of one.
Both numbers are bounded by the402's rate limits rather than verified per caller (a client that rotates its client name mints callers), so they can overstate demand and can never prove it. A metric quoted in an argument about whether to keep going should carry its own error bars.
Pageviews are Cloudflare's, as of 2026-09-22T02:00:48.240Z: 90 across the site in the last seven days, 2 on the explorer, 5 on the methodology, 0 on this page and 0 on the sandbox. They are adaptively sampled, so they are an estimate and the402 does not round them into a total it did not measure.
The rules being measured
What a verdict is, what evidence it rests on, who may run paid checks and how a correction happens are set out in how the402 decides. This page measures that document; it does not restate it.