# Measure the job before the token.

**Product research note · 7 September 2026 · Pilot hypotheses, not revenue claims or a token offering.**

A verification system can produce many artifacts and still fail to solve a useful problem. The harder question is whether another person or program uses those artifacts to complete work that would otherwise be slower, more expensive, or less dependable. For Proof21, the proposed unit of value is a bounded job with an independent consumer—not a counter that goes up whenever a producer signs another report.

The current website, research articles, and offline demo are learning and evaluation materials. They are not evidence of paying customers, measured fraud reduction, a live checking API, or an available P21 token. A credible pilot should make that starting point explicit and define what would count as progress before collecting flattering numbers.

## Choose one decision that someone already needs to make

A narrow pilot might help an operator compare a payout instruction with supported execution evidence. Another could help a marketplace reproduce a reviewer selection or a creator reproduce a DMT-derived trait. These are distinct jobs with different consumers, evidence, and failure costs. A single impressive total of “verifications” would hide those differences.

For the first job, name the person or system that consumes the result and the action it informs. Document the existing workflow: where evidence comes from, who checks it, which ambiguities require escalation, and what gets repeated during a dispute. Observe that process with permission rather than assume every team has the same problem.

Paul Graham's founder essay on doing things that do not scale offers a practical argument for recruiting and learning from early users directly. It is founder guidance, not a controlled study or a forecast for Proof21. The useful application here is to learn one workflow in detail before assuming a broad market will accept a generic evidence product. [1]

## Compare like with like

Measure the existing process and the proposed assisted process on comparable cases. Record case difficulty, supported evidence type, missing data, operator experience, and review conditions. A faster result obtained by skipping a required check is not the same outcome. A result that another consumer cannot understand or reproduce may simply move the work elsewhere.

Useful measurements include time spent gathering evidence, time spent reviewing it, successful independent reproduction, unresolved cases, and incorrect acceptance or rejection under a defined test policy. Keep the denominator visible. Ten successful checks from ten selected easy cases say something different from ten successful checks among a hundred eligible attempts.

The NIST AI RMF Playbook's Measure function emphasizes selecting appropriate metrics, documenting test sets and methods, and assessing behavior in conditions relevant to deployment. Those principles support disciplined evaluation; citing them does not certify a product or establish compliance. [2] A Proof21 pilot should publish its measurement scope and what was not measured, including any limits on independent review.

## Count failures as work, not as disappearing rows

An unavailable source consumes operator time. An authentic negative report may prevent an inappropriate next action but still require investigation. An unsupported case may need a different tool. Count these outcomes separately rather than drop them from a chart of successful jobs. Otherwise the pilot can reward the very concealment that portable evidence is supposed to discourage.

Separate synthetic fixtures from operational observations. Fixtures are valuable for checking known failure paths because their expected answers are controlled. They do not establish the frequency of those failures in a real population. A demonstration with intentionally inserted defects cannot support a claim about an actual customer's fraud rate or money saved.

Also separate discovery, trial, repeat use, and payment. A page view is not a service call; a service call is not necessarily an accepted result; a subsidized experiment is not necessarily repeat willingness to pay. Record the relationships rather than report the largest available number as adoption. Do not infer a partnership from the use of a public specification.

## Make the cost model inspectable

A proposed service needs an operational cost model that includes evidence retrieval, computation, storage, support, failed attempts, and any optional commitment service. Costs depend on the selected profile and implementation. Do not assume every check requires a Bitcoin transaction, and do not promise free operation because a deterministic calculation itself is small.

A simple pilot ledger can attribute consumed resources to a named job and record which charges are measured, allocated, or estimated. Keep one-time integration work separate from recurring processing, but do not hide it. State whether review labor and unresolved-case handling are included. This is an accounting design for a future pilot, not a published price list or a benchmark.

Payment transport is another layer. x402 describes an HTTP-based flow for payment requirements and paid resource access. It can inform a future charging interface, but it does not decide whether a particular verification is valuable, what price a customer will accept, or whether the underlying evidence is sufficient. Those questions must remain visible in the product experiment. [3]

## Introduce incentives only when the mechanism needs them

A token proposal adds questions about what the token does, who needs it, how incentives affect behavior, and which new risks or dependencies it introduces. Those questions should be answered explicitly rather than treating a token as the default prerequisite for reading a report. Nothing about an offline deterministic check inherently requires a new tradable asset.

Rewarding the number of reports can encourage redundant reports. Rewarding favorable outcomes can discourage honest negative findings. Rewarding participation can attract activity that disappears when the subsidy ends. These are design risks to test, not empirical claims about any named project. A pilot should distinguish work demanded by a consumer from work produced mainly to earn an incentive.

Proof21's product hypothesis is stronger when useful evidence can be evaluated on its own merits. A future economic mechanism would need its own specification, review, and explicit approval. This release does not define token supply, promise rewards, solicit an investment, or claim that P21 has launched. Bitcoin and DMT compatibility are technical design choices, not substitutes for customer evidence.

## A decision at the end of the pilot

Before starting, record the conditions for continuing, changing scope, or stopping. Examples include an independent consumer reproducing the supported result, a documented reduction in duplicated review work, and unresolved cases being routed correctly. Choose actual thresholds with the pilot participant rather than invent universal targets in a launch article.

The final report should preserve unfavorable observations and explain which assumptions survived. A small pilot can justify another narrow experiment without proving an enormous market. The next useful step is a working producer and independent consumer with clear evidence and permissions. More marketing, more artifacts, or a new incentive should not be allowed to stand in for that result.

## Frequently asked questions

### Is generating many receipts evidence of product demand?

Not by itself. Demand requires an identifiable consumer and a useful task. Distinguish generated artifacts from independently used results, repeat use, and actual payment, and report the denominator and any subsidy.

### Does evaluating Proof21 require buying P21?

No. The current materials and offline demo do not require a wallet, payment, or token. P21 is project shorthand here; this release is not a token launch or an investment offer.

### Can an offline fixture demonstrate production savings?

No. It can demonstrate a supported behavior under controlled inputs. Claims about operational savings need a defined baseline, comparable real observations, full cost boundaries, and transparent treatment of failures and uncertainty.

<!-- p21-ecosystem-payments-v09 -->

## Multi-asset accounting and treasury boundaries

Let customers choose among explicitly supported assets; do not force a NAT purchase or a swap before every call. Price service entitlements with an expiring quote, exact atomic units and a stated rounding rule. Record both the received asset quantity and the service entitlement. Stablecoin labels do not eliminate issuer, depeg, bridge or network risks; review those per route and provide a pause policy.

Keep payment settlement, service-credit consumption, the action under inspection and any treasury conversion separate. Settle directly into approved merchant recipients. A later conversion is optional, separately authorized and preferably batched; failed conversion cannot erase settled customer credit or change an evidence finding. Adding a rail is implementation work, not a token partnership or an automatic claim of wallet compatibility.

[Binance assets and methods](https://developers.binance.com/en/docs/products/onchainpay-x402/basics/9.supported-payment-methods) · [Binance integration](https://developers.binance.com/en/docs/products/onchainpay-x402/introduction) · [Coinbase facilitator](https://docs.cdp.coinbase.com/x402/seller/facilitator) · [PayAI assets](https://docs.payai.network/x402/reference) · [Virtuals ACP](https://os.virtuals.io/acp/concepts) · [NAT Ethereum listing](https://www.bitmart.com/en-US/support/articles/7923014477723/360001026214/49446319153179) · [NAT Solana record](https://solscan.io/token/FbKRaqBzupLry3V7QujpNghwrHgxutB4MY11M8aeyVa1) · [NAT BNB Chain contract](https://bscscan.com/token/0x600e3b55d5368c32a94f9372563318adb6a3f882)


## Primary references and next reading

- [1: Paul Graham — Do Things that Don't Scale](https://paulgraham.com/ds.html)
- [2: NIST AI RMF Playbook — Measure](https://airc.nist.gov/airmf-resources/playbook/measure/)
- [3: x402 — HTTP 402 payment flow](https://docs.x402.org/core-concepts/http-402)
- [Proof21 pilot design](https://proof21.xyz/docs/partners/pilot/)
- [Proof21 economics scope](https://proof21.xyz/docs/partners/economics/)
