A working definition — versioned, hash-anchored, free to adopt

What counts as a real human review?

Regulators now require a licensed clinician behind AI-generated clinical decisions. Certifying bodies audit organizations. But the requirement lives at the level of the single determination — and no one has defined what counts there. This is a working definition: five tests a review must pass to be called real, written so that anyone — a vendor, a regulator, a reporter — can check each one.

Version 1.0 · July 19, 2026 · canonical text · SHA-256 d731ce0a…8bf4d38 — the definition is anchored in the same registry it prescribes.

The five tests

TEST 1

Licensed and matched

The reviewer holds an active license appropriate to the decision, verified against the primary source at review time, and matched to the case's jurisdiction and specialty.

Check it: NPI lookup against CMS NPPES →
TEST 2

Independent

The reviewer is organizationally independent of the developer of the AI system under review, and compensation does not depend on the direction of the verdict.

Why it matters: a vendor grading its own homework is governance, not review →
TEST 3

Engaged — measured, not assumed

A defined sample is independently double-reviewed and inter-reviewer agreement statistics are disclosed. A process that cannot show its disagreement rate cannot show its engagement.

The evidence for this test: unmeasured sign-off is the documented failure mode →
TEST 4

Receipted

Every determination produces a tamper-evident receipt at signing — content hash, verdict, credential class, timestamp — verifiable by any third party without trusting the reviewing organization.

Try one: verify any receipt in the public registry →
TEST 5

Durable

Receipts are anchored so they outlive the operator — committed to a public, independently verifiable timestamp and retained for the full audit period. Proof that dies with the vendor is not proof.

The mechanism: registry root countersigned into Bitcoin daily →

What does not count

Where this sits in the landscape

Each of these is real and useful — and each stops one level above the determination. This definition is the missing bottom layer, not a competitor to any of them.

Joint Commission + CHAI

Voluntary Responsible Use of AI certification (2026) for its 22,000+ accredited organizations — governance playbooks: policy, oversight committees, lifecycle, vendor oversight. Organization-level.

URAC

Health Care AI accreditation with AI/ML transparency and bias-testing standards in v8.0+. Accredits the program, not the determination.

NCQA

AI provisions entering Health Plan Accreditation and utilization-management standards for 2026. Plan-level.

CMS WISeR

Requires licensed clinicians behind every non-payment recommendation and human review of every denial — the requirement IS per-determination, but no test of what counts, and no verifiability mechanism. The gap this page fills.

CMS SaMS / "O1" (proposed)

The CY2027 OPPS rule creates the first Medicare payment category for clinical AI — paid per use, to facilities — and flags per-click billing as a program-integrity concern with no verification mechanism proposed. Tests 4 and 5 are the answer; see our O1 page.

The definition, anchored

Canonical textreal-review-v1.txt
SHA-256d731ce0ad779ec950a9167e91cf1e7d1b2afd84d7afb5af73874bd03d8bf4d38
Anchored2026-07-19 · public registry · root countersigned into Bitcoin daily
What this definition claims — and does not. It claims accountability and verifiability, not accuracy. Published evidence on human-AI teaming is mixed; unmeasured human review can underperform either party alone — which is exactly why Test 3 requires measurement rather than presence. It is published by the operators of a physician review network: the conflict is disclosed, and the definition is free for anyone to adopt, meet, or exceed — including our competitors. If a future version changes a word, the hash changes, and both versions stay verifiable forever.