Method · Operating system

How a claim becomes
a defensible verdict.

Six stages. Each one has a specific job and a specific failure mode it prevents. Together they produce the one deliverable that matters: a written verdict the person who commissioned it can attach to a board memo, a vendor file, or an investment decision — and defend.

The six stages

Capture

State the claim in testable terms. Identify what is asserted, what it implies, what hardware or evidence is invoked, and what decision hangs on it. A claim that cannot be stated in testable terms is flagged as structurally unfalsifiable at this stage.

Decompose

Break the evidence structure apart. What did the hardware actually run — QPU, simulator, emulator, or hybrid? What does the pipeline compute? What comparison establishes that the result is meaningful? Where does post-processing begin and where does it end?

Benchmark

Establish the baseline. What does a null result produce — the same pipeline on random data, or a classically-solvable instance? A result that cannot be distinguished from its own null is not a quantum result, regardless of where it ran.

Falsify

Run the test the claim should have faced on day one. Substitute randomized data. Apply the five failure-mode checklist. Compute the permutation null for time-series data. Document what the result would have to look like for the claim to be false — if there is no such result, stop here.

Grade

Score the five scorecard categories. Each produces a finding. Together they produce a verdict: ignore, monitor, test, fund, or reject. Grades are not confidential opinions — they are conclusions attached to the specific evidence that produced them.

Translate

Convert the technical findings into the one-page board memo. Plain English. The verdict, the three most important reasons, and what changes if new evidence arrives. The decision-maker should be able to read this in ninety seconds and act on it.

The five scorecard categories

Five questions. One verdict.

Claim specificity What exactly is being asserted? The claim in testable terms, the evidence cited, and the decision attached to it. Vague claims receive a structural flag before any grading begins.
Evidence provenance Where did the evidence come from? Hardware vs. simulator vs. emulator vs. hybrid. Who ran it, when, on what backend, and what were the job parameters? Hardware claims without job IDs are treated as unverified.
Baseline adequacy What comparison was used? Does the result survive a synthetic null — random data through the same pipeline? Does classical post-processing alone reproduce it? What would a broken or trivial result look like?
Falsification strength What alternative explanations were tested? Timeline artifacts, pipeline artifacts, post-selection, statistical theater, unfalsifiable formulation. The five failure modes, each with its specific tell and the test that catches it.
Decision relevance What action does the surviving evidence justify? Ignore, monitor, test, fund, or reject — with the specific reasoning for each. A verdict without reasoning is an opinion. A verdict with reasoning is a defensible position.
The five failure modes

What goes wrong and how we catch it.

Timeline artifacts

Results built on queue-submission timestamps or file timestamps instead of actual QPU execution times. The tell: per-job execution timestamps from the IBM dashboard contradict the timeline in the paper.

Pipeline artifacts

"Quantum" signal that is actually classical post-processing, extraction bugs, or verification-oracle brute force. The test: randomized data through the same pipeline. If it still works, the hardware wasn't the source.

Post-selection

Endpoints, exclusions, and control definitions chosen after seeing the data. The tell: exclusion criteria that perfectly partition results into "effect" and "no effect" categories, with reasoning constructed after the fact.

Statistical theater

Drift, autocorrelation, and uncorrected multiple comparisons dressed up as significance. The test: circular-rotation permutation with FDR correction. At autocorrelation r=0.65, a t-test fires false positives 40% of the time.

Unfalsifiable claims

Results that no outcome could have killed — which means no outcome can confirm them. The tell: a model that accommodates effects before, during, after, and without the intervention, with no advance specification of which pattern would constitute failure. A claim that cannot be killed cannot be confirmed. This is graded as a structural defect, not a scientific finding.

Limits of the review

What a claim audit is not.

Not included

Formal regulatory certification. Legal or investment advice. Penetration testing or security assurance. Full PQC migration planning or implementation. Unlimited review of an entire vendor portfolio. Review of materials not shared under the engagement scope.

Good fit

One vendor claim, result, or board question with a specific decision attached. PoC result review for teams running NISQ experiments. Pre-publication adversarial review of quantum research before it is used in decision-making. Board-memo generation from existing quantum assessments.