Six stages. Each one has a specific job and a specific failure mode it prevents. Together they produce the one deliverable that matters: a written verdict the person who commissioned it can attach to a board memo, a vendor file, or an investment decision — and defend.
State the claim in testable terms. Identify what is asserted, what it implies, what hardware or evidence is invoked, and what decision hangs on it. A claim that cannot be stated in testable terms is flagged as structurally unfalsifiable at this stage.
Break the evidence structure apart. What did the hardware actually run — QPU, simulator, emulator, or hybrid? What does the pipeline compute? What comparison establishes that the result is meaningful? Where does post-processing begin and where does it end?
Establish the baseline. What does a null result produce — the same pipeline on random data, or a classically-solvable instance? A result that cannot be distinguished from its own null is not a quantum result, regardless of where it ran.
Run the test the claim should have faced on day one. Substitute randomized data. Apply the five failure-mode checklist. Compute the permutation null for time-series data. Document what the result would have to look like for the claim to be false — if there is no such result, stop here.
Score the five scorecard categories. Each produces a finding. Together they produce a verdict: ignore, monitor, test, fund, or reject. Grades are not confidential opinions — they are conclusions attached to the specific evidence that produced them.
Convert the technical findings into the one-page board memo. Plain English. The verdict, the three most important reasons, and what changes if new evidence arrives. The decision-maker should be able to read this in ninety seconds and act on it.
Results built on queue-submission timestamps or file timestamps instead of actual QPU execution times. The tell: per-job execution timestamps from the IBM dashboard contradict the timeline in the paper.
"Quantum" signal that is actually classical post-processing, extraction bugs, or verification-oracle brute force. The test: randomized data through the same pipeline. If it still works, the hardware wasn't the source.
Endpoints, exclusions, and control definitions chosen after seeing the data. The tell: exclusion criteria that perfectly partition results into "effect" and "no effect" categories, with reasoning constructed after the fact.
Drift, autocorrelation, and uncorrected multiple comparisons dressed up as significance. The test: circular-rotation permutation with FDR correction. At autocorrelation r=0.65, a t-test fires false positives 40% of the time.
Results that no outcome could have killed — which means no outcome can confirm them. The tell: a model that accommodates effects before, during, after, and without the intervention, with no advance specification of which pattern would constitute failure. A claim that cannot be killed cannot be confirmed. This is graded as a structural defect, not a scientific finding.
Formal regulatory certification. Legal or investment advice. Penetration testing or security assurance. Full PQC migration planning or implementation. Unlimited review of an entire vendor portfolio. Review of materials not shared under the engagement scope.
One vendor claim, result, or board question with a specific decision attached. PoC result review for teams running NISQ experiments. Pre-publication adversarial review of quantum research before it is used in decision-making. Board-memo generation from existing quantum assessments.