The Elephant in the Quantum Laboratory

2026-08-19 · 6,949 words · Singular Grit Substack · View on Substack

Quantum computing has demonstrated "astonishing components".

Quantum computing has demonstrated “astonishing components”. What it has not yet demonstrated, under a strict systems-evidence standard, is independently reproduced, verified, end-to-end fault-tolerant logical computation.

Thesis statement. Quantum-computing experiments have made genuine advances in physical control, encoding, error detection and correction, fault-tolerant components, and small encoded algorithms. But these achievements should not be conflated with the stronger engineering and scientific claim that a useful logical quantum computer has been demonstrated. This essay adopts an explicitly stricter systems-evidence criterion: sustained fault-tolerant computation; error management whose cost remains explicit rather than being hidden by selective survival; an independently specified task; an auditable verification method; and independent reproduction of the substantive result on separately operated hardware. Under that criterion, the published demonstrations reviewed here through 19 August 2026 do not establish all of those requirements in one end-to-end experiment. The central problem is therefore not whether quantum mechanics works. It is whether quantum computing has crossed from remarkable apparatus to reproducible computation under a standard strong enough to support external reliance.

Keywords: quantum computing; logical qubits; quantum error correction; fault tolerance; reproducibility; verification; benchmarking; post-selection; quantum advantage; random circuit sampling; scientific replication


There is an elephant in the quantum laboratory. It is extremely expensive, maintained at improbable temperatures, surrounded by excellent physicists, and referred to by almost every name except the obvious one.

The obvious name is reproducible computation.

That phrase requires care. Quantum experiments are not generally irreproducible. Many physical effects used in quantum information science have been repeatedly observed. Gate operations can be characterised. Entanglement can be produced. Error syndromes can be extracted. Encoded states can be prepared. Logical memories can outperform particular physical baselines. Error suppression can improve with code distance. None of that should be dismissed, and a scientifically serious critique gains nothing by pretending otherwise.

The harder claim is different.

A computer is not merely a physical system that behaves interestingly. It is a system to which an independently specified problem can be submitted and from which an answer can be obtained under a stable abstraction: the answer should not depend on knowing it in advance; the computation should not require the experimenter to discard inconvenient runs until the desired population survives; the error model should remain meaningful outside the calibration episode that produced it; and another competent operator, using a distinct machine that implements the same computational specification, should be able to reproduce the substantive result within declared statistical bounds.

That is the elephant.

The field has made real progress towards logical quantum computation. Google Quantum AI reported below-threshold surface-code memories in which logical error decreased as code distance increased, including a distance-7 memory whose lifetime exceeded that of its best physical constituent qubit (Google Quantum AI and Collaborators, 2025). Bluvstein and colleagues demonstrated programmable logical processing with neutral atoms and later a more elaborate neutral-atom architecture combining repeated error correction, logical operations, teleportation, and mechanisms for entropy removal (Bluvstein et al., 2024; Bluvstein et al., 2026). Butt and colleagues demonstrated a fault-tolerant universal logical gate toolbox on small error-detecting codes and used it to execute a three-logical-qubit Grover search (Butt et al., 2026). Paetznick and colleagues reported logical error improvements against selected physical-circuit baselines using trapped ions, with error detection, correction, and post-selection playing important roles (Paetznick et al., 2026).

Two 2026 preprints push closer to the centre of the argument. Perlin et al. (2026) report end-to-end execution of QAOA and HHL instances using fault-tolerant components and the [[7, 1, 3]] Steane code on Quantinuum processors, including active error-correction cycles and circuits reaching 12 logical qubits. Martiel et al. (2026) report a 70-qubit encoded hard-sampling circuit with a device-dependent fidelity certificate, using 97 physical qubits, syndrome post-selection, and a reported 95% confidence lower bound of 0.284 on the hard state’s fidelity. These are precisely the kinds of results a rigorous critique must confront rather than ignore.

They are important results. They also make the central distinction more precise, not less.

A logical memory is not yet a logical computer. A universal gate set demonstrated on a tiny code is not yet sustained universal fault-tolerant computation. A benchmark whose target statistic is known is not the same thing as solving an independently specified unknown problem. A computation whose result is beyond practical classical calculation raises a different problem: how is correctness established without simply trusting the device? And a result reproduced on repeated runs of the same apparatus is not the same evidential object as reproduction across independently operated machines.

The quantum-computing debate becomes confused because these categories are frequently allowed to slide into one another.

The purpose of this essay is to stop the sliding.

Figure 1. Conceptual evidence ladder. Each higher rung requires the lower rungs but adds a distinct requirement. Current experiments have reached several advanced intermediate rungs, including below-threshold memory and small fault-tolerant logical operations. The final rung is not a claim about a single component; it is an end-to-end reproducibility criterion. Figure is an original synthesis, not an empirical dataset.

1. What counts as a logical qubit?

The phrase “logical qubit” is technically legitimate long before it is commercially or computationally sufficient.

In quantum error correction, a logical qubit is quantum information encoded into a larger physical Hilbert space so that errors affecting physical components can be detected or corrected without directly measuring the encoded logical information. A code is commonly described by parameters [[n, k, d]], where n is the number of physical qubits, k the number of encoded logical qubits, and d the code distance. Roughly, a distance-d code can detect up to d − 1 errors and, under appropriate assumptions, correct up to floor((d − 1)/2) errors.

For surface-code families, a useful approximation to logical error scaling below threshold is:

ε_d ≈ A(p / p_th)^((d + 1)/2)

where:-

ε_d is the logical error rate at code distance d;

-

p is an effective physical error rate;

-

p_th is the threshold error rate; and

-

A absorbs architecture- and noise-dependent factors.

The significance of the threshold is straightforward. If p < p_th, increasing d should suppress logical error rather than amplify the damage caused by adding more imperfect components. Google’s Willow experiments reported precisely this kind of below-threshold scaling, with an error-suppression factor greater than two when code distance increased by two (Google Quantum AI and Collaborators, 2025).

That is a genuine milestone.

But a logical qubit can mean several different operational things:-

an encoded state;

-

an error-detecting logical state;

-

an error-corrected logical memory;

-

a logical qubit with one or more fault-tolerant operations;

-

a set of logical qubits supporting a universal fault-tolerant gate set;

-

a logical processor capable of sustained computation with repeated entropy removal and bounded logical error;

-

a useful logical computer whose result can be independently specified, verified, and reproduced.

Calling all seven simply “logical qubits” is technically possible and rhetorically hazardous.

The first six are physical and architectural milestones. The seventh is the computational claim that matters outside the laboratory.

This essay therefore uses a deliberately stricter term: operational logical qubit. An operational logical qubit is not merely encoded. It must participate in sustained fault-tolerant computation in which logical performance remains meaningfully bounded over the computation, error-management overhead is accounted for rather than hidden through aggressive data rejection, and the resulting computation can enter an external verification and replication protocol.

That definition is not offered as consensus nomenclature. It is a proposed systems-evidence standard. It is deliberately stricter than the coding-theory use of logical qubit, because the evidential burden of claiming a useful computer is larger than the burden of showing that information has been encoded.

Here, useful does not mean commercially profitable, faster than every classical alternative, or ready for mass deployment. It means computationally nontrivial and externally meaningful: the logical machinery must be doing work beyond merely demonstrating the component whose performance is under test, and the result must be capable of supporting a claim outside the calibration exercise itself.

2. The first horn: experiments whose answer is already known

There is nothing wrong with a known-answer test. Every serious engineering discipline uses them.

If a processor is supposed to implement a gate G and we prepare a state whose expected transformation under G is known, the output can be compared against the expectation. If a randomised benchmarking sequence is deliberately constructed so that the ideal composite operation returns to a known reference condition, its decay under repeated gates can estimate an average error quantity. If a small quantum circuit can be simulated classically, the observed output distribution can be compared with the ideal distribution.

These procedures are useful because they are tests.

The conceptual error begins when a test of a component is quietly promoted into evidence of general computation.

Suppose a machine receives instances C_1, C_2, ..., C_m for which an ideal statistic S(C_i) can be computed in advance. The experiment measures an observed statistic Ŝ(C_i) and reports an error:

E = (1/m) Σ_i |Ŝ(C_i) − S(C_i)|.

A small E shows agreement on the tested ensemble, under the experimental conditions used. It does not prove that the machine will correctly evaluate an arbitrary C drawn from a broader class, still less that it will solve an unknown problem whose result cannot be independently evaluated.

Benchmarking is therefore an inference problem. Its validity depends on what is being inferred from what.

Polloreno et al. (2025), in developing a theory of direct randomised benchmarking, make the point with admirable technical precision: standard randomised benchmarking is limited in what it directly measures, and the common practice of rescaling “error per Clifford” to infer “error per native gate” can be an unreliable extrapolation. Their solution is not to abandon benchmarking but to tighten the relation between the measured protocol and the parameter one claims to estimate.

That is exactly the discipline the larger field requires.

A benchmark can establish a benchmark. It cannot, by grammar alone, become an application.

The same distinction appears in quantum-advantage experiments based on random circuit sampling. The machine samples bit strings from a circuit-dependent distribution. For circuits that remain classically tractable, ideal or near-ideal probabilities can be computed and compared with the observed samples. At sufficiently large scales, however, exact classical verification becomes expensive—the very point of the exercise. Bouland et al. (2019) analysed the complexity-theoretic foundations and verification conditions for random circuit sampling. The result is not that verification is impossible; it is that verification itself becomes a sophisticated computational object with assumptions and scaling behaviour that must be stated explicitly.

One can write a simplified cross-entropy-style statistic as:

F_XEB = 2^n × (1/N) Σ_i P_U(x_i) − 1

where:-

n is the number of qubits;

-

N is the number of observed samples;

-

x_i is an observed output string; and

-

P_U(x_i) is the ideal probability assigned to that output by the target circuit U.

For the standard linear-XEB normalisation and the usual Porter–Thomas idealisation, uniform sampling has expected score near zero, while ideal sampling has expected score near one, subject to finite-sample effects and the assumptions of the circuit ensemble. Better agreement with the target distribution generally raises the statistic; the statistic is not itself a universal distance measure between distributions.

But notice the exquisite circularity that must be avoided: to evaluate P_U(x_i), one needs knowledge about the ideal circuit distribution. At small or structured scales, classical computation supplies it. At very large scales, one must rely on partial verification, tractable subinstances, structural checks, theoretical hardness assumptions, interactive certification, or other proxies.

Again, none of these is illegitimate. What would be illegitimate is to describe a proxy as though it were identical to direct verification.

A man who weighs himself on a scale has measured his weight. A man who measures the spring, estimates the floor, models the humidity, and then infers what the scale probably would have said has done something more interesting. He has not done the same thing.

3. The second horn: experiments whose answer cannot be cheaply checked

The second horn is more intellectually serious.

Suppose a quantum computer eventually performs a calculation that is genuinely outside practical classical reach. We should want this. A machine that only solves problems already easy for classical computers would be a very expensive educational toy.

But as the classical cost of solving the problem rises, the classical cost of naive verification may rise with it.

This creates the verification problem:

Given a claimed output y for an instance x, can a verifier determine, with high confidence and substantially less computational work than solving the original problem, that y is a correct output of the intended quantum computation?

For NP-type classical problems, verification may be dramatically easier than solution. Factoring is a familiar example: finding factors can be hard, but multiplying proposed factors is easy. In other settings, particularly sampling problems, there may be no short deterministic certificate for each sample.

The field is fully aware of this. Verification is not an invention of critics. It is an active research programme.

Mahadev (2018) gave a landmark protocol by which a classical polynomial-time verifier can interactively verify a quantum computation under a cryptographic hardness assumption related to Learning With Errors. Other protocols use limited quantum capabilities, trap constructions, measurement-based verification, or special circuit structure. The existence of such work destroys one lazy criticism and strengthens a better one.

The lazy criticism is: “A quantum computation beyond classical reach can never be verified.” That is false as a general statement.

The better criticism is: if a claimed computational result is not directly classically checkable, the experiment must specify which verification protocol closes the evidential gap, what assumptions that protocol uses, what soundness guarantee it supplies, and whether the protocol itself has been implemented at the claimed scale.

Verification does not disappear when quantum advantage appears. It becomes part of the computer.

The distinction can be formalised.

Let C be an intended quantum computation. Let P_C denote its ideal output distribution. Let machine M produce an empirical distribution Q_M,C. A natural measure of distributional error is total variation distance:

D_TV(P_C, Q_M,C) = (1/2) Σ_x |P_C(x) − Q_M,C(x)|.

If D_TV is small, the machine’s distribution is close to the target. But estimating D_TV directly over an exponentially large outcome space is itself generally difficult. For n qubits there may be 2^n possible bit strings. Full distribution reconstruction therefore scales disastrously in the generic case.

One can estimate selected observables instead. One can verify structured properties. One can use interactive protocols. One can employ cryptographic assumptions. One can certify particular classes of output. All are legitimate, but each proves something different.

The dangerous move is to leave the verifier implicit.

If a machine produces an answer that no one can independently calculate, and no scalable verification protocol is attached to the computation, then “the machine returned y” and “y is correct” are different propositions.

A laboratory may possess extraordinary confidence in the first. Science requires an argument for the second.

The strongest 2026 counterexamples narrow the criticism

The two-horn formulation is useful, but it must not be mistaken for a theorem saying that every quantum experiment is condemned either to a known answer or to unverifiability. By 2026, that would be too crude.

Perlin et al. (2026) address one side of the problem by taking error-corrected algorithms further into the computational pipeline. Their preprint reports QAOA and HHL executions using only fault-tolerant components on [[7, 1, 3]] Steane-code logical qubits, with active QEC and measurement-dependent feedback. The largest QAOA circuit reported uses 12 logical qubits mapped to 97 physical qubits and 2,132 physical two-qubit gates. The authors describe the performance as near break-even rather than as a decisive application advantage. That wording matters. It is evidence of increasingly complete fault-tolerant execution; it is not evidence that an independently valuable, classically inaccessible answer has been obtained and independently replicated.

Martiel et al. (2026) attack the other horn more directly. Their doped-Clifford-sampling construction is designed to combine computational hardness with an experimentally accessible fidelity certificate. The reported experiment uses 70 encoded computational qubits and 97 physical qubits. Its circuit contains 468 T gates, and the authors derive a 95%-confidence fidelity lower bound of 0.284 for the hard state. That is much stronger than saying, “the answer is hard, therefore trust the machine.”

It also exposes exactly why careful terminology is necessary. The certificate is explicitly device dependent. The error suppression relies on syndrome post-selection. In the detailed report, the encoded Clifford stage gains roughly a 29-fold state-fidelity improvement at the cost of roughly an 860-fold reduction in effective sampling rate. The authors also calibrate for the specific circuit, select hardware, use Pauli twirling, and discard shots in which identified non-Markovian errors occur. None of those choices makes the experiment invalid. They define what the experiment actually establishes.

The correct conclusion is therefore not that verification is absent. It is that verification has become a first-class engineering object whose assumptions, acceptance rate, hardware dependence, and resource cost must travel with the claim.

This is progress against the elephant.

It is not yet independent replication of an end-to-end fault-tolerant computational result. Martiel et al. provide a sophisticated certificate on one experimental implementation; Perlin et al. provide increasingly complete fault-tolerant algorithm execution. The systems-evidence standard proposed here asks for the conjunction: sustained logical computation, non-hidden resource accounting, an independently specified nontrivial task, an appropriate verification method, and reproduction of the substantive result by an independent execution environment. The 2026 results close important gaps without closing all of them at once.

This distinction is also why the phrase “nothing in quantum computing is reproducible” should be rejected. It is too broad to survive contact with the literature. The defensible criticism is narrower: the field has not yet supplied independent, end-to-end replication of the strongest verified logical-computation claims under a predeclared systems-level criterion.

4. Reproducibility is not the same as correctness

A further confusion must be removed. Reproducibility and validity are not synonyms.

Two machines can agree and both be wrong.

One machine can be right once and impossible to reproduce.

The strongest evidence comes when validity and reproducibility are jointly established.

Figure 2. Conceptual matrix separating reproducibility from validity. Agreement between machines is evidence of reproducibility, not by itself of correctness. Verification against an independently justified target establishes validity. The strongest computational claim requires both. Figure is an original synthesis, not an empirical dataset.

For probabilistic computation, reproducibility cannot mean obtaining the identical bit string on every run. That would misunderstand quantum measurement as badly as demanding that a fair coin land heads in the same sequence in two laboratories.

Instead, reproducibility means agreement in the relevant statistical object.

Suppose two independently operated machines M_1 and M_2 execute the same specified computation C and produce empirical distributions Q_1 and Q_2. A reproducibility criterion might require:

D_TV(Q_1, Q_2) ≤ ε_rep

with confidence at least 1 − δ, for a predeclared ε_rep and δ.

But even this is insufficient. If both are biased by the same flawed compilation rule, both may agree with one another while disagreeing with the intended distribution P_C. We therefore also require validity:

D_TV(P_C, Q_i) ≤ ε_val

or an alternative sound verification condition appropriate to the problem class.

This yields an important triangle:

specification → execution → verification

Reproducibility adds a fourth edge:

independent execution → agreement under the same specification

A mature computational claim needs all four.

The experimental papers reviewed here establish some of these edges, sometimes with considerable sophistication. They do not establish all of them together at useful fault-tolerant scale under the systems-evidence criterion defined above.

5. The post-selection problem

Post-selection deserves particular attention because it is both mathematically legitimate and rhetorically dangerous.

Suppose an experiment produces an outcome y and an acceptance flag A. Instead of reporting the unconditional output distribution Q(y), the analysis reports:

Q(y | A = 1).

If the acceptance probability is:

α = Pr(A = 1),

then, on average, approximately 1/α raw trials are required for each accepted trial.

If α = 0.5, the overhead is mild. If α = 0.1, the experiment needs roughly ten raw attempts per retained outcome. If α = 0.001, the post-selected computation is effectively consuming about a thousand attempts for every reported accepted result, before accounting for other overheads.

The conditional distribution may indeed be much cleaner than the unconditional distribution. That is the point. But computational accounting must include the rejection rate.

More importantly, post-selection can change the claim being made.

Butt et al. (2026), for example, explicitly report logical-state fidelities that improve with post-selection and note cases in which only a fraction of runs are accepted. Their paper is commendably clear about this. Paetznick et al. (2026) likewise describe a scalable method combining error detection and post-selection to reduce logical error rates. Bluvstein et al. (2026) use post-selection in parts of their neutral-atom logical experiments, including confidence-based decoder selection.

The problem is not that these authors conceal post-selection. The problem is what readers, investors, journalists, and sometimes neighbouring technical claims may infer from a post-selected fidelity number.

If a logical processor is to become a computer, rejected runs are not metaphysical accidents. They are part of the cost model.

A useful metric therefore needs at least three numbers:-

conditional logical error among accepted runs;

-

acceptance probability α; and

-

total physical resource cost per accepted logical operation or completed algorithm.

A beautifully small conditional error with vanishing acceptance may be an excellent physics result and a poor computer. Martiel et al. (2026) make the trade-off unusually visible: their reported fidelity improvement is accompanied by a large reduction in effective sampling rate. That transparency is preferable to pretending that rejected shots are free.

The distinction is not unkind. It is arithmetic.

6. Error detection is not error correction

A related ambiguity occurs between detecting errors and correcting them.

A distance-2 code can detect a single error but cannot generally identify and correct an arbitrary single-qubit Pauli error. Such codes can nevertheless be extremely useful in experiments because an error syndrome can be used to reject a run. This can substantially raise the fidelity of the retained data.

But detection plus rejection is not operationally equivalent to correction plus continuation.

The difference becomes decisive in long computations.

If each logical layer has acceptance probability α < 1 and the protocol requires all L layers to survive independently, then a naive acceptance probability behaves like:

α_total ≈ α^L.

Even α = 0.99 becomes α^1000 ≈ 0.000043 after a thousand layers under this simplified independent-layer model. Real architectures can correlate acceptance events, correct rather than reject many faults, restart modules locally, or use constructions in which post-selection is confined to state preparation. The narrower lesson is mathematical: any architecture whose end-to-end acceptance probability really does multiply as α^L with fixed α < 1 pays an exponential survival penalty. A scalable claim must therefore show how its architecture avoids, localises, amortises, or explicitly prices that regime.

Fault tolerance exists precisely because long computations require entropy to be removed while the logical state continues.

This is why recent work on repeated correction, logical teleportation, qubit reset, real-time decoding, and constant-entropy operation matters. It is also why a demonstration of an error-detecting code with spectacular post-selected fidelity should not be described as though it has already solved the problem that active correction was invented to solve.

7. Below threshold is necessary, not sufficient

The threshold theorem is one of the deepest reasons quantum computing remains scientifically plausible.

In idealised form, if physical operations are sufficiently accurate and noise satisfies appropriate assumptions, arbitrarily long quantum computation can be performed with polylogarithmic overhead by increasing the amount of error correction. The threshold transforms quantum computing from a hopeless analogue amplification problem into a potentially scalable digital architecture.

But crossing a component-level threshold is not identical to demonstrating the full assumptions of scalable fault-tolerant computation.

The Google surface-code result is instructive. It reports below-threshold memory behaviour and a logical lifetime exceeding physical constituents. It also reports rare correlated error events and an apparent logical error floor in high-distance repetition-code experiments, with some catastrophic bursts occurring about once per hour and their physical origin not yet understood (Google Quantum AI and Collaborators, 2025).

This matters because ideal threshold scaling depends critically on the structure of noise. Correlated failures can defeat simple independence assumptions. Leakage can create time-correlated effects. Calibration drift can alter effective error channels. Crosstalk can create spatial correlations. Classical decoding can become a bottleneck. A theoretical asymptotic guarantee is only as relevant as the degree to which the physical system satisfies the guarantee’s assumptions.

A useful way to express the issue is to decompose logical failure probability into local and correlated components:

p_L ≈ p_local(d) + p_corr(d, t, architecture).

The first term may fall rapidly with distance. The second may not.

If p_local becomes extremely small while p_corr approaches a floor, increasing distance eventually buys little. The experiment has not failed; it has discovered the next enemy.

That is excellent science.

It is not yet the same thing as a production computer.

8. The calibration boundary

Every advanced experimental platform is calibrated. Classical computers are calibrated too, in the broader engineering sense. The difference is how much of the effective computation depends on a device-specific calibration state that changes over time.

Quantum processors can require calibration of qubit frequencies, pulse amplitudes, pulse phases, crosstalk compensation, readout discriminators, coupler settings, dynamical decoupling sequences, decoder priors, leakage handling, compilation choices, and hardware-specific mappings.

The final executable object is therefore not merely an abstract circuit C. It is more like:

E = Compile(C, H, θ, N, D, t)

where:-

H describes hardware connectivity and native operations;

-

θ represents calibration parameters;

-

N represents the effective noise environment;

-

D represents decoder and mitigation choices; and

-

t represents time, because all of the above may drift.

This is not a defect unique to quantum computing. It is simply more severe there.

The reproducibility question is therefore whether two independent implementations of the same logical specification C can produce compatible results even though their physical realisations E_1 and E_2 differ.

That is what abstraction means in computing.

A C program is not reproducible because two identical motherboards happen to emit the same voltage trace. It is reproducible because different conforming systems implement the same semantic operation.

Quantum computing will have matured when the logical layer acquires comparable independence from the peculiar biography of the hardware beneath it.

9. The software reproducibility problem is already measurable

The hardware problem is difficult to quantify across the whole literature because access to machines, calibration logs, pulse-level controls, and historical device states is uneven.

But software reproducibility has now been studied directly.

Köster et al. (2026) examined a curated sample of 127 quantum-computing papers and conducted a larger automated analysis of nearly 5,000 papers. In their manual sample, only 24.4% provided code artefacts. Of the papers that did provide code, 64.5% failed to execute successfully in a clean environment. Their broader automated analysis found code-availability rates of a similar order and frequent absence of machine-readable environment specifications.

This study does not show that the underlying quantum results were false. It shows something more mundane and, for computing, still serious: in the sampled literature, software-level reproducibility was poor even before one confronts cryogenics, trapped ions, atom loss, decoder drift, or inaccessible calibration states.

One cannot demand that quantum hardware become reproducible while treating its classical software stack as disposable laboratory ephemera.

The executable environment is part of the experiment.

10. The verification escape route—and why it must become operational

A fair critic must acknowledge the strongest answer quantum computing has to the “unknown answer” problem: verification protocols.

Theoretical computer science has not ignored this problem. Mahadev’s protocol showed, under cryptographic assumptions, that a classical verifier can interact with a quantum prover and verify a BQP computation (Mahadev, 2018). Measurement-based verification protocols offer other routes. Special-purpose experiments have also made verification operational rather than merely theoretical. Ringbauer et al. (2025) demonstrated efficiently verifiable measurement-based quantum random sampling on trapped-ion processors at proof-of-principle scale. Liu et al. (2025) used a trapped-ion processor in a certified-randomness protocol in which a classical client challenged and checked a remote quantum server, while explicitly limiting the security claim to a restricted class of realistic near-term adversaries and stated assumptions. These results matter precisely because they show that verification can be designed into an experiment. They do not, however, amount to sustained, application-scale, fault-tolerant logical computation reproduced across independent machines.

This means the future of quantum computing need not require blind faith in an oracle.

But there is a crucial distinction between a theorem that a scalable verification protocol exists, an experimentally demonstrated verification method for a structured task, and a large fault-tolerant quantum computation whose externally meaningful result has been independently reproduced under such a protocol. Ringbauer et al. (2025), Liu et al. (2025), and Martiel et al. (2026) show that the middle category is no longer hypothetical.

The last category is the systems milestone proposed here.

A mature experiment claiming useful beyond-classical computation should therefore publish, before execution:-

the logical problem specification;

-

the success criterion;

-

the verification protocol;

-

the soundness and completeness parameters;

-

the accepted hardware and software deviations;

-

the treatment of rejected runs;

-

the total resource accounting;

-

the stopping rule;

-

the statistical analysis plan; and

-

the independent replication protocol.

The phrase “quantum advantage” should then refer not merely to a fast device output but to the entire verified computational pipeline.

Otherwise the field risks winning the race to produce answers before it has built the institution capable of deciding whether those answers deserve belief.

11. Statistical replication must be predeclared

Quantum outputs are probabilistic, so replication requires statistics.

Suppose an observable has true expectation μ and bounded observations in [0, 1]. If we estimate μ with sample mean μ̂, Hoeffding’s inequality gives:

Pr(|μ̂ − μ| ≥ ε) ≤ 2 exp(−2Nε²).

To achieve error at most ε with failure probability at most δ, it is sufficient that:

N ≥ ln(2/δ) / (2ε²).

The specific bound may be conservative, and many experiments use more efficient estimators, Bayesian models, likelihood methods, or domain-specific tests. The point is methodological: replication tolerances are not prose. They are numbers.

Before an independent laboratory runs the experiment, one should know what constitutes replication.

For example:-

Which observable or distributional property must agree?

-

Within what tolerance?

-

At what confidence level?

-

After how many trials?

-

With what treatment of post-selection?

-

With what allowed recalibration?

-

With what allowed changes to compiler, decoder, or pulse stack?

Without such declarations, “replication” can become a moving target whose bullseye is painted around the arrow after it lands.

Pre-registration is not common in experimental physics in the same form used in some biomedical and social sciences, and it need not be imported mechanically. But beyond-classical computational claims would benefit from precommitted verification criteria because the hypothesis space—circuits, calibration choices, decoders, mitigation procedures, rejection rules, and metrics—is large.

The more flexible the analysis, the more important the audit trail.

12. What would count as the missing demonstration?

The missing experiment is not mysterious.

It would look something like this.

First, a computational task would be specified independently of the machine and frozen before execution. The task would be nontrivial enough that success demonstrates more than a component benchmark.

Second, the computation would run on logical qubits with sustained fault-tolerant operation. Error correction would operate during the computation. Any post-selection would be explicitly included in the resource cost and would not be allowed to convert vanishing success probability into a misleading conditional fidelity.

Third, verification would be part of the protocol. If the answer is classically checkable, it would be checked. If it is not, a sound interactive, cryptographic, structural, or problem-specific verification method would be declared in advance.

Fourth, the entire logical specification, software stack, compiler transformations, decoder policy, calibration procedure, stopping rule, and statistical analysis would be archived sufficiently for audit.

Fifth, an independent group would execute the same logical specification on separately operated hardware. It need not be the same physical architecture. Indeed, cross-architecture replication would be stronger because shared hardware-specific systematic errors would be less likely.

Sixth, the two executions would be compared under a predeclared statistical criterion.

Seventh, the result would be valid as well as reproducible.

That would be a computational event rather than merely an experimental event.

It would also be vastly more persuasive than another press release announcing a larger qubit count.

13. Why the present experiments still matter

The argument above should not be mistaken for technological nihilism.

The current experiments matter enormously because each attacks one of the conditions that a real fault-tolerant computer must satisfy.

Below-threshold error correction matters because without it logical scaling fails.

Decoder latency matters whenever feed-forward decisions depend on decoded syndrome information; a practical architecture must show that its classical control path keeps pace with the operations that require it.

Logical teleportation matters because movement of quantum information without movement of its physical errors is central to modular fault-tolerant architectures.

Entropy removal matters because long computations cannot merely accumulate physical disorder.

Universal logical gate sets matter because Clifford-only computation is not computationally universal.

Post-selection studies matter because they expose how much performance can be recovered by identifying bad runs—and, properly reported, reveal the cost of doing so.

Verification protocols matter because beyond-classical computation creates an epistemic problem that ordinary deterministic computing often hides.

The error is not in doing these experiments.

The error is in allowing the nouns to outrun the verbs.

A logical qubit that stores is not yet a logical qubit that computes usefully. A processor that performs a small encoded algorithm is not yet a fault-tolerant general-purpose computer. A benchmark is not an application. A distributional proxy is not full verification. Repetition on one apparatus is not independent replication. A low conditional error after rejection is not a free logical operation. A theoretical verification protocol is not yet an operational verification stack.

One may admire every rung of a ladder without pretending one has reached the roof.

14. The danger is epistemic, not cinematic

When people hear that an unverifiable quantum computer could be “dangerous”, they often imagine science-fiction scenarios: autonomous machines, broken cryptography, inscrutable superintelligence, and other profitable forms of anxiety.

The immediate danger is duller and therefore more plausible.

It is institutional dependence on computational outputs whose evidential status is weaker than their social authority.

Suppose a future quantum service is used to optimise a drug candidate, value a financial portfolio, select a material, solve an inverse problem, or generate a scientific prediction. Suppose further that the calculation is claimed to lie beyond practical classical reach.

If the result cannot be directly checked, the user needs a verification mechanism.

If the result cannot be independently reproduced, the user needs an audit mechanism.

If the computation depends sensitively on proprietary calibration, compilation, post-selection, and mitigation procedures, the user needs disclosure sufficient to distinguish a computational result from a machine-specific artefact.

Otherwise the service is not merely a computer. It is an authority.

Science has traditionally been suspicious of authorities that cannot be cross-examined.

A black box does not become more epistemically respectable because the box is cold.

15. A proposed standard for claims

A useful way to improve the discussion is to separate five claim levels.

Level 1: Physical-control claim

The system demonstrates specified physical operations with measured fidelities or coherence properties.

Level 2: Encoded-information claim

The system encodes logical information and demonstrates error detection, error correction, or logical memory behaviour.

Level 3: Fault-tolerant component claim

The system demonstrates that specified logical operations have fault-tolerant structure and outperform appropriate physical or non-fault-tolerant baselines under declared conditions.

Level 4: Logical-computation claim

The system executes a nontrivial algorithm using a universal logical operation set with repeated error management and honest accounting of rejected runs and physical overhead.

Level 5: Reproducible verified-computation claim

The system executes an independently specified task whose output is verified under a predeclared protocol and whose substantive result is independently reproduced on separately operated hardware within declared statistical bounds.

Under this taxonomy, the literature reviewed here contains examples of Levels 1–4 in various experimental forms. The disputed systems-evidence boundary is Level 5.

Level 5 is not a standard dictionary definition of computer. It is the standard proposed here for a quantum result on which an external scientist, engineer, regulator, or customer should be entitled to rely without trusting the originating apparatus as an oracle.

16. The elephant, stated plainly

The strongest scientifically defensible version of the criticism is not that quantum computing has produced nothing real. That statement would be easy to refute and would insult excellent experimental work.

The stronger criticism is this:

Under the systems-evidence criterion defined in this essay, the published demonstrations reviewed through 19 August 2026 do not establish a useful, independently reproduced, end-to-end logical computation in which the task is independently specified, error-management costs are fully accounted for, sustained fault-tolerant operation supports the computation, and the correctness of a classically nontrivial result is established by an appropriate verification procedure.

That is a narrower claim.

It is also much harder to evade.

The field currently oscillates between two comfortable experimental regimes. In one, the expected answer or target statistic is sufficiently known that the device can be benchmarked against it. This is necessary engineering, but it does not by itself prove general computational usefulness. In the other, the task is designed to outrun straightforward classical reproduction, at which point verification becomes a central part of the scientific claim rather than an optional appendix.

Between these regimes lies the real engineering problem: build a logical abstraction that can compute what was not pre-known, verify what cannot cheaply be recomputed, and reproduce the substantive result on an independent system.

That is not an unreasonable standard.

It is the standard that makes a computer more than an apparatus.

The irony is that quantum computing does not need lower standards because it is difficult. It needs higher standards because it is difficult.

The more extraordinary the hardware, the less extraordinary the epistemology should be.

The day an independently specified, verified, fault-tolerant logical computation is reproduced across independently operated machines will be a genuine turning point. It will deserve more attention than another qubit-count record because it will establish something deeper than scale.

It will establish trust without requiring faith.

Until then, the elephant remains in the room—not because nothing has been achieved, but because the final achievement is precisely the one a computer is eventually required to make ordinary.


References

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., Ebadi, S., Cain, M., Kalinowski, M., Hangleiter, D., Bonilla Ataides, J. P., Maskara, N., Greiner, M., Vuletić, V., & Lukin, M. D. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3

Bluvstein, D., Geim, A. A., Li, S. H., Evered, S. J., Bonilla Ataides, J. P., Baranes, G., Gu, A., Manovitz, T., Xu, M., Kalinowski, M., Majidy, S., Kokail, C., Maskara, N., Trapp, E. C., Stewart, L. M., Hollerith, S., Zhou, H., Gullans, M. J., Yelin, S. F., . . . Lukin, M. D. (2026). A fault-tolerant neutral-atom architecture for universal quantum computation. Nature, 649, 39–46. https://doi.org/10.1038/s41586-025-09848-5

Bouland, A., Fefferman, B., Nirkhe, C., & Vazirani, U. (2019). On the complexity and verification of quantum random circuit sampling. Nature Physics, 15, 159–163. https://doi.org/10.1038/s41567-018-0318-2

Butt, F., Pogorelov, I., Freund, R., Steiner, A., Meyer, M., Monz, T., & Müller, M. (2026). Demonstration of measurement-free universal logical quantum computation. Nature Communications, 17, 995. https://doi.org/10.1038/s41467-026-68533-x

Google Quantum AI and Collaborators. (2025). Quantum error correction below the surface code threshold. Nature, 638, 920–926. https://doi.org/10.1038/s41586-024-08449-y

Köster, D., Franz, M., Zec, B., Hoess, N., Ramsauer, R., & Mauerer, W. (2026). Works on my QPU: Reproducibility in quantum computing research [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.08348

Liu, M., Shaydulin, R., Niroula, P., DeCross, M., Hung, S.-H., Kon, W. Y., Cervero-Martín, E., Chakraborty, K., Amer, O., Aaronson, S., Acharya, A., Alexeev, Y., Berg, K. J., Chakrabarti, S., Curchod, F. J., Dreiling, J. M., Erickson, N., Foltz, C., Foss-Feig, M., . . . Pistoia, M. (2025). Certified randomness using a trapped-ion quantum processor. Nature, 640, 343–348. https://doi.org/10.1038/s41586-025-08737-1

Mahadev, U. (2018). Classical verification of quantum computations. In 2018 IEEE 59th Annual Symposium on Foundations of Computer Science (FOCS) (pp. 259–267). IEEE. https://doi.org/10.1109/FOCS.2018.00033

Martiel, S., Chung, J.-U., Seif, A., Ghosh, S., Hincks, I., Deshpande, A., Fefferman, B., Gambetta, J. M., & Javadi-Abhari, A. (2026). Sampling hard circuits with verifiably high fidelity [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2607.25941

Paetznick, A., Reichardt, B. W., da Silva, M. P., Ryan-Anderson, C., Aasen, D., Bello-Rivas, J. M., Campora, J. P., III, Chao, R., Chernoguzov, A., van Dam, W., Dreiling, J. M., Foltz, C., Frachon, F., Gaebler, J. P., Gatterman, T. M., Grans-Samuelsson, L., Gresh, D., Hayes, D., Hewitt, N., . . . Svore, K. M. (2026). Improved quantum processor logical error rates via correction and detection. Nature, 654, 349–355. https://doi.org/10.1038/s41586-026-10628-y

Perlin, M. A., He, Z., Armenakas, A. A., Andres-Martinez, P., Hao, T., Herman, D., Jin, Y., Mayer, K., Self, C., Amaro, D., Ryan-Anderson, C., & Shaydulin, R. (2026). Fault-tolerant execution of error-corrected quantum algorithms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.04584

Polloreno, A. M., Carignan-Dugas, A., Hines, J., Blume-Kohout, R., Young, K., & Proctor, T. (2025). A theory of direct randomized benchmarking. Quantum, 9, 1848. https://doi.org/10.22331/q-2025-09-05-1848

Ringbauer, M., Hinsche, M., Feldker, T., Faehrmann, P. K., Bermejo-Vega, J., Edmunds, C. L., Postler, L., Stricker, R., Marciniak, C. D., Meth, M., Pogorelov, I., Blatt, R., Schindler, P., Eisert, J., Monz, T., & Hangleiter, D. (2025). Verifiable measurement-based quantum random sampling with trapped ions. Nature Communications, 16, 106. https://doi.org/10.1038/s41467-024-55342-3


← Back to Substack Archive