The Quantum Bullseye Was Painted After the Shot

2026-09-10 · 5,705 words · Singular Grit Substack · View on Substack

Known answers, post-selection, logical qubits, and the expensive art of discovering what you already knew

Keywords: quantum computing; quantum error correction; logical qubits; post-selection; error detection; error mitigation; selection bias; survivorship bias; target leakage; benchmarking; fault tolerance; scientific method.

There is a wonderfully modern way to perform a miracle. First decide what the miracle is supposed to look like. Then perform the experiment often enough to accumulate a respectable cemetery of failures. Remove the corpses. Photograph the survivors. Finally, announce that nature has yielded.

In less ecclesiastical language, part of this procedure is called post-selection.

Post-selection is not, by itself, misconduct, error, or bad science. It is a legitimate mathematical and experimental operation. Quantum information theory uses it openly. Error-detection experiments use it deliberately. Good papers tell us how much data were discarded, why, what acceptance probability remained, and what conditional fidelity was obtained. There is no scandal in stating a conditional probability as a conditional probability.

The scandal begins when three propositions borrow one another’s evening clothes: we can detect an error; we can reject a run containing an error; and we can make a computation continue correctly despite errors. Those are different achievements. Calling all three “error correction” is rather like calling a smoke alarm, a fire exit, and a fire brigade the same technology because each has encountered smoke.

There is a second problem, and it is more fundamental. Many demonstrations use tasks whose answers are already known. That is often necessary for characterization. Yet if the known target can influence calibration, circuit choice, decoder selection, confidence thresholds, exclusions, stopping rules, mitigation, or interpretation, then the experiment is no longer a clean test of whether the machine can discover an unknown answer. It is a test of whether an apparatus, embedded in a human-and-classical-computer pipeline that knows the destination, can be made to arrive there.

That may be useful engineering. It is not the same epistemic claim.

My argument is simple: known-answer quantum benchmarks are legitimate as benchmarks, but they become weak evidence of autonomous computational capability when target information is allowed to supervise the choices used to generate or select the reported result. Post-selection compounds the problem because it can convert “the machine often fails” into “the accepted runs have splendid fidelity.” A serious demonstration should therefore separate tuning from testing, report unconditional as well as conditional performance, and include preregistered or cryptographically blind challenge instances.

Quantum computers are entitled to exotic physics. They are not entitled to exotic logic.

The answer may be known to the examiner. It must not be allowed to teach the candidate how to pass the examination.

1. A demonstration is not a rehearsal

A known-answer benchmark can characterize a quantum device, but agreement with a known target does not by itself establish general computational capability when knowledge of that target participates in selecting the procedure.

Google’s 2023 surface-code experiment is a useful example of what a disciplined claim looks like. It compared logical performance across code distances and reported that a distance-5 surface-code logical qubit modestly outperformed an ensemble of distance-3 logical qubits over the measured cycles. The point was scaling of error suppression with code size, an essential ingredient of quantum error correction, not discovery of an unknown mathematical fact.1 IBM researchers likewise reported multi-round subsystem error correction with decoder-dependent logical errors on leakage post-selected data, explicitly identifying the conditioning.2 These are characterization experiments, and characterization is real science.

The logical problem appears when characterization is promoted into something larger. Calibration asks whether an apparatus behaves as expected against a reference. Validation asks how a frozen apparatus behaves on instances that did not determine its configuration. Discovery asks it to return information not already available to the process judging and tuning it. Those questions overlap, but they are not identical.

Take the vulgar example 3 × 7 = 21. If everyone knows 21, one can run a noisy device a thousand times, adjust controls, compare decoders, move thresholds, exclude unstable periods, reject ugly syndromes, and celebrate retained outputs agreeing with 21. One may have learned a great deal about the hardware. But one has not yet shown that the same end-to-end system will return the correct answer when nobody operating the pipeline knows what the answer is.

A thermometer is calibrated against known temperatures. Nobody objects. But if its calibration continues until it reports the expected temperature and those same readings are then offered as independent proof of predictive accuracy, calibration has eaten validation.

The question is not “Was the benchmark answer known?” Known answers are indispensable. The question is: Was the answer quarantined from the degrees of freedom whose performance is supposedly being tested?

2. The answer can leak without being typed into a qubit

Target leakage does not require anyone to feed the correct bit string directly into the processor. Information leaks whenever knowledge of the desired answer influences tuning, selection, filtering, decoding, or evaluation.

Other quantitative fields know this problem intimately. Machine learning separates training, validation, and test data because choosing a model on the test set destroys the independence of the test. Econometrics worries about specification search. Clinical trials preregister outcomes because a hundred possible endpoints provide a hundred opportunities to discover a flattering story. The names differ; the arithmetic does not.

Quantum experiments contain many legitimate choices: physical-qubit selection, calibration windows, pulse parameters, decoder families, decoder hyperparameters, syndrome weighting, leakage classification, readout correction, post-selection thresholds, circuit compilations, shot exclusions, and stopping rules. Every one can be defensible. But if researchers repeatedly inspect performance against a known target while choosing among them, the final configuration is selected because of its agreement with that target. Its measured error on the same benchmark is therefore in-sample.

Bluvstein and colleagues’ neutral-atom logical-processor work usefully exposes one such trade-off. Their “sliding-scale error detection” varies a confidence threshold, producing a trade-off between fidelity and acceptance probability. Their GHZ-state results distinguish levels of post-selection and explicitly note post-selection overhead.3 That is how conditional performance should be discussed: as a trade-off, not as an alchemical conversion of rejected trials into non-events.

Suppose one evaluates 50 decoder settings against a known logical state and publishes the setting with the lowest observed error. Even if all 50 are respectable, the minimum is optimistically biased as an estimate of future performance. The winner has been selected partly for fitting peculiarities of the evaluation sample. Calling the winning configuration “the decoder” afterward does not retroactively make it the only configuration tested.

This is target leakage in its civilized form. Nobody need falsify a number. Nobody need act dishonestly. The protocol itself can manufacture optimism.

A clean test requires an informational wall: tune on one set of known instances, freeze the pipeline, and evaluate on new instances whose answers played no role in selection.

3. Correction, detection, and rejection: three nouns, three mechanisms

Error correction, error detection, and error rejection are operationally distinct, and the distinction matters most when claims concern scalable fault tolerance.

Quantum error-correcting codes encode logical information across multiple physical qubits. Syndrome measurements reveal information about faults without directly revealing the encoded logical state. A decoder interprets those syndromes and determines a correction or frame update so computation can continue. Error detection asks the easier question: did something inconsistent with the code occur? Post-selection may then discard the run.

IBM has described early experiments in these terms, explaining that post-selection throws away runs in which the code reports an error and that simply tossing such data does not itself perform decoding.4 Smith, Brown, and Bartlett’s work on “exclusive decoders” deliberately aborts decoding instances deemed too uncertain and analyzes abort probability as part of performance.5 In 2026, Chertkov and colleagues reported an approach explicitly designed to achieve error detection without postselection, emphasizing avoidance of the exponential cost that post-selection can introduce.6

A fault-tolerant computer cannot answer every troublesome syndrome by refusing to finish. That is not resilience; it is good manners at a disastrous dinner party. Large computations contain vast numbers of opportunities for faults. If the response to detected faults is routinely “discard the entire computation,” acceptance can collapse as depth and width increase.

Let A mean acceptance and C a correct returned answer. Then P(C ∣ A) is not an end-to-end success probability. A method can make P(C ∣ A) spectacular by making P(A) microscopic. A minimal useful-output probability per raw attempt is

P(C ∩ A) = P(C ∣ A)P(A).

A rejected run need not be called a logical error; often that would be technically wrong. It is an abort or rejection. But it remains an attempted computation and a resource cost. The scientifically honest report therefore separates conditional logical error, acceptance probability, abort probability, and cost per accepted result. The rejected trials do not belong in the wrong numerator; neither do they belong under the carpet.

That still does not capture time, energy, qubit count, retries, or correlated failures, but at least it prevents the rejected population from disappearing through a grammatical trapdoor.

Any headline about a low logical error rate should provoke an immediate question: low conditional on what, with what acceptance probability, at what retry cost, and does computation continue when the condition fails?

Figure 1. A known answer can guide tuning, selection and interpretation. A stronger computational demonstration withholds the target, requires a committed output, and reveals the answer only afterward.

4. The Texas sharpshooter has acquired a dilution refrigerator

When many configurations are tried against known answers and the best survivor is reported, the result inherits the logic of the Texas sharpshooter fallacy and the statistics of multiple testing.

Selection among many analyses changes the distribution of the selected result. If one runs 100 independent null tests at a nominal five-percent threshold, chance alone produces about five nominally significant results on average. Quantum experiments do not generally express decoder search in p-values, but the selection principle is identical. Try enough qubit subsets, thresholds, calibrations, decoders, circuit mappings, mitigation settings, and exclusion criteria, and some configuration will look unusually fine on the data used to select it.

The Texas sharpshooter fires at a barn and paints the target around the tightest cluster afterward. The quantum version has better funding, colder equipment and a supplementary-information file. It may use maximum-likelihood decoding, syndrome confidence, cross-entropy scores, leakage classifiers and enough Greek letters to frighten a classics department. Yet sophistication does not make a target antecedent.

The methodological cure is boring, which is why it works. Define the metric before the challenge. Freeze the decoder. Freeze the acceptance rule. Freeze the number of attempts. Freeze what counts as a failed run. Receive unseen instances. Commit outputs. Only afterward compare with hidden answers. If multiple pipelines are to be compared, preserve an untouched final test set.

This is not hostility to quantum computing. It is ordinary experimental hygiene. Drug trials, machine-learning competitions and cryptographic evaluations all understand that a test ceases to be a test once it becomes part of the training loop.

Quantum computing will become more credible, not less, when its dramatic claims are subjected to designs in which the bullseye existed before the shot.

5. Why 3 × 7 = 21 is useful—and insufficient

Known-answer circuits are valuable diagnostics. The mistake is treating successful reproduction of a known target as equivalent to solving a genuinely unknown instance.

Quantum benchmarking requires references. Randomized benchmarking estimates gate behaviour through circuits with predictable aggregate properties. State tomography reconstructs states from known measurement settings. Stabilizer experiments prepare known code states because logical failure must be measurable. Cross-entropy benchmarking compares samples with reference probabilities where calculation remains feasible. None of this is inherently circular.

The word “computation” does not require that humanity be ignorant of the answer. Known-answer tests are indispensable. What matters is whether the answer is unavailable to the part of the procedure whose generalization is being evaluated. If the entire experimental ecosystem knows the answer and is allowed to optimize itself against that answer, agreement establishes something weaker: the ecosystem can reproduce a target under supervised conditions.

There is a clean analogy with machine learning. Training accuracy matters. Validation accuracy matters. Neither substitutes for a genuinely untouched test. A neural network with 99.99% training accuracy and miserable test accuracy is not rescued by announcing that the training examples were difficult. It has overfit.

Likewise, a quantum system can be exquisitely engineered to reproduce benchmark behaviour and still fail to establish robustness for unknown tasks. Classical decoders and analysis are not cheating—they are part of fault-tolerant architecture—but their degrees of freedom belong inside the validation boundary.

The appropriate conclusion from known-answer success is “this system passed this benchmark under this fixed protocol,” not “therefore the system possesses the same reliability on unknown computations.”

6. The cryptographic challenge: an answer the laboratory cannot flatter

A hidden-answer challenge provides a powerful way to distinguish supervised reproduction from genuine out-of-sample computational performance.

Cryptography supplies the conceptual machinery. A challenger can generate a secret, construct a public instance from it, and withhold the secret from the party being tested. Verification occurs only after a committed response. Quantum-verification research has long studied how a verifier can test quantum computations, including blind and verifiable protocols where the server lacks information that would trivialize verification.78

The exact challenge must be chosen fairly. One should not demand a problem beyond the claimed scale merely to guarantee failure. Construct a family of instances at the scale the device claims to handle. Keep solutions secret from the laboratory and from any classical system involved in tuning. Require the laboratory to publish or cryptographically commit outputs before solutions are revealed.

Most importantly, specify the accounting. If the machine aborts, that is an abort, not an absent observation. If it requests another instance, that consumes an attempt. If post-selection rejects 99 instances before one answer is returned, the score must say so. If a decoder confidence threshold was fixed before the challenge, good. If it was changed after seeing which answers were correct, the blind test has been punctured.

The challenge need not be RSA factoring. Indeed, demanding a problem beyond the claimed scale would be theatre rather than science. A hidden key, a held-out labelled syndrome set, a hidden circuit parameter, or another cheaply verifiable structured problem can suffice. The essential property is informational: the target must be unavailable to the decision process that determines how the result is produced and accepted.

The virtue of a blind challenge is almost comic in its simplicity: one cannot tune toward an answer one does not know.

7. Conditional fidelity is not a free lunch

Post-selection can legitimately improve the quality of retained results, but that improvement must be priced in acceptance probability and total resources.

Bluvstein et al. explicitly show fidelity-versus-acceptance trade-offs in logical experiments.9 Smith et al. treat post-selection as an engineering resource and analyze regimes in which abort probabilities may scale favourably, rather than pretending aborts do not count.10 This is the correct technical conversation. Post-selection can be useful for state preparation, heralded operations, near-term logical circuits, and protocols where retry is inexpensive. The issue is accounting.

Let an experiment require R raw attempts to obtain N accepted samples. The empirical acceptance rate is  = N/R. If accepted fidelity is , reporting alone discards the operational fact that roughly 1/ attempts are required per accepted sample. At scale, runtime, qubit-time volume and energy may be dominated by rejected work.

Nor are rejections necessarily independent. Bursts, leakage, drift and correlated errors can produce temporal structure. Google’s surface-code work investigated rare correlated events and error floors, illustrating why independent-noise stories can be inadequate.11 A post-selection rule that performs beautifully under one noise regime may behave differently when rare events dominate.

There is also a semantic hazard. “Logical error rate” sounds unconditional to a non-specialist. If the number is calculated only on accepted trials, the adjective belongs in the sentence. “Post-selected logical error rate at acceptance probability a” is less suitable for a press release, but science has survived worse inconveniences.

Conditional performance is valuable information. It becomes misleading only when the condition is hidden, minimized, or rhetorically allowed to evaporate.

Figure 2. Correction continues despite faults; detection identifies faults; rejection/post-selection conditions the result on surviving an acceptance rule. Conditional fidelity must be read together with acceptance probability and resource cost.

8. A logical qubit is not a magic adjective

Encoding information logically is necessary for scalable fault tolerance, but the label “logical” does not itself prove that the encoded object is operationally superior under every relevant metric.

Google’s 2023 work demonstrated the important direction of scaling: the larger surface code modestly improved logical performance.12 The neutral-atom work of Bluvstein et al. demonstrated logical operations at substantial scale and showed that error detection could improve algorithmic performance, while openly reporting the role of post-selection.13 More recent neutral-atom work has explicitly highlighted cases where surface-code analysis used no post-selection, which is exactly the sort of distinction that should be visible when assessing progress.14 Quantinuum has likewise publicized later logical-memory experiments reporting error rates with no post-selection and separately reporting the additional improvement obtained with a small amount of post-selection.15

These distinctions matter because “logical” describes an encoding and processing level, not a certificate issued by nature. A logical qubit can be worse than a physical qubit if gates, measurements, syndrome extraction and decoding add more errors than the code removes. The break-even problem exists precisely because redundancy has costs.

A persuasive fault-tolerance result therefore asks whether logical performance improves as code protection increases under an operationally meaningful, consistently defined metric. Better still, it demonstrates that improvement without relying on a rejection rate that destroys throughput. Better again, it does so over circuits deep enough to expose correlated faults, leakage and decoder limitations. Best of all, the frozen system then performs on unseen tasks.

This is not moving the goalposts. These were the goalposts all along. Fault tolerance was invented because useful quantum computations are too long to survive physical noise by optimism.

“Logical qubit” should begin the methodological questions, not end them.

9. The rejected run is still a physical event

A discarded run may vanish from a plotted fidelity, but it does not vanish from the laboratory, clock, power budget, or probability model.

Post-selected and exclusive-decoder research explicitly treats abort probability and repetition overhead as resources.16 This is essential because a computation that succeeds conditionally may require repeated preparation, execution and measurement before an accepted answer appears. In small demonstrations the overhead can be tolerable. In large algorithms it can dominate.

Suppose a logical subroutine has a 90% acceptance probability. That sounds excellent. Chain 100 independent such stages under a crude “any rejection aborts the whole job” rule and total acceptance is 0.9100, about 2.7 × 10−5. At 99% per stage the same toy calculation yields about 36.6% after 100 stages and about 4.3 × 10−5 after 1,000. Real architectures need not behave this way; local recovery, modularity and scalable decoding can alter the picture profoundly. The toy arithmetic exists to expose the accounting principle.

A rejected shot consumed control pulses. It occupied hardware. It experienced noise. It generated syndrome data. It consumed wall-clock time. If a claimed advantage depends on pretending these resources were free, the advantage is a bookkeeping artefact.

The same applies statistically. If the rejected population differs systematically from the accepted one—as it must when rejection is based on estimated error likelihood—then accepted results describe a selected subpopulation. That may be exactly what the protocol intends. But inference about all attempted computations cannot simply inherit the conditional statistic.

The proper denominator is not a matter of taste. It is part of the scientific claim.

10. Error mitigation is not absolution

Classical post-processing can extract useful estimates from noisy quantum data, but mitigation should not be confused with having physically executed a low-error fault-tolerant computation.

Near-term quantum computing has developed a family of error-mitigation techniques precisely because hardware errors remain substantial. Post-selection is one member of a broader family of strategies that trade samples, assumptions, or classical computation for improved estimates. Smith et al. explicitly discuss post-selection in this resource-aware manner.17

There is nothing ignoble about mitigation. Science routinely corrects instruments, subtracts backgrounds and models noise. The problem arises when a corrected estimator is rhetorically treated as though the physical process itself possessed the corrected fidelity. A photograph sharpened after exposure may reveal real information; it does not mean the lens had the resolution of the processed image.

For computational claims, the distinction is sharper. If classical processing uses a known ideal answer to decide which quantum outputs are plausible, the “quantum result” is no longer an independent output of the quantum device. It is the output of a hybrid inference system supplied with target information. If, by contrast, the classical decoder uses only syndrome information and a noise model fixed independently of the unknown logical answer, that is ordinary fault-tolerant decoding. The information boundary is decisive.

This gives a useful diagnostic question: Could the same classical post-processing be run unchanged if the correct logical answer were hidden? If yes, the method may be a legitimate component of blind validation. If no—if the procedure needs the answer to decide what counts as a good result—then the answer is part of the algorithmic pipeline and the demonstration cannot establish independent discovery of it.

The sin is not classical assistance. The sin is smuggling the marking scheme into the examination and then praising the candidate for knowing the answers.

11. What would actually convince me?

The critique is constructive: quantum-computing claims can be made substantially stronger with preregistration, held-out challenges, complete accounting, and replication.

These safeguards are routine elsewhere and technically compatible with quantum experiments. Verification and blind-computation protocols show that quantum outputs can be tested under information constraints.1819 Contemporary QEC papers already report acceptance probabilities, no-post-selection analyses, decoder comparisons and resource costs in varying degrees.20212223 The field therefore possesses the ingredients.

A serious high-stakes demonstration should publish, before receiving the challenge set, the hardware version, circuit family, compiler, decoder, calibration policy, permitted recalibration schedule, acceptance rule, maximum attempts, scoring rule, treatment of aborts, and primary endpoint. The challenge instances should then be generated independently, with solutions hidden. Outputs should be committed before revelation. Every attempted instance should appear in the denominator.

A second challenge set should test robustness after no further tuning. An independent laboratory should reproduce the protocol where practical. If the task permits classical verification, verification should be mechanical. If it does not, an interactive or cryptographic verification protocol should be considered. Confidence intervals should include uncertainty from finite samples and, where selection occurred during development, the final claim should rest on data untouched by that selection.

For logical qubits specifically, reports should separate: physical error rates; logical error rates without post-selection; logical error rates conditional on each post-selection rule; acceptance probability; retry overhead; decoder training data; decoder test data; leakage treatment; correlated-error behaviour; and scaling with code distance and circuit depth.

This sounds demanding only because publicity has accustomed us to easier examinations.

A machine that survives such a test deserves the adjective “impressive” without assistance from the marketing department.

12. The strongest version of the argument—and its limit

The rigorous criticism is not “known-answer experiments are invalid.” It is that known-answer experiments have limited evidential scope unless target information is isolated from the process being evaluated.

The best QEC work itself demonstrates why the distinction is necessary. Google reports scaling behaviour rather than pretending its 2023 experiment was a general-purpose unknown-problem solver.24 Bluvstein et al. report post-selection overhead and acceptance trade-offs.25 IBM identifies leakage post-selection.26 Smith et al. explicitly model abort probabilities.27 Chertkov et al. treat avoidance of post-selection as an achievement.28 Later demonstrations increasingly emphasize when results are obtained without post-selection.2930

It would therefore be intellectually lazy to accuse every logical-qubit experiment of committing the same fallacy. Some are careful characterization studies. Some use post-selection transparently. Some perform active correction. Some distinguish conditional and unconditional metrics. The critique should land where the evidence permits it to land.

But that restraint cuts both ways. A characterization experiment cannot become a proof of useful fault-tolerant computation merely because the press release is impatient. A post-selected fidelity cannot become an unconditional success rate because the conditional clause is typographically inconvenient. A decoder optimized on known answers cannot be called independently validated on those same answers. And a computation whose target is known throughout tuning cannot establish the same thing as a blind computation whose target is withheld.

The logical hierarchy is straightforward:

known-answer calibration < held-out known-answer validation < blind challenge with frozen protocol < blind challenge with full resource accounting and independent replication.

Each step removes a degree of freedom by which expectation can contaminate evidence.

Quantum computing does not need weaker standards because it is difficult. Difficulty is the reason stronger standards are necessary.

13. A proposed “unknown-answer” standard for quantum claims

The field would benefit from a simple label distinguishing characterization from demonstrations in which the target is genuinely hidden during execution and analysis.

Existing papers already distinguish physical from logical error, corrected from detected errors, and post-selected from non-post-selected results. Extending that discipline to the information status of the target would cost little and clarify much.

I propose that prominent computational demonstrations report four flags alongside the primary result.

First: Target status. Was the ideal answer known to the experimental team during execution and analysis, known only to an independent verifier, or computationally unavailable to all parties?

Second: Protocol status. Were decoder, mitigation, acceptance, exclusion and stopping rules frozen before the evaluated data were collected?

Third: Selection status. What fraction of attempted runs were rejected, and for what reasons? Were rejected runs counted as failures for any end-to-end metric?

Fourth: Replication status. Was the result reproduced on a held-out challenge set or by an independent team?

These flags would not diminish benchmark experiments. They would classify them correctly. A beautifully executed known-answer characterization could remain exactly that. A blind challenge would be recognized as stronger evidence for generalization. The public would no longer be asked to infer experimental design from adjectives such as “breakthrough,” a word whose statistical distribution appears mysteriously independent of hardware error rates.

Classification is not cynicism. It is how science stops unlike claims from being compared as though they were identical.

14. Conclusion: stop grading the exam with the answer sheet open

Quantum error correction is real, difficult, important work. Precisely for that reason, it should not be defended with weaker logic than we would tolerate in an undergraduate methods course.

The literature demonstrates genuine progress: surface-code scaling has crossed important experimental milestones; neutral-atom systems have implemented sophisticated logical circuits; decoders have improved; leakage is being addressed; experiments increasingly distinguish post-selected from non-post-selected performance.3132333435 None of this requires mockery.

What deserves mockery is the intellectual shortcut by which conditional success becomes success, detection becomes correction, reproduction becomes discovery, and a known-answer benchmark becomes proof that a machine can solve an unknown problem.

If I know that 3 × 7 = 21, I can use that fact to calibrate a device. I can use it to debug a decoder. I can use it to test readout. I can use it to estimate a conditional error rate. All perfectly respectable.

What I cannot do, without further evidence, is use my ability to steer an experimental pipeline toward 21 and then announce that the pipeline has independently demonstrated its ability to discover answers.

For that, give it something I do not know—or at least something the people tuning it do not know.

Generate a secret key outside the laboratory. Construct a challenge from it. Freeze the hardware and analysis rules. Hand over the public instance. Let the machine run. Count every attempt. Count every rejection. Require an answer before revealing the secret. Then compare.

If the quantum computer succeeds, splendid. Open the champagne, publish the Nature paper, hire a poet for the press release. I will happily read all three.

If it succeeds only after the answer is known, after thresholds are moved, after troublesome runs are discarded, after decoders are selected against the target, and after the denominator has been taken behind the laboratory and quietly shot, then what has been demonstrated is not the thing being advertised.

It is a rehearsal performed with the script open.

The truth is rarely pure and never simple; quantum computing sometimes appears to have adopted this as an engineering specification. The machinery is difficult, the mathematics is difficult, and the noise is vicious. But the central experimental principle is almost offensively easy:

Do not let the answer teach the test how to pass itself.

And if the machine must throw away its failures before it looks clever, tell us how many bodies are in the bin.

That is not anti-quantum. It is the minimum price of calling the exercise science.

A machine that succeeds on a frozen, blind challenge has earned our attention. A machine that succeeds after the answer has been allowed to supervise thresholds, exclusions or selection has earned something else: another development run.

The bullseye must be painted before the shot. Otherwise one has not discovered accuracy. One has decorated noise.



References


-

Google Quantum AI. (2023). Suppressing quantum errors by scaling a surface code logical qubit. Nature, 614, 676–681. https://doi.org/10.1038/s41586-022-05434-1↩︎

-

Sundaresan, N., Yoder, T. J., Kim, Y., Li, M., Chen, E. H., Harper, G., Thorbeck, T., Cross, A. W., Córcoles, A. D., & Takita, M. (2023). Demonstrating multi-round subsystem quantum error correction using matching and maximum likelihood decoders. Nature Communications, 14, 2852. https://doi.org/10.1038/s41467-023-38247-5↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

IBM Quantum. (2022). How IBM Quantum is advancing quantum error correction. IBM Quantum.↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Chertkov, E., Potter, A. C., Hayes, D., & Foss-Feig, M. (2026). Error detection without postselection in adaptive quantum circuits. Physical Review Research, 8, 023057. https://doi.org/10.1103/6b14-7wkt↩︎

-

Morimae, T. (2012). Verification for measurement-only blind quantum computing. arXiv:1208.1495.↩︎

-

Barz, S., Fitzsimons, J. F., Kashefi, E., & Walther, P. (2013). Experimental verification of quantum computations. arXiv:1309.0005.↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Google Quantum AI. (2023). Suppressing quantum errors by scaling a surface code logical qubit. Nature, 614, 676–681. https://doi.org/10.1038/s41586-022-05434-1↩︎

-

Google Quantum AI. (2023). Suppressing quantum errors by scaling a surface code logical qubit. Nature, 614, 676–681. https://doi.org/10.1038/s41586-022-05434-1↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

Bluvstein, D., Geim, A. A., Li, S. H., Evered, S. J., Bonilla Ataides, J. P., et al. (2026). A fault-tolerant neutral-atom architecture for universal quantum computation. Nature, 649, 39–46. https://doi.org/10.1038/s41586-025-09848-5↩︎

-

Quantinuum. (2026). Real Time Error Correction at Increased Scale. Corporate technical report/blog; cited as a company-reported result, not peer-reviewed evidence.↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Morimae, T. (2012). Verification for measurement-only blind quantum computing. arXiv:1208.1495.↩︎

-

Barz, S., Fitzsimons, J. F., Kashefi, E., & Walther, P. (2013). Experimental verification of quantum computations. arXiv:1309.0005.↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Bluvstein, D., Geim, A. A., Li, S. H., Evered, S. J., Bonilla Ataides, J. P., et al. (2026). A fault-tolerant neutral-atom architecture for universal quantum computation. Nature, 649, 39–46. https://doi.org/10.1038/s41586-025-09848-5↩︎

-

Quantinuum. (2026). Real Time Error Correction at Increased Scale. Corporate technical report/blog; cited as a company-reported result, not peer-reviewed evidence.↩︎

-

Google Quantum AI. (2023). Suppressing quantum errors by scaling a surface code logical qubit. Nature, 614, 676–681. https://doi.org/10.1038/s41586-022-05434-1↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

Sundaresan, N., Yoder, T. J., Kim, Y., Li, M., Chen, E. H., Harper, G., Thorbeck, T., Cross, A. W., Córcoles, A. D., & Takita, M. (2023). Demonstrating multi-round subsystem quantum error correction using matching and maximum likelihood decoders. Nature Communications, 14, 2852. https://doi.org/10.1038/s41467-023-38247-5↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Chertkov, E., Potter, A. C., Hayes, D., & Foss-Feig, M. (2026). Error detection without postselection in adaptive quantum circuits. Physical Review Research, 8, 023057. https://doi.org/10.1103/6b14-7wkt↩︎

-

Bluvstein, D., Geim, A. A., Li, S. H., Evered, S. J., Bonilla Ataides, J. P., et al. (2026). A fault-tolerant neutral-atom architecture for universal quantum computation. Nature, 649, 39–46. https://doi.org/10.1038/s41586-025-09848-5↩︎

-

Quantinuum. (2026). Real Time Error Correction at Increased Scale. Corporate technical report/blog; cited as a company-reported result, not peer-reviewed evidence.↩︎

-

Google Quantum AI. (2023). Suppressing quantum errors by scaling a surface code logical qubit. Nature, 614, 676–681. https://doi.org/10.1038/s41586-022-05434-1↩︎

-

Bluvstein, D., Evered, S. J., Geim, A. A., Li, S. H., Zhou, H., Manovitz, T., et al. (2024). Logical quantum processor based on reconfigurable atom arrays. Nature, 626, 58–65. https://doi.org/10.1038/s41586-023-06927-3↩︎

-

Smith, S. C., Brown, B. J., & Bartlett, S. D. (2024). Mitigating errors in logical qubits. Communications Physics, 7, 386. https://doi.org/10.1038/s42005-024-01883-4↩︎

-

Bluvstein, D., Geim, A. A., Li, S. H., Evered, S. J., Bonilla Ataides, J. P., et al. (2026). A fault-tolerant neutral-atom architecture for universal quantum computation. Nature, 649, 39–46. https://doi.org/10.1038/s41586-025-09848-5↩︎

-

Quantinuum. (2026). Real Time Error Correction at Increased Scale. Corporate technical report/blog; cited as a company-reported result, not peer-reviewed evidence.↩︎


← Back to Substack Archive