Saturday, August 22, 2026 · Weekly Paper, Week 7, Part 3
Cheap to Believe, Expensive to Run: What It All Adds Up To
The first two papers in this series end in the same place: to believe the answer you must first trust assumptions about how the hardware’s noise distorted it. The third takes a different route and builds the certificate into the circuit itself. IBM Research, with two University of Chicago coauthors, ran a 70-qubit circuit of depth 70 inside an error-detecting code, doped it with 468 T gates at locations that preserve the code’s checks, and derived a fidelity of at least 0.284 with 95 percent confidence, on far weaker noise assumptions than the usual proxies require. It cost about 99.94 percent of the shots. What follows is how that certificate is built, why 468 small gates turn an easy circuit hard, what the bound actually proves, the two different things classical hardness means here, and what all three papers add up to when read together. This is a first reading of one paper, written to learn in public. It is my own opinion, and it is not a valuation and not financial advice. Do your own research.
There is no molecule here, no physical question waiting on the result. The circuit exists to be hard to simulate. What makes it worth a notebook entry is that it pursues three things sampling demonstrations have historically had to trade against each other: classical hardness, error suppression, and a fidelity certificate you can check.1
The certificate is in the circuit
Start with the problem the certificate solves. You want to know how well the machine prepared the state it was meant to prepare, and the usual answer is a fidelity proxy such as cross-entropy benchmarking, which infers fidelity from statistical properties of the output samples. That works only under assumptions: noise sufficiently weak, uniform across the chip, and independent in time.1 Real hardware does not oblige.
The certificate then arrives in two steps. First, strip out the gates that make the circuit classically hard. What remains is a Clifford circuit, and its fidelity can be measured directly using a technique that samples random stabilizers of the target state and makes no assumptions at all about the hardware noise.2 That measurement gives 0.32. Second, put the hard gates back in locations chosen so the code’s checks see the same thing before and after, and bound how far the fidelity could have fallen:
where $F_{1}$ is the measured Clifford fidelity, $F_{2}$ the fidelity of the hard circuit, and the subtracted term is the probability that a fault which previously stabilized the state now damages it.1
What makes it hard is a different resource, which the field calls magic, supplied here by 468 single-qubit T gates. Add enough of them and the stabilizer shortcut becomes exponentially expensive.
Entanglement alone is not enough. Magic is what kills the shortcut.
Which creates the tension the paper has to resolve. The circuit can be certified only because the code’s checks work, so scattering new gates through it carelessly would break the very thing that was going to prove the answer. So the T gates are not scattered. They sit only in locations where they commute with the checks, which is to say where a Z fault at that spacetime location would have evaded those checks anyway. The circuit gets harder to simulate while the certificate stays intact.
There is a neat hardware detail underneath. On IBM machines these rotations are handled by shifting a reference frame in software rather than firing a physical pulse, so they add essentially no noise of their own. The experiment supports that directly: the acceptance rate is statistically indistinguishable before and after doping.1
Figure 1: The certificate in two steps. The
undoped circuit's fidelity is measured directly, then the hard gates are added
at locations that preserve the code's checks, so the fidelity of the classically
hard circuit can be bounded from below.
At 468 T gates, measuring the fidelity directly is no longer tractable, since the classical machinery that direct fidelity estimation needs becomes part of the problem the experiment exists to escape. So they measure the easier Clifford cousin, getting 0.32. Then a Monte Carlo sweep across noise strengths and polarizations puts the maximum estimated loss from doping, within their modeled Pauli noise, at 0.013. Take the pessimistic end of the measurement, subtract the pessimistic end of the loss, and 0.284 at 95 percent confidence is what remains.1
They then test that bridge rather than asserting it. With five hard gates the direct measurement still works and it agrees. With seventy-five, classical simulation is still possible and an independent proxy agrees, on stronger assumptions. Doping with Clifford S gates instead provides a control. The syndrome distributions barely move across all of it.1
What remains assumed is worth naming. That noise behaves as Pauli noise after twirling. That the doping gates add no noise, which they check but do not prove. And that the noise environment is the same across circuits, which they address by interleaving every shotA single execution of a quantum circuit, ending in one readout. Estimating anything useful takes many of them. across an eight-hour run.1 So the certificate sits between a measurement and a model: weaker than measuring the hard state directly, which is not tractable, and stronger than a proxy that needs well-behaved noise.
One detail I find quietly persuasive. With readoutThe hardware process that converts a qubit measurement into classical data, usually a reported 0 or 1. mitigation applied they estimate the true state fidelity at 0.57, roughly double the headline. They built the bound on the raw 0.32 instead.1
Two kinds of hard
"Classically hard" means two different things here, and they should not be conflated. One is asymptotic: the circuit family is universal in the worst case, and assuming two complexity-theoretic conjectures, no classical sampler reproduces its output distribution on average.1 Those conjectures are unproven, and similar ones underpin other sampling-based advantage proposals, going back to the argument that classically simulating certain commuting circuits would collapseThe state update associated with a measurement outcome. In an ideal projective measurement, the state is projected into the eigenspace corresponding to that outcome. the polynomial hierarchy.45
The other is empirical. A search over 100 million random bipartitions of the 70-qubit graph state found no cut below rank 30, which the authors take as evidence that a useful low-width tensor-network decomposition will be hard to find, and stabilizer methods are conjectured to need a term count with 33 digits.1 One obstruction is more interesting than the rest. Noise usually helps the classical side, since a sufficiently noisy circuit can be approximated by discarding correlations, and running inside the code cut the effective two-qubit error rate tenfold. The error detectionHow often a qubit can flag that an error has just occurred. Catching errors makes them far easier to correct. is not only buying fidelity, it is buying hardness.
The error detection is not only buying fidelity, it is buying hardness.
Then they concede the point. Better algorithms may come, and their stated goal was to show that hardness guarantees can coexist with trusted execution, not to claim a permanent barrier.1
The price
Keeping only the runs where nothing was detected discards about 99.94 percent of them. The hard 468-T-gate sampling run produced 2,051 accepted samples in 16.1 minutes.1 Stated as an exchange rate, the encoding buys 29 times better state fidelity at the cost of an 860-fold drop in effective sampling rate.
The authors name the consequence themselves. Error detectionHow often a qubit can flag that an error has just occurred. Catching errors makes them far easier to correct. alone cannot scale indefinitely, because the post-selection overhead grows with the machine. And verification that assumes nothing whatever about the device, rather than merely assuming less, remains unsolved.1
My takeaway. This is the weakest of the three papers on producing something worth knowing, and the strongest on making a result checkable. It does not need the noise to be weak or uniform. It needs a weaker and more structural set of assumptions: Pauli noise after twirling, effectively noiseless doping gates, and a comparable noise environment across circuits. The certificate remains device dependent, and within those assumptions it hands you a number with a confidence interval. That is a real change in kind, and it costs almost every shotA single execution of a quantum circuit, ending in one readout. Estimating anything useful takes many of them. you take.
What the three add up to
Read together, the three answer one question in three different currencies. With Algorithmiq, trust comes from accumulated evidence, testing the assumptions behind the answer rather than the answer itself.6 With Qedma, from a second opinion, repeating selected points on a competitor’s trapped-ion hardware.7 Here, from the circuit itself.
Rigor, reach and affordability are available in pairs. Not yet as a set.
Then a pattern shows up that no single paper can state. In the first two, the deepest headline regime is ultimately carried by a lower-overhead heuristic rather than by the most controlled estimator. This paper inverts that. Its bound does reach the hard regime, and it costs nearly all of the data. So rigor, reach and affordability are available in pairs. Not yet as a set.
Figure 2: My synthesis of the tradeoff across the
three results. Each reaches two of the three properties an advantage demonstration
would ideally combine. The missing edge is the finding.
The investor’s read
After three papers my answer is that the advantage is real and that it is narrow. The narrowness is the finding, not a complaint. Each result ran on a problem chosen to sit where classical shortcuts give out, on hardware the problem was shaped to fit. That does not make it fake, since every experiment picks a system it can run and all three document the choosing. It does mean the demonstrations tell you where the frontier is, not that it has moved closer to something anyone needs done.
I began this series saying IBM had not convinced me its hardware was moving the application frontier. One of the three moved me. The driven-magnet result changed a physical inference the classical calculations in that study could not resolve, the reverse of the chemistry work where a processor supplied samples and a classical solver drew the conclusion.
What binds all of it is noise. None of the three headline results comes from raw hardware alone. Two recover observablesA measurable property represented mathematically by a Hermitian operator. One measurement returns one of its allowed outcomes; repeated measurements can be used to estimate its expectation value. with error mitigation, and this one detects errors and discards the runs that fail. Different fixes, same bottleneck: physical error rates. Lower them and all three improve at once, each in its own currency: longer dynamics, deeper error-bounded estimates, more surviving runs. That is why error suppression is where I am looking next.
One thing I did not expect. Two of the three depended critically on private-company partners. Qedma supplied the mitigation stack behind the driven-magnet result, and Algorithmiq led the modeling, circuit design and classical benchmarking in the other, with IBM supplying the hardware and, in Algorithmiq’s case, the error mitigation too. The public market offers several ways to own the hardware layer. Some of the capabilities deciding what it can demonstrate are being built inside companies that are not publicly traded.
Where I land
None of it adds up to a moatA durable advantage that protects a company from competitors, like the moat around a castle.. A hard family of circuits is not a permanent barrier around one 70-qubit instance, and the authors never claim otherwise. What it adds up to is a field that has stopped asking whether a quantum computer can beat a classical one and started asking a harder question: how would you know?
That is progress of a less exciting kind than the headlines suggest, and a more durable kind. My reading goes two ways from here. Narrower applications where an advantage might matter inside this decade, and papers from the privately held companies sitting inside results like these. Not a position. A reading list.
Sources & notes
S. Martiel, J.-U. Chung, A. Seif, S. Ghosh, I. Hincks, A. Deshpande, B. Fefferman, J. M. Gambetta and A. Javadi-Abhari, "Sampling hard circuits with verifiably high fidelity," arXiv:2607.25941 (2026).
S. T. Flammia and Y.-K. Liu, "Direct Fidelity Estimation from Few Pauli Measurements," Physical Review Letters106, 230501 (2011).
S. Aaronson and D. Gottesman, "Improved simulation of stabilizer circuits," Physical Review A70, 052328 (2004).
M. J. Bremner, R. Jozsa and D. J. Shepherd, "Classical simulation of commuting quantum computations implies collapse of the polynomial hierarchy," Proceedings of the Royal Society A467, 459 (2010).
S. Aaronson and L. Chen, "Complexity-theoretic foundations of quantum supremacy experiments," 32nd Computational Complexity Conference (CCC 2017), 22:1 (2017), arXiv:1612.05903.
S. V. Barron, B. Mitchell, V. Tripathi et al., "Observable Estimation in the Absence of Classical Verification," arXiv:2607.25998 (2026).
E. Leviatan, T. Watad, R. Perry et al., "Resolving Structure in Prethermal Floquet Dynamics with Precision Quantum Computation," arXiv:2607.24937 (2026).