Saturday, August 22, 2026 · Weekly Paper, Week 7, Part 1
Nothing Left to Check Against: When Classical Verification Runs Out
For much of the recent quantum-advantage debate a result has been judged by whether it matches a classical simulation. That test only works where a classical computer can still follow, which is exactly the place these machines are trying to leave. A July 2026 preprint from IBM Research and Algorithmiq says so on its first page, calling the standard circular, and then takes on the harder question: how a quantum estimate earns trust when nothing classical can check it. What follows is the quantity they chose to measure, why its regime has no classical answer, the heuristic they lean on and the evidence they stack behind it, and the reversal at the end, where the method with genuine error bars turns out to run only at the depths a classical reference can still reach. This is a first reading of one paper, written to learn in public. It is my own opinion, and it is not a valuation and not financial advice. Do your own research.
IBM published three advantage claims within days of each other, and this is the first of three notebook entries working through them. It is the right one to start with, because it makes the verification problem itself the central question rather than asking what one particular machine managed to compute.
What pulled me in was the wording. A quantum computation is considered trustworthy when it agrees with a classical one, the authors write, while the point of the quantum computation is to go beyond what classical calculation can reliably provide.1 That is not a criticism from outside the field. It is the field describing its own standard, in a paper whose job is to replace it.
The quantity they chose to measure
The experiment measures something called an operator Loschmidt echo, and the name is more forbidding than the idea. Take a measurable property of a system, call it $O$. Let the system evolve forward in time. Give it a small kick. Let it evolve backward. Then ask how much of the original property survived the round trip. Formally,
$$
S_{\delta} = \frac{1}{2^{n}} \operatorname{Tr}\!\left( U^{\dagger} O U \cdot V_{\delta}^{\dagger} \cdot U^{\dagger} O U \cdot V_{\delta} \right),
$$
The older cousin of this quantity is the state Loschmidt echo, a standard probe of quantum chaos.2 The authors chose the operator version for a practical reason worth understanding. The state echo shrinks exponentially as the system gets bigger, which puts it out of experimental reach at any interesting size. The operator echo, at small kicks, falls only polynomially with how many qubits the kick touches, so it stays measurable on noisy hardware even when the dynamics are complicated.1 At small $\delta$ it also connects to a quantity physicists already care about, the out-of-time-order correlator $C$, through $S_{\delta} \approx 1 - \frac{\delta^{2}}{2} C$, which measures how fast an initially simple observableA measurable property represented mathematically by a Hermitian operator. One measurement returns one of its allowed outcomes; repeated measurements can be used to estimate its expectation value. spreads across the system.3
Why this particular regime has no classical answer
The circuit was not chosen for beauty. It is a Floquet Ising model on the heavy-hex lattice, meaning the qubitsThe basic unit of a quantum computer. Like a 'bit' in a normal computer, but instead of being only 0 or 1 it can be 0, 1, or a blend of both at once. sit on IBM’s own wiring pattern and get kicked by the same repeated block of gates over and over. A few scattering sites stop the forward and backward halves from canceling exactly, and the parameters were tuned so information moves quickly through some regions and slowly through others. The authors call the result semi-scrambling: the observableA measurable property represented mathematically by a Hermitian operator. One measurement returns one of its allowed outcomes; repeated measurements can be used to estimate its expectation value. neither sits still nor spreads ballistically.1
That middle ground is precisely where classical shortcuts break, and they break in different places. Tensor networks with belief propagation approximate the state by capping how much entanglementA quantum link where two qubits' states become tied together, so acting on or measuring one affects the other. It is a key resource for quantum computing. it carries.4 They converge nicely while the scattering is weak, then stop converging as it grows. A Monte Carlo version of Pauli propagation, which follows the observable through the circuit instead of the state, behaves almost oppositely.5 It does better once the dynamics scramble hard and struggles where interference between paths still matters. The two agree at both extremes of the parameter range and disagree through the middle, which is exactly where the experiment lives.
The cost of closing that gap is severe. The tensor-network runs reached a bond dimension of 980, distributed across eight NVIDIA H200 GPUs, at about 160 GPU-hours for each scattering strength at that bond dimension.1 As a resource estimate, the authors ask what bond dimension would be needed to reach a global state fidelity of 0.8. The answer runs from about three thousand at the weakest scattering to around ten quadrillion at the strongest, a figure the paper says lies far beyond any conceivable tensor-network simulation.1
Figure 1: What it would cost the classical side
to keep up. The bond dimension needed to reach a global state fidelity of 0.8
rises by about thirteen orders of magnitude across the parameter range; the largest
simulation actually run reached 980.
An estimate, and then a case for believing it
With no classical answer available, the first thing the authors do is simple. Noise drags the measured echo down, so they measure the same circuit with the kick switched off, where the ideal answer is exactly one, and divide it out:
This is called global rescaling, and it rests on one assumption: that noise attenuates the signal by roughly the same factor no matter how hard you kick the system.1 Everything then turns on whether that assumption holds on the circuit that matters.
Not proof. An accumulation of independent checks, which is how classical methods earned their own standing.
So they attack it. At shallow depths, where a converged tensor network still provides a trustworthy reference, the rescaled experimental signal tracks the classical ground truth, with a small positive bias. Then they change the noise on purpose. They slow the two-qubit gatesA gate that acts on two qubits at once, such as the CNOT. Much harder to perform accurately than a single-qubit gate, and the real test of a machine. from 96 to 128 nanoseconds, repeat the whole experiment on a second processor, and inject two artificial noise models on top. Across those five settings the raw signals vary by a factor of two, while the corrected estimates agree to within 0.03.1
That is persuasive, and it is worth naming what kind of evidence it is. Not proof. An accumulation of independent checks, which is how classical many-body methods earned their own standing. Density functional theory was never certified either; it earned confidence through agreement among independent methods, checks against small cases with exact answers, and convergence with respect to controllable parameters.6 The authors make this comparison themselves, and it is the most disarming argument in the paper.
The price: where the error bars actually live
Persuasive is not the same as bounded. The rescaled estimate carries no quantitative statement of how far it could be from the truth, and the authors say so plainly. For that they turn to probabilistic error cancellation, which learns a model of the device’s noise and then statistically inverts it, producing an unbiased estimate with a real error bar.7
It does not remove the need for trust. It relocates it.
Here is the reversal, and it is the reason this paper stayed with me. Probabilistic error cancellation does not remove the need for trust. It relocates it. The paper’s own words are that the method transforms the problem of validating the observableA measurable property represented mathematically by a Hermitian operator. One measurement returns one of its allowed outcomes; repeated measurements can be used to estimate its expectation value. estimation into the problem of validating the noise model, and that the resulting confidence interval is meaningful only to the extent that the learned model represents the errors actually occurring.1 So the question stops being whether the answer matches a classical calculation and becomes whether the noise model is right.
To their credit, they test the model rather than assume it, checking that twirling really does turn the noise into simple stochastic errors, that its sparsity assumption holds, and that the device behaves without memory once leakage is filtered out.1
But notice where the error-bounded method was actually run. It is applied at circuit depthsThe number of sequential layers of gates in a quantum circuit. Greater depth generally means a longer computation and more exposure to noise. of two and four layers, because that is where a converged tensor network exists to benchmark it against. The regime the whole experiment was built for is six layers, and reaching it with the bounded method is discussed as an extension that follows from modest reductions in hardware error rates.1 The error bars, in other words, sit where you did not need them.
Figure 2: The error bars and the interesting
regime do not overlap. The bounded estimate is demonstrated where a classical
reference still exists, and reaching the target depth is presented as an
extension contingent on lower hardware error rates.
Nothing here changes a revenueThe total money a company brings in from sales, before any costs are subtracted. line. What it changes is what IBM is selling. The interesting product is not the processor, it is the apparatus of trust wrapped around it: the noise characterization, the deliberate noise manipulation, the cross-device repeat, the hierarchy of consistency checks. That is what turns an unverifiable number into a defensible one, and what a customer would be buying when the answer cannot be checked any other way.
It is also expensive, and the credibility of the whole construction currently rests on tests IBM performs on IBM hardware against IBM’s own model of its noise. The bounded method at four layers alone consumed roughly four hours of quantum-processor runtime for a single perturbation setting.1 That is not a criticism of the science, which is careful. It is an observation about how much an outside party can independently confirm, which is very little.
Hold onto one thing, because it runs through all three of these papers. Every cost named above falls as the physical error rate falls, and the authors point at exactly that when they describe reaching the target regime with bounded error.
Where this leaves the question
As a physicist, I find the honesty here more impressive than the result. A paper that opens by calling the field’s own benchmarking standard circular, and then spends most of its length attacking its own estimate rather than defending it, is doing science the way I would want it done.
But notice what has and has not been established. A method now exists for making an unverifiable quantum estimate defensible. Nothing yet says the estimate was worth having. The echo was measured on a model built to be hard, not on a question anyone was independently waiting on, and the authors do not pretend otherwise.
Which raises the obvious next thing to ask. If you can now trust an answer no classical computer can check, does the answer matter? Part two takes up the strongest case IBM has for saying yes: a driven quantum magnet where the larger-system quantum data changed a physical inference the classically accessible sizes could not resolve.
Sources & notes
S. V. Barron, B. Mitchell, V. Tripathi, F. Grieco, I. Rosen, F. Pietracaprina, D. Materia, A. Seif, D. Wanisch, R. L. Panadés-Barrueta, E. van den Berg, J.-U. Chung, A. Eddins, S. Ferracin, G. García-Pérez, J. Goold, L. C. G. Govia, H. Haas, I. Hincks, J. C. Hoke, Z. Holmes, S.-u. Lee, Y. Kim, S. Majumder, S. Maniscalco, S. Montangero, D. Puzzuoli, T. Prosen, J. Raftery, R. Rivera Cardoso, M. Rossmannek, M. Rudolph, B. Saxberg, L. Shirizly, K. Siva, J. Skanes-Norman, I. Siloi, K. C. Smith, B. Sokolov, M. Takita, Y. Teng, M. T. Tan, J. Tindall, Z. Zimborás, M. A. C. Rossi, M. T. Tran, S. N. Filippov and A. Kandala, "Observable Estimation in the Absence of Classical Verification," arXiv:2607.25998 (2026).
T. Gorin, T. Prosen, T. H. Seligman and M. Žnidarič, "Dynamics of Loschmidt echoes and fidelity decay," Physics Reports435, 33 (2006).
B. Swingle, "Unscrambling the physics of out-of-time-order correlators," Nature Physics14, 988 (2018).
J. Tindall and M. Fishman, "Gauging tensor networks with belief propagation," SciPost Physics15, 222 (2023).
M. S. Rudolph, T. Jones, Y. Teng, A. Angrisani and Z. Holmes, "Pauli propagation: A computational framework for simulating quantum systems," arXiv:2505.21606 (2025).
R. O. Jones, "Density functional theory: Its origins, rise to prominence, and future," Reviews of Modern Physics87, 897 (2015).
E. van den Berg, Z. K. Minev, A. Kandala and K. Temme, "Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors," Nature Physics19, 1116 (2023).