← Back

The Answer the Classical Methods Could Not Settle

Part one ended on a complaint. A method for trusting an unverifiable quantum estimate is only worth having if the estimate itself is worth having. This paper is the answer to that, and the strongest of IBM’s three recent advantage claims. Qedma, IBM Quantum, RIKEN and BlueQubit drove a magnet on 51 and 74 qubits and produced a conclusion the classical calculations in the study could not settle: whether a strange subharmonic rhythm, one swing for every four kicks of the drive, persists as the system grows or fades away as an artifact of small size. What follows is the physics of a driven magnet that refuses to scramble, the finite-size argument the whole result turns on, what the two error-mitigated points added, what the correction costs, and the fact that the problem was selected to sit where classical methods struggle. This is a first reading of one paper, written to learn in public. It is my own opinion, and it is not a valuation and not financial advice. Do your own research.

This is the second of three notebook entries on IBM’s recent advantage claims, and the only one of them where the machine produced a physical result rather than a method for trusting one.1

A magnet that keeps time

Picture a grid of qubits wired so that each behaves like a tiny magnetic needle, tugging on its neighbors. Now hit the whole grid with the same block of gates over and over, at a steady rhythm. In physics this is called a Floquet drive, and the expected outcome is dull. A generic interacting system tends to absorb energy from a periodic drive and has nowhere to dump it, so the grid heats until every local measurementA physical process that produces a classical outcome and updates the quantum state. In an ideal projective measurement, the state is left in an eigenstate associated with the observed outcome. sits at featureless infinite-temperature values. All structure gone.

Some systems delay that fate for a long stretch. This is the prethermal regime, and it has a rigorous foundation: for a fast enough drive, heating can be exponentially slow, leaving a long window in which the system behaves as though governed by an approximate conserved energy.2 The experiment lives inside that window.

What the authors find there is stranger than a system merely holding still. The magnetization, which is simply the average of each qubit’s up-or-down measurement,

$$ M = \frac{1}{N_{q}} \sum_{q=1}^{N_{q}} \langle Z_{q} \rangle , $$

does not settle. It swings, completing roughly one cycle for every four kicks of the drive. The system is being driven at one rhythm and answering at a quarter of it, a subharmonic response.1 Stranger still, it does this while getting more complicated rather than less. The rhythm is not a frozen leftover surviving in a quiet corner. State entanglementA quantum link where two qubits' states become tied together, so acting on or measuring one affects the other. It is a key resource for quantum computing. and operator complexity keep growing around it, rapidly at first and then more slowly.

The mechanism behind these particular oscillations is not yet understood. The paper reports the effect and then states plainly that a theoretical understanding of the mechanism is still missing, singling out a strong dependence on how many neighbors each qubit has as the part most needing explanation.1 An experiment that produces a physical effect whose mechanism is not understood is doing physics, not running a benchmark.

prethermal value one swing, four drive cycles 0 10 20 30 drive cycle magnetization prethermal value four drive cycles 0102030 drive cycle (magnetization, schematic)
Figure 1: Schematic. A driven magnet that should scramble into featurelessness instead keeps a rhythm, completing about one oscillation for every four kicks of the drive. The period is drawn to scale; the amplitudes are illustrative.

The question you cannot answer with a small system

Here is where the experiment stops being a demonstration and starts being an argument. Small groups of anything are twitchy. Flip five coins and you can get four heads. So the real question is not whether the swing exists at 30 qubitsThe basic unit of a quantum computer. Like a 'bit' in a normal computer, but instead of being only 0 or 1 it can be 0, 1, or a blend of both at once., it is whether the swing survives as the system grows, or fades to nothing as a finite-size artifact.

To settle that you watch how the oscillation’s size shrinks with system size and ask what it is shrinking toward. The authors fit

$$ A(N_{q}) = a\, e^{-\sqrt{N_{q}}/\xi} + c , $$
Real at 30 qubits, absent in the limit. That is the difference the fit has to settle.

where $A$ is the measured amplitude, $N_{q}$ the number of qubits, $\xi$ a length scale over which the finite-size effect dies away, and $c$ the leftover: the amplitude that remains once the system is infinitely large.1 Everything hangs on $c$. A nonzero $c$ supports an oscillation that persists as the system grows. If $c$ is zero, the oscillation is a finite-size effect that fades at scale: real at 30 qubits, absent in the limit.

Exact classical simulation reaches 21, 28 and 35 qubits here, because every qubit you add doubles the state space. Fit those three sizes alone and $c$ comes out at $7.3 \times 10^{-3}$, with a 68 percent interval running from $-3.7 \times 10^{-3}$ to $1.2 \times 10^{-2}$.1 That interval contains zero, and it even dips negative, which is impossible for an amplitude. In plain terms, from the classically checkable sizes alone the oscillation might vanish entirely in the large-system limit.

What the machine added

Add the error-mitigated measurementsA physical process that produces a classical outcome and updates the quantum state. In an ideal projective measurement, the state is left in an eigenstate associated with the observed outcome. at 51 and 74 qubitsThe basic unit of a quantum computer. Like a 'bit' in a normal computer, but instead of being only 0 or 1 it can be 0, 1, or a blend of both at once. and $c$ becomes $9.6 \times 10^{-3}$, with an interval from $8.1 \times 10^{-3}$ to $1.1 \times 10^{-2}$.1 Zero now sits outside that interval, which is strong evidence that the oscillation persists in the large-system limit of the heavy-hex ladder geometries studied.

That is the whole result, and it is worth being precise about why it matters. The quantum processor did not reproduce something already known more quickly. Its two extra data points are what make the physical conclusion possible, and without them the honest answer would remain "cannot tell." Compare that to the chemistry results that first made me skeptical of IBM’s application claims, where the processor generated samples and a classical solver extracted the conclusion. Here it runs the other way round.

zero classical sizes only (21, 28, 35) with 51 and 74 qubits −5 0 5 10 15 asymptotic oscillation amplitude (× 10⁻³) zero classical sizes only with 51 and 74 qubits −505 1015 asymptotic amplitude (× 10⁻³)
Figure 2: The classically simulable sizes cannot distinguish a fading artifact from a persistent effect, because their interval crosses zero and even runs negative. Adding the two error-mitigated points moves the estimate clear of it.

The price of the correction

None of this comes from raw hardware. Unmitigated, the 51-qubit execution loses accuracy by about the fourth cycle, and the reported dynamics run to thirty.1 Almost everything interesting happens after software repairs the measurementA physical process that produces a classical outcome and updates the quantum state. In an ideal projective measurement, the state is left in an eigenstate associated with the observed outcome..

Two different corrections do that work and the difference matters. The first is formally unbiased provided its characterized model of the device noise is accurate, built on probabilistic error cancellation.3 It agrees with classical results out to roughly eleven cycles and stays affordable to sixteen. Past that the sampling cost becomes prohibitive and a cheaper method takes over, one the authors themselves describe as heuristic and capable of introducing bias.1 The headline late-cycle results live there.

Partial, but not IBM checking IBM.

So they build a ladder of checks: the two estimators against each other where both run, the cheaper one against small systems with exact answers, the noise model validated separately, different noise amplifications compared. And then a rung that is not IBM at all. Selected cycles were repeated on Quantinuum trapped-ion hardware using a simplified scheme that needs no characterization of the device, and the two agree.1 Trapped ionsA qubit made from a single electrically charged atom held in place by electromagnetic fields and controlled with lasers. fail differently from superconducting qubitsA qubit made from tiny electrical circuits chilled to near absolute zero, where they lose all electrical resistance. and Quantinuum is not IBM, so that agreement is not one vendor marking its own homework. It is partial, covering selected cycles rather than the full run, but it is the first evidence in this story that does not come from the machine under question.

And the problem was chosen

One limit the authors document but do not list as a limitation. Before running anything they scanned parameter space on a system small enough to solve exactly. Most settings were poor candidates: some heated too fast to leave a usable signal, others built complexity too slowly to be interesting. They selected a point near the tradeoff frontier between a readable magnetization and high entanglementA quantum link where two qubits' states become tied together, so acting on or measuring one affects the other. It is a key resource for quantum computing., and they write that the choice should challenge both major families of classical simulation. The circuits are described as hardware-native, laid out on IBM’s own heavy-hex lattice.1

Evidence about a problem chosen to suit the instrument, not about problems in general.

That is normal. Every experiment picks a system it can actually run, and the paper is transparent about all of it. But it changes what the result proves. This is evidence about a problem chosen to suit the instrument, not about problems in general.

My takeaway. This is the one of the three that produced something worth knowing rather than something worth measuring. The conclusion exists because the machine reached sizes exact simulation cannot, which is the reverse of the pattern that made me skeptical. It is also the one whose most interesting result rests on a heuristic correction, on a problem selected to sit where classical methods struggle.

The investor’s read

The authors bound their own claim more tightly than most coverage will. The scaling is phenomenological. The systems studied are ladders at most two heavy-hex plaquettes wide, so the large-system inference concerns that geometry rather than genuinely two-dimensional matter. Better classical algorithms may still reproduce the result, and they invite the community to try. Most striking, they distinguish this from an asymptotic complexity separation based on error mitigation, which they say is essentially excluded on fundamental grounds.1

The costs point the same way part one did. One of the largest tensor-network calculations consumed about 600 wall-clock hours and lost reliability around the twelfth cycle, while the mitigated cycle-30 estimate required well under an hour of amortized quantum-processor time.1 Those are different accounting units, so I would not turn them into a speedup ratio. And the cheap number belongs to the heuristic. The formally unbiased estimator becomes statistically prohibitive beyond cycle 16, and its unbiasedness was always conditional on the noise model being accurate.

For still larger problems the authors say error-corrected processors will be required, with mitigation likely continuing on top of them.1 Mitigation is part of the bridge to error correctionTechniques that combine many shaky physical qubits into fewer reliable ones, so a long calculation stays correct., and the authors expect it to keep mattering into the early fault-tolerant era.

Where this leaves the question

So the answer to part one’s complaint is yes, with conditions. A quantum machine produced a result the classical calculations in this study could not settle, in a regime chosen to be hard, using a correction whose most controlled version stops before the deepest regime.

Which leaves one thing unexamined. In both papers so far, the deepest headline regime is ultimately carried by a lower-overhead heuristic rather than by the most controlled estimator available. The third of IBM’s claims takes a different route entirely, deriving its guarantee from the structure of the circuit rather than from a model of the hardware. Part three takes that apart, and asks what all three together are worth.

Sources & notes

  1. E. Leviatan, T. Watad, R. Perry, L. Broers, M. Z. Mullath, O. Alberton, I. Arad, Y. Atia, E. Bairey, S. Barkan, M. Ben Dov, A. Berkovitch, E. van den Berg, I. Cohen, O. Golan, I. Gurwich, A. Haber, B. A. Katzir, O. Kenneth, R. Levi, Y. Y. Lifshitz, Y. Lukovsky, R. Melcer, A. Meyer, B. Muratov, A. Panahi, G. Schul, T. Shnaider, M. Shutman, A. Seif, T. Shirakawa, A. Sinay, V. P. Su, H. Tepanyan, O. Trebitch, A. Zubida, D. Aharonov, H. Gharibyan, A. Kandala, S. Yunoki and N. H. Lindner, "Resolving Structure in Prethermal Floquet Dynamics with Precision Quantum Computation," arXiv:2607.24937 (2026).
  2. D. A. Abanin, W. De Roeck and F. Huveneers, "Exponentially Slow Heating in Periodically Driven Many-Body Systems," Physical Review Letters 115, 256803 (2015).
  3. E. van den Berg, Z. K. Minev, A. Kandala and K. Temme, "Probabilistic error cancellation with sparse Pauli-Lindblad models on noisy quantum processors," Nature Physics 19, 1116 (2023).