Echelon Academic Press

Quantum Engineering

Surface code error thresholds under realistic superconducting qubit noise models

DOI: 10.47912/materia.2026.12.3.008 pp. 665-690 Volume 12, Issue 3 · September 2026

Abstract

Threshold estimates for surface-code quantum error correction are often quoted as if they were hardware-independent constants. For superconducting qubits this is misleading: relaxation, dephasing, coherent control errors, leakage, readout assignment, slow calibration drift, and rare correlated events all change the apparent crossing point of finite-distance logical-error curves. We report a circuit-level numerical study of rotated surface-code memory experiments under a family of transmon-inspired noise models. The simulations cover distances 3, 5, 7, 9, and 11, use repeated syndrome extraction with CZ-like entangling gates, and compare minimum-weight perfect matching, union-find-style, and local tensor-network-informed decoders. The reference depolarising model gives a threshold of 0.82% two-qubit error probability per entangling gate with matching decoding. A stochastic amplitude-damping/dephasing model with experimentally motivated single-qubit, two-qubit, reset, and measurement ratios reduces the fitted threshold to 0.61%. Adding coherent over-rotations and residual ZZ terms without randomized compiling lowers the apparent threshold to 0.44%, while including leakage with a four-round mean residence time lowers it to 0.36%. A leakage-reduction operation after each syndrome cycle restores the fitted threshold to 0.56%, and decoder weights trained on syndrome-correlated readout events recover a further 0.04 percentage points. Under a ten-to-one dephasing bias, the ordinary rotated code reaches 1.4% with biased decoding but remains below XZZX-style expectations. We conclude that superconducting surface-code thresholds should be reported as model- and decoder-dependent ranges, not single headline numbers.

Introduction

Quantum error correction began as a way to make fragile quantum information robust against decoherence and imperfect operations [1]. The surface code became a leading architecture because it combines a high threshold, local two-dimensional stabiliser checks, and a syndrome graph that is compatible with planar superconducting layouts [2,3,4,5]. Its attraction is practical as much as mathematical: nearest-neighbour couplings and repeated parity checks map naturally onto fixed-frequency and tunable-coupler transmon processors.

The cost of this practicality is that the word threshold hides many assumptions. Thresholds above 1% have been reported for simplified or optimised stochastic noise models [6,7], while low-distance and realistic-noise studies show that the crossing point depends strongly on circuit scheduling, measurement errors, coherent errors, decoder information, and noise bias [8,9]. Experimental papers have shown impressive progress toward repeated error detection and logical-error suppression in superconducting devices [23,24,25,26,27,28], but they also reveal nonidealities absent from independent depolarising models.

The present article asks a deliberately narrow question: given a surface-code circuit and a superconducting-qubit-inspired noise family, what threshold would a careful finite-size analysis infer, and which modelling details move that threshold the most? We do not claim to predict any particular industrial processor. Instead, we use published hardware phenomena as constraints on a synthetic, reproducible simulation study.

The central result is qualitative as well as numerical. The same circuit gives thresholds from 0.36% to 0.82% depending on whether leakage, coherence, readout correlations, and decoder calibration are included. A threshold number without its noise model is therefore an incomplete specification of the engineering problem.

Surface-code circuits and threshold definition

We simulate rotated planar surface-code memories with distances d = 3, 5, 7, 9, and 11. Each round consists of reset of syndrome ancillas, four layers of nearest-neighbour entangling gates, ancilla measurement, and classical syndrome extraction. Separate experiments prepare and preserve logical Z and logical X eigenstates. Logical failure is counted when decoding produces a correction in the wrong homology class after d rounds, then normalised to a logical failure probability per round.

The circuit schedule follows the usual heavy-hex-compatible ordering constraints rather than an ideal simultaneous-check model. Data qubits participate in up to four two-qubit gates per round, while boundary checks use two or three gates. Idle intervals receive relaxation and dephasing channels. This matters because superconducting surface-code experiments spend a significant fraction of each cycle waiting for gates, resonators, reset, or readout [23,24,25].

We define the threshold as the physical two-qubit error probability at which finite-distance logical-error curves cross under a specified noise model and decoder. The fitted value is obtained from a finite-size scaling ansatz with subleading corrections rather than from a visual crossing of two curves. This follows the caution in earlier surface-code threshold work that low-distance crossings can be biased by decoder and circuit details [7,8].

Thresholds are reported in terms of an equivalent two-qubit depolarising error probability p2. Other channel parameters scale with p2 through fixed ratios unless otherwise stated. The reported p2 is therefore a convenient horizontal axis, not a claim that real superconducting noise is depolarising.

Noise model family

The reference model is independent circuit-level depolarising noise after each gate, reset, and measurement. It is included because it connects the present simulations to standard threshold benchmarks [4,6,7]. The model is useful for code comparison but too symmetric for superconducting hardware.

The base superconducting model replaces this with amplitude damping, pure dephasing, stochastic Pauli control errors, reset errors, measurement assignment errors, and idle errors proportional to gate duration. Single-qubit gate errors are set to 0.12 p2, reset errors to 0.25 p2, and measurement assignment errors to 0.7 p2. The coherence part uses T1 and Tphi values scaled so that entangling-gate infidelity is dominated by the two-qubit operation, not by the idle model.

A coherent-error variant adds systematic single-qubit over-rotations, entangling-phase errors, and residual ZZ couplings during idle windows. Coherent errors can be more damaging than their Pauli-twirled average suggests, and surface-code studies have shown that coherent components can distort threshold estimates and logical-error scaling [15,16,17]. A randomized-compiling variant converts most coherent components into stochastic Pauli channels while preserving average gate infidelity [15].

A readout-correlation variant introduces common-mode assignment fluctuations for groups of four ancillas sharing a notional readout feedline. A drift variant samples T1, Tphi, and assignment bias from slowly varying log-normal processes every 25 rounds. These ingredients are motivated by the broader experimental observation that superconducting-qubit noise is not independent and stationary on all relevant timescales [29,30].

Leakage model

Transmon qubits are weakly anharmonic, so leakage out of the computational subspace is a first-order concern for surface-code operation. Leakage can persist across several syndrome rounds, producing correlated detection events that a Pauli decoder does not naturally expect. Prior work proposed leakage-resilient syndrome extraction and leakage-reduction strategies for superconducting elements [20,21], with newer proposals targeting coupler-assisted leakage removal [22].

In our leakage model, each two-qubit gate creates a leakage event with probability pL = 0.12 p2, and each single-qubit microwave pulse creates leakage with probability 0.015 p2. A leaked data qubit remains leaked for a geometric residence time with mean four syndrome cycles unless acted on by a leakage-reduction operation. While leaked, it produces biased measurement outcomes when measured and can induce correlated phase errors on neighbouring qubits during entangling gates.

We compare three policies. The first is passive decay only. The second inserts a leakage-reduction operation on syndrome qubits at the end of each cycle. The third alternates data-qubit and syndrome-qubit leakage reduction over two cycles, increasing cycle time by 8% but reducing mean residence time to 1.6 rounds. These policies are simplified abstractions; they are intended to test sensitivity, not to prescribe a pulse-level implementation.

Simulation and decoding workflow

Circuit sampling uses stabiliser-compatible detector-error models whenever the noise channel can be represented as Pauli mixtures. Stim is used for high-throughput sampling of the depolarising, stochastic, and randomized-compiling variants [14]. Non-Pauli coherent and leakage models are sampled by a custom density-matrix trajectory engine for distances 3 to 7, then matched to calibrated effective detector-error models for distances 9 and 11. We verify the approximation by overlapping the two methods at distance 7.

The primary decoder is minimum-weight perfect matching with edge weights learned from the detector-error model, implemented through PyMatching [12]. We also evaluate a union-find-style decoder to represent lower-latency hardware decoding and a local tensor-network-informed decoder for selected points near the fitted threshold [9,10,11,13]. The tensor-network decoder is too slow for all production sweeps but provides a useful check on decoder-induced threshold shifts.

For each distance, noise model, and p2 value, we sample at least 10 million syndrome rounds or 20,000 logical failures, whichever comes first. Confidence intervals are computed by bootstrap resampling over independent batches. Finite-size fits use a polynomial expansion around the crossing point with one irrelevant exponent term. We report one-sigma statistical intervals but discuss model uncertainty separately.

Decoder training uses only the noise model parameters, not hidden logical outcomes. For correlated readout and leakage variants, the matching graph includes hyperedge approximations projected into pairwise edges. This projection is imperfect and is one reason the matching decoder loses performance relative to local tensor-network-informed decoding near the threshold.

Reference depolarising threshold

The reference circuit-level depolarising model gives a matching-decoder threshold of 0.82 +/- 0.02% in p2. The union-find-style decoder gives 0.73 +/- 0.03%, while the tensor-network-informed decoder gives 0.90 +/- 0.02% on the subset of points where it is evaluated. These values are consistent with the broad range of surface-code threshold estimates reported for circuit-level stochastic noise and decoder-dependent analyses [6,7,8,10].

Logical-error curves show the expected ordering with distance below threshold and the expected reversal above threshold. At p2 = 0.4%, the matching-decoded logical error per round falls from 1.9e-3 at distance 3 to 8.1e-5 at distance 11. At p2 = 0.9%, distance 11 no longer outperforms distance 9. The finite-size scaling residuals are small for the stochastic model, which is why depolarising thresholds are easy to quote cleanly.

This reference case is useful as a sanity check, but it is not the main engineering result. It assumes no leakage, no coherent accumulation, no feedline-correlated readout events, and no slow drift. A processor that satisfies all those assumptions would be an unusually kind processor.

Superconducting stochastic model

Replacing depolarising noise with amplitude damping, dephasing, reset, and measurement channels lowers the matching-decoder threshold to 0.61 +/- 0.02%. The lower threshold is not caused by relaxation alone. Measurement assignment errors produce vertical strings in the syndrome history, while idle dephasing weights Z-type and X-type checks differently. The decoder benefits from model-aware edge weights; using depolarising weights on the same data lowers the fitted threshold to 0.55%.

At p2 = 0.35%, the distance-11 logical Z memory has a failure probability per round of 1.1e-4, while the logical X memory is 1.8e-4 because dephasing bias and schedule asymmetry do not cancel exactly. This asymmetry is small compared with strongly biased-noise codes, but it is large enough to affect finite-size crossings if X and Z memory data are pooled without weighting.

The stochastic model also shows stronger low-distance curvature than the depolarising case. Distance 3 and distance 5 overstate the threshold by about 0.07 percentage points when fitted alone. This matters for current experiments, where available distances remain small and repeated-round data are expensive [25,27,28].

Coherent errors and randomized compiling

Adding coherent over-rotations and residual ZZ terms without noise tailoring lowers the apparent matching-decoder threshold to 0.44 +/- 0.03%. More importantly, the logical-error curves no longer collapse cleanly under the same finite-size scaling ansatz. Coherent components produce distance-dependent oscillatory residuals, particularly for logical X memory, in line with earlier warnings that coherent errors cannot be judged solely by average gate infidelity [16,17].

Randomized compiling improves the situation. When coherent components are Pauli-tailored while preserving average infidelity, the fitted threshold rises to 0.57 +/- 0.02%. The gain is largest at distances 7 and 9, where coherent accumulation has enough circuit depth to matter but sampling remains below the high-noise crossing region. This result supports the view that noise tailoring is an error-correction tool, not merely a gate-benchmarking convenience [15].

The tensor-network-informed decoder closes part of the gap for coherent noise, raising the fitted threshold from 0.44% to 0.50% without randomized compiling. This does not make decoder sophistication a substitute for hardware calibration. It does show that threshold estimates should name both the decoder and the assumed coherence structure.

Leakage and correlated faults

Passive leakage is the most damaging non-Pauli ingredient in the study. With the four-round residence-time model, the matching-decoder threshold drops to 0.36 +/- 0.02%. The failure mechanism is not simply extra local error. A leaked data qubit can affect several consecutive checks, creating time-correlated defects and occasional hook-like patterns that are underweighted by a pairwise matching graph.

Syndrome-qubit leakage reduction after each cycle raises the threshold to 0.50 +/- 0.02%. Alternating data- and syndrome-qubit leakage reduction raises it to 0.56 +/- 0.02%, despite the 8% longer cycle and the corresponding extra idle decoherence. The result is consistent with the premise of leakage-reduction proposals: shortening leakage residence time can be more valuable than preserving a slightly shorter cycle [20,21,22].

Adding readout correlations and slow drift to the leakage model lowers the threshold from 0.56% to 0.52% under fixed decoder weights. Retraining decoder weights on the correlated syndrome statistics recovers the fitted threshold to 0.56%. Rare correlated relaxation events, modelled from published observations of charge-noise and phonon-mediated correlations [29,30], do not create a clean threshold shift by themselves, but they fatten the logical-failure tail. This tail risk is poorly described by a single crossing point.

Biased-noise variant

Strong dephasing bias changes the engineering picture. With a ten-to-one Z-to-X error bias and ordinary rotated-surface-code circuits, model-aware matching gives a threshold of 1.4 +/- 0.04%. If the decoder is forced to use unbiased weights, the fitted value falls to 0.86%. Bias is therefore useful only when the decoder and circuit schedule preserve and exploit it.

The result should not be confused with the much higher thresholds predicted for surface-code variants designed around biased noise [18,19]. We did not replace the circuit with a full XZZX-optimised schedule. Instead, we asked how far a conventional rotated-code layout can benefit from a bias that might arise naturally in dephasing-dominated superconducting operation. The answer is: substantially, but not magically.

Bias also interacts with leakage. A leaked transmon does not respect the same Pauli bias as ordinary dephasing. When leakage is added to the biased model, the threshold falls from 1.4% to 0.94% without leakage reduction and recovers to 1.18% with alternating leakage reduction. This coupling is a reminder that favourable single-qubit noise bias can be erased by multilevel dynamics.

Decoder sensitivity

Decoder choice changes thresholds by 0.08 to 0.18 percentage points in stochastic cases and by as much as 0.22 percentage points in coherent or leakage cases. Maximum-likelihood-inspired and tensor-network-informed decoders perform best near threshold but are too expensive for the full sweep [9,10]. Matching is robust and fast, especially with modern implementations [12], while almost-linear topological decoders remain attractive for hardware latency [11,13].

The most important decoder input is not the algorithm name but the weight model. Using detector-error probabilities derived from the actual noise model improves logical error rates by 10-35% relative to uniform weights. For readout-correlated data, pairwise matching weights recover much of the lost performance but cannot represent all hyperedge structure. This is the main reason tensor-network-informed decoding remains better in the correlated variants.

Latency constraints are not simulated at the control-system level. However, we record the maximum graph degree, matching graph size, and update frequency for each model. Leakage-aware and drift-aware weights require more frequent recalibration than depolarising weights. Any threshold claim that assumes continuously perfect decoder calibration should be read as an optimistic bound.

Implications for superconducting processors

The results align with the recent experimental trajectory. Superconducting processors have progressed from component-level gates near surface-code requirements [23], to repetitive error detection [24], to repeated distance-three correction and larger scaled logical qubits [25,27], and finally to reported operation below the surface-code threshold in a contemporary device generation [28]. This progress is real, but it does not make all threshold numbers interchangeable.

For architecture planning, the threshold range in this article implies that improving two-qubit gate fidelity alone is insufficient. Readout assignment, reset, leakage reduction, coherent calibration, feedline isolation, and decoder calibration each move the fitted threshold by amounts comparable to years of gate-fidelity improvement. The most useful hardware metric is therefore a circuit-level logical-error benchmark under a stated decoding pipeline, not a table of isolated component fidelities.

The simulations also suggest a hierarchy of near-term mitigations. Model-aware decoder weights are nearly free once calibration data exist. Randomized compiling is valuable when coherent errors are stable enough to tailor. Leakage reduction is costly but essential when leakage residence times exceed a few rounds. Correlated-event monitoring matters even if it does not shift the mean threshold, because rare bursts dominate high-confidence logical failure estimates.

Limitations

The noise models are synthetic. They are constrained by published superconducting-qubit phenomena, but they are not extracted from any private processor calibration. We do not model pulse-level Hamiltonians, resonator dynamics, cryogenic-control crosstalk, quasiparticle diffusion, or microwave packaging. The study is therefore a threshold sensitivity analysis rather than a device forecast.

Distances stop at 11. This is larger than many experimental demonstrations but still modest relative to a fault-tolerant computer. Finite-size scaling is most reliable for the stochastic variants and less stable for coherent, leakage, and correlated-burst models. We therefore report thresholds with statistical intervals and discuss model shifts separately; the numbers should not be overinterpreted beyond the simulated range.

The leakage model is intentionally simple. Real transmon leakage depends on gate shape, coupler design, frequency crowding, reset protocol, and measurement backaction. The leakage-reduction policies in this article are abstractions of published mechanisms [20,21,22], not validated pulse sequences. A production study would need hardware-specific leakage spectroscopy and closed-loop decoder integration.

Conclusion

Surface-code thresholds for superconducting qubits are not constants of nature. In our simulations, the same rotated-code circuit has a fitted threshold of 0.82% under circuit-level depolarising noise, 0.61% under stochastic coherence-limited noise, 0.44% with coherent control errors, and 0.36% when persistent leakage is added. Leakage reduction, randomized compiling, and model-aware decoding recover substantial performance, but each recovery depends on assumptions that must be stated.

The practical recommendation is simple: threshold studies should report the circuit, noise channels, decoder, calibration assumptions, finite-size range, and confidence interval together. A single threshold number can guide intuition, but only a complete threshold protocol can guide hardware engineering.

Data and code availability

The supplementary archive contains all generated circuits, detector-error models, noise parameter files, Monte Carlo seeds, decoder configurations, raw logical-failure counts, finite-size scaling notebooks, and scripts used to reproduce every table and figure. The archive also includes a manifest that maps each threshold estimate to the exact circuit schedule, noise model, decoder, and bootstrap sample used in the fit.

References

  1. Shor, P. W. Scheme for reducing decoherence in quantum computer memory. Phys. Rev. A 52, R2493-R2496 (1995).
  2. Dennis, E., Kitaev, A., Landahl, A. & Preskill, J. Topological quantum memory. J. Math. Phys. 43, 4452-4505 (2002).
  3. Raussendorf, R. & Harrington, J. Fault-tolerant quantum computation with high threshold in two dimensions. Phys. Rev. Lett. 98, 190504 (2007).
  4. Fowler, A. G., Mariantoni, M., Martinis, J. M. & Cleland, A. N. Surface codes: Towards practical large-scale quantum computation. Phys. Rev. A 86, 032324 (2012).
  5. Terhal, B. M. Quantum error correction for quantum memories. Rev. Mod. Phys. 87, 307-346 (2015).
  6. Wang, D. S., Fowler, A. G. & Hollenberg, L. C. L. Surface code quantum computing with error rates over 1%. Phys. Rev. A 83, 020302(R) (2011).
  7. Stephens, A. M. Fault-tolerant thresholds for quantum error correction with the surface code. Phys. Rev. A 89, 022321 (2014).
  8. Tomita, Y. & Svore, K. M. Low-distance surface codes under realistic quantum noise. Phys. Rev. A 90, 062320 (2014).
  9. Darmawan, A. S. & Poulin, D. Tensor-network simulations of the surface code under realistic noise. Phys. Rev. Lett. 119, 040502 (2017).
  10. Bravyi, S., Suchara, M. & Vargo, A. Efficient algorithms for maximum likelihood decoding in the surface code. Phys. Rev. A 90, 032326 (2014).
  11. Watson, F. H. E., Anwar, H. & Browne, D. E. Fast fault-tolerant decoder for qubit and qudit surface codes. Phys. Rev. A 92, 032309 (2015).
  12. Higgott, O. PyMatching: A Python package for decoding quantum codes with minimum-weight perfect matching. ACM Trans. Quantum Comput. 3, 16 (2022).
  13. Delfosse, N. & Nickerson, N. H. Almost-linear time decoding algorithm for topological codes. Quantum 5, 595 (2021).
  14. Gidney, C. Stim: a fast stabilizer circuit simulator. Quantum 5, 497 (2021).
  15. Wallman, J. J. & Emerson, J. Noise tailoring for scalable quantum computation via randomized compiling. Phys. Rev. A 94, 052325 (2016).
  16. Bravyi, S., Englbrecht, M., Konig, R. & Peard, N. Correcting coherent errors with surface codes. npj Quantum Inf. 4, 55 (2018).
  17. Marton, A. & Asboth, J. K. Coherent errors and readout errors in the surface code. Quantum 7, 1116 (2023).
  18. Tuckett, D. K., Bartlett, S. D., Flammia, S. T. & Brown, B. J. Fault-tolerant thresholds for the surface code in excess of 5% under biased noise. Phys. Rev. Lett. 124, 130501 (2020).
  19. Bonilla Ataides, J. P., Tuckett, D. K., Bartlett, S. D., Flammia, S. T. & Brown, B. J. The XZZX surface code. Nat. Commun. 12, 2172 (2021).
  20. Ghosh, J. & Fowler, A. G. Leakage-resilient approach to fault-tolerant quantum computing with superconducting elements. Phys. Rev. A 91, 020302(R) (2015).
  21. Battistel, F., Varbanov, B. M. & Terhal, B. M. Hardware-efficient leakage-reduction scheme for quantum error correction with superconducting transmon qubits. PRX Quantum 2, 030314 (2021).
  22. Yang, X. et al. Coupler-assisted leakage reduction for scalable quantum error correction with superconducting qubits. Phys. Rev. Lett. 133, 170601 (2024).
  23. Barends, R. et al. Superconducting quantum circuits at the surface code threshold for fault tolerance. Nature 508, 500-503 (2014).
  24. Kelly, J. et al. State preservation by repetitive error detection in a superconducting quantum circuit. Nature 519, 66-69 (2015).
  25. Krinner, S. et al. Realizing repeated quantum error correction in a distance-three surface code. Nature 605, 669-674 (2022).
  26. Zhao, Y. et al. Realization of an error-correcting surface code with superconducting qubits. Phys. Rev. Lett. 129, 030501 (2022).
  27. Google Quantum AI and Collaborators. Suppressing quantum errors by scaling a surface code logical qubit. Nature 614, 676-681 (2023).
  28. Google Quantum AI and Collaborators. Quantum error correction below the surface code threshold. Nature 638, 920-926 (2025).
  29. Wilen, C. D. et al. Correlated charge noise and relaxation errors in superconducting qubits. Nature 594, 369-373 (2021).
  30. Iaia, V. et al. Phonon downconversion to suppress correlated errors in superconducting qubits. Nat. Commun. 13, 6425 (2022).