Small Group Tutorials

Here to help students catch up, keep up, and move ahead. Book a consultation here.

Quantum Mathematics Learning Guide 24: Randomised Benchmarking, Gate Fidelity and Error Budgets

Randomised benchmarking estimates how rapidly repeated gate sequences lose the ability to return a quantum state to an expected reference outcome. It is deliberately less descriptive than process tomography, but often more robust to state-preparation and measurement errors.

Quantum process tomography tries to reconstruct a channel. Randomised benchmarking asks a narrower operational question: on average, how much error accumulates per gate or per Clifford under a specified random sequence ensemble?

The central mathematical observation is that random conjugations can “twirl” a broad class of errors into an effective depolarising description for the averaged survival signal. Under standard assumptions, sequence fidelity then decays approximately exponentially with sequence length.

Random sequence → inverse recovery → survival probability → exponential fit → decay parameter → average error scale → compare against an error budget.

1. Why not just use process tomography?

Process tomography is information-rich but sensitive to assumptions about state preparation and measurement. If the initial state or detector is miscalibrated, the reconstructed gate can absorb those errors.

Randomised benchmarking is designed so that many SPAM effects appear mainly in the prefactor and offset of a decay curve rather than in the decay constant itself, under the model assumptions.

This does not make RB completely SPAM-free. State preparation and measurement still affect statistical precision and can interact with leakage or sequence-dependent effects. The claim is relative robustness, not immunity.

2. The basic Clifford-sequence experiment

Choose a sequence length m. Sample m Clifford gates independently from the benchmarking group. Compute one final recovery Clifford that ideally inverts the entire product.

Prepare a reference state, apply the noisy sequence and recovery, then measure whether the expected reference outcome is obtained.

Repeat for many random sequences at the same m and for many values of m. Average the survival probability over sequences.

3. Why Clifford gates are convenient

Clifford gates map Pauli operators to Pauli operators under conjugation. This finite algebraic structure makes randomisation and classical tracking efficient.

The Clifford group forms a unitary 2-design, sufficient for the standard twirling argument that converts average noise into an effective depolarising channel at the level relevant to basic RB.

Other benchmarking groups and variants exist. The gate ensemble is part of the protocol and must be reported.

4. Depolarising channel and its decay parameter

In d dimensions, write an effective depolarising channel as

𝓓_p(ρ)=pρ+(1-p)I/d.

Applying it m times gives

𝓓_p^m(ρ)=p^mρ+(1-p^m)I/d.

Thus any traceless component of the state shrinks by p each step. The averaged survival signal inherits a p^m dependence.

5. The standard RB fit

A common model is

F(m)=A p^m+B.

A and B absorb state preparation, measurement and recovery conventions under the standard gate-independent noise model. p describes the average contraction.

For a qubit with ideal binary readout and large m, B is often near 1/2, but it should usually be fitted or calibrated rather than fixed blindly.

6. From p to average gate infidelity

For a d-dimensional depolarising channel, average gate fidelity relative to identity is

F_avg=p+(1-p)/d.

Therefore the average infidelity is

r=1-F_avg=(d-1)(1-p)/d.

For a qubit, r=(1-p)/2.

If p=0.998, the corresponding qubit average error scale is 0.001, or 0.1% per benchmarked Clifford under the basic interpretation.

7. Per Clifford is not automatically per physical gate

A randomly chosen Clifford may compile into several native gates. If the average number of native gates per Clifford is g, dividing the Clifford error by g is only an approximation when errors are small and composition is sufficiently simple.

Coherent errors, gate-dependent noise and unequal native-gate fidelities can make the conversion nonlinear.

Whenever possible, benchmark the native gate set directly or use interleaved and cycle benchmarking methods matched to the hardware.

8. Worked decay example

Assume the fitted qubit model is

F(m)=0.47(0.996)^m+0.50.

The depolarising parameter is p=0.996, so the average Clifford infidelity is

r=(1-0.996)/2=0.002.

That is 0.2% per Clifford under the fitted RB model.

At m=100, p^m≈0.6698. The predicted survival is approximately 0.47×0.6698+0.50≈0.8148.

9. Why exponential decay emerges

Under gate-independent Markovian noise, each ideal random Clifford conjugates the error into a different orientation. Averaging over a 2-design removes directional information and leaves an effective depolarising action.

Repeated depolarising action multiplies the traceless component by p each time, producing p^m.

When noise is strongly gate dependent, temporally correlated or non-Markovian, deviations from a single exponential can appear.

10. Sequence-to-sequence variation

At one sequence length, different random Clifford sequences can produce different survival probabilities because the errors do not average identically in every finite sequence.

The experiment therefore samples both measurement shots and random sequences. Statistical uncertainty must account for both sources.

Using thousands of shots on only one random sequence does not estimate the group average well. Sequence diversity is a real experimental resource.

11. Interleaved randomized benchmarking

To estimate the error of a specific gate G, first run reference RB and obtain p_ref. Then run an interleaved experiment in which G is inserted between random Clifford gates, obtaining p_int.

Under a simple depolarising approximation, the isolated gate decay is estimated by

p_G≈p_int/p_ref.

The corresponding average infidelity estimate is (d-1)(1-p_G)/d. Rigorous interleaved-RB analyses include bounds because gate dependence can spoil exact factorisation. [2]

12. Worked interleaved example

Suppose qubit reference RB gives p_ref=0.997 and interleaved RB gives p_int=0.993.

The simple ratio estimate is

p_G≈0.993/0.997≈0.99598796.

The corresponding qubit infidelity is approximately

(1-p_G)/2≈0.002006.

So the gate contributes roughly 0.20% average infidelity under the simple model. A publication-quality result should include interleaved-RB bounds or a model-adequacy discussion rather than only this ratio.

13. Coherent over-rotation can hide inside a small average infidelity

Consider a systematic unitary over-rotation by a small angle ε. Its average infidelity scales approximately as ε² for small ε.

But repeating the same coherent error m times can accumulate an angle mε before randomisation or cancellation. Certain circuits can therefore experience much larger worst-case error than an average infidelity suggests.

Randomized benchmarking compresses much of this structure into one decay rate. Complementary diagnostics are needed to distinguish coherent from stochastic error.

14. Pauli stochastic error behaves differently

If each gate suffers an independent Pauli error with small probability q, the error behaves more like an incoherent random process. Repeated application accumulates differently from a fixed coherent misrotation.

Two noise channels can have the same average gate fidelity while producing very different long-circuit behaviour.

This is why an error budget should contain more than one scalar whenever coherent accumulation, leakage or correlated noise matters to the application.

15. Unitarity benchmarking

Unitarity is a metric designed to quantify how coherent the noise is on the traceless operator subspace. Highly unitary noise behaves more like a coherent rotation; strongly contractive noise behaves more stochastically.

Combining average infidelity with unitarity gives a richer picture of whether error mitigation should focus on calibration of coherent controls or reduction of stochastic decoherence.

The exact unitarity protocol has its own randomisation and fitting assumptions and should not be reduced to “another fidelity number.”

16. Leakage

A computational qubit may leave the intended two-level subspace. Standard RB models assume a fixed d-dimensional trace-preserving channel on the computational space.

Leakage can produce offsets, multi-exponential behaviour or apparently non-trace-preserving dynamics in the reduced subspace.

Leakage-aware benchmarking measures both survival in the computational subspace and logical decay within it. A good error budget separates leakage probability from ordinary in-subspace gate infidelity.

17. Simultaneous benchmarking and crosstalk

A gate can perform differently when neighbouring qubits are driven simultaneously. Benchmarking one qubit in isolation may therefore understate crosstalk in realistic parallel circuits.

Simultaneous RB runs independent random sequences on several subsystems at once and compares decay rates against isolated measurements.

An increased error under simultaneous operation is evidence of context dependence, though further diagnostics are needed to locate the physical mechanism.

18. Cycle benchmarking

Cycle benchmarking targets the error of an entire layer or cycle of gates rather than individual compiled Cliffords. Random Pauli dressing can convert coherent errors into a form whose average fidelity is estimated efficiently.

This can align the benchmark more closely with the unit of parallel hardware execution used by a compiler.

The relevant error unit should match the architecture: gate, cycle, layer or algorithmic primitive.

19. Error per Clifford versus error per layer

An algorithm executes a schedule. Several gates may occur in parallel, some qubits idle, and crosstalk may couple operations.

Adding isolated per-gate infidelities assumes independence and often over- or under-estimates real circuit failure.

An error budget should therefore distinguish primitive calibration metrics, isolated RB metrics, simultaneous metrics and circuit-level validation.

20. Build an error budget from named mechanisms

A useful error budget lists mechanisms rather than one total percentage. Example entries include:

  • energy relaxation during gates;
  • dephasing during gates and idle periods;
  • coherent amplitude or phase miscalibration;
  • leakage out of the computational subspace;
  • crosstalk from simultaneous operations;
  • readout error;
  • state-preparation error;
  • drift between calibrations;
  • compiler approximation and synthesis error.

Not every term is additive in the same metric. The budget is an accounting framework first; combination requires a model.

21. Worked decoherence floor estimate

Suppose a one-qubit gate takes τ=20 ns, with relaxation time T₁=100 μs and pure-dephasing time Tφ=150 μs.

For a rough small-time scale, relaxation probability is approximately τ/T₁=2×10−4. Pure-dephasing scale is approximately τ/Tφ≈1.33×10−4.

These are mechanism probabilities or decay scales, not directly equal to average gate infidelity. The mapping to fidelity depends on channel form and conventions.

Comparing a measured RB error with a decoherence-limited channel model can nevertheless reveal whether control error dominates over unavoidable T₁/T₂ exposure.

22. Fit diagnostics matter

Do not report p from a nonlinear fit without checking residuals. Systematic curvature in log-scale decay, oscillatory residuals or length-dependent variance can indicate model failure.

Compare single- and multi-exponential models only with enough data to justify the extra parameters. A better in-sample fit can simply overfit sequence noise.

Bootstrap random sequences and measurement shots to estimate uncertainty in p when analytic approximations are doubtful.

23. Drift can masquerade as gate dependence

If long sequences are measured later in the day than short sequences, calibration drift can create a false length dependence.

Randomise or interleave sequence lengths in acquisition order. Record timestamps and calibration state.

Benchmarking is an experiment in statistical design as much as a group-theory calculation.

24. Common misconception: RB measures the exact error channel

No. Basic RB intentionally compresses the noise into an average decay parameter. Different channels can produce similar p values.

25. Common misconception: SPAM never affects RB

SPAM is largely absorbed into A and B under standard assumptions, but it still affects signal contrast, uncertainty and any model violation involving leakage, drift or context-dependent preparation and measurement.

26. Common misconception: a 99.9% gate fidelity means a 1000-gate circuit has 36.8% success

Naively multiplying 0.9991000 assumes independent identical stochastic failure with a binary “success” interpretation. Quantum errors can interfere coherently, cancel, leak, correlate and affect observables differently.

The product can be a rough heuristic under a specific stochastic model, not a universal circuit-fidelity law.

27. Worked synthesis problem

A qubit reference RB experiment fits p_ref=0.9985. An interleaved experiment for gate G fits p_int=0.9965.

Step 1: Reference Clifford error. r_ref=(1-0.9985)/2=0.00075, or 0.075%.

Step 2: Interleaved ratio. p_G≈0.9965/0.9985≈0.997997.

Step 3: Gate error estimate. r_G≈(1-0.997997)/2≈0.0010015, about 0.10%.

Step 4: Interpretation. This is an average infidelity estimate under the interleaved depolarising approximation, not a full error channel.

Step 5: Error budget. Compare against decoherence estimates, leakage measurements, coherent calibration checks and simultaneous RB before assigning the entire 0.10% to one physical mechanism.

28. Practice set

  1. What is the basic RB sequence structure?
  2. Why are Clifford gates useful?
  3. Write the standard decay model.
  4. For a qubit, convert p to average infidelity.
  5. Why is error per Clifford not automatically error per native gate?
  6. What does interleaved RB estimate?
  7. Why can coherent and stochastic errors with similar average fidelity behave differently?
  8. What is leakage?
  9. Why run simultaneous RB?
  10. What should fit residuals be used for?
  11. Why should sequence lengths be randomised in acquisition time?
  12. Name four components of a useful hardware error budget.

Answers

  1. Random benchmark gates followed by a recovery that ideally returns the state to a known outcome.
  2. They form a finite efficiently trackable 2-design and map Paulis to Paulis.
  3. F(m)=Ap^m+B.
  4. r=(1-p)/2.
  5. A Clifford compiles into multiple native gates with potentially unequal and coherent errors.
  6. The average error of a selected interleaved gate relative to a reference benchmark, under stated assumptions.
  7. Coherent errors can accumulate directionally while stochastic errors diffuse.
  8. Population leaving the computational Hilbert subspace.
  9. To detect crosstalk or context-dependent degradation during parallel operations.
  10. To test whether the single-exponential model is adequate.
  11. To prevent temporal drift from becoming confounded with sequence length.
  12. Examples: decoherence, coherent miscalibration, leakage, crosstalk, readout error, preparation error and drift.

Sources and further study

[1] Easwar Magesan, J. M. Gambetta and Joseph Emerson, Scalable and robust randomized benchmarking of quantum processes, Physical Review Letters 106, 180504 (2011). A foundational scalable RB protocol and analysis.

[2] Easwar Magesan and colleagues, Efficient measurement of quantum gate error by interleaved randomized benchmarking, Physical Review Letters 109, 080505 (2012). The interleaved RB protocol and error bounds.

[3] J. M. Chow and colleagues, Randomized benchmarking and process tomography for gate errors in a solid-state qubit, Physical Review Letters 102, 090502 (2009). An experimental comparison of process tomography and RB.

[4] Joel J. Wallman and colleagues, Estimating the coherence of noise. This work develops unitarity as a measure of coherent versus stochastic error structure.

Batch 06 series navigation

Educational note: RB metrics are protocol- and model-dependent. Report the gate ensemble, compilation, random-sequence count, shots, fit model, leakage treatment, acquisition order and uncertainty before interpreting a decay parameter as hardware performance.