Skip to content

A Benchmark Ladder for Quantum Field Simulators

A benchmark ladder should increase the scientific burden one controlled capability at a time: local algebra, free dynamics, interacting equilibrium, real-time response, gauge constraints, scattering or another QFT observable, regulator scaling, and finally a held-out scientific result with matched resources. A simulator occupies the highest rung for which all lower requirements and their failure tests pass. No rung by itself establishes a current quantum advantage.

Required background. Resource and continuum certification supplies accuracy-matched logical and classical costs.

Helpful background. Contraction, truncation, and continuum error certification supplies independent classical errors. Benchmark theories for truncation methods supplies exact and cross-basis fixtures.

Convention and regulator card. A benchmark record freezes target Hamiltonian, regulator, sector, state, observable, accuracy, confidence, and resource accounting before execution. “Pass” means a preregistered numerical tolerance with covariance and a negative control. Training observables used for calibration do not count as held out. Current platform and advantage status are never encoded in the durable rung definitions.

  1. Encoded algebra. Reproduce local commutators or finite group algebra, operator matrix elements, Hermiticity, and exact constraints on a small subspace.
  2. Free QFT dynamics. Recover analytic dispersion, vacuum covariances, unequal-time phases, and conservation laws while varying encoding and algorithmic cutoffs.
  3. Interacting finite-regulator equilibrium. Match exact diagonalization or controlled classical results for energies plus held-out correlators at nonzero coupling.
  4. Interacting real-time response. Match a time series, response function, or quench over a predeclared window, with later times held out.
  5. Physical gauge sector. When applicable, preserve Gauss constraints and match a gauge-invariant spectrum and observable across link cutoffs.
  6. QFT process observable. Extract scattering, flux, particle production, inclusive response, or another renormalized finite-regulator quantity with preparation and measurement errors included.
  7. Cross-regulator continuum trend. Follow a line of constant physics while varying aa, LL, dlocd_{\rm loc}, and time window separately; predict a withheld regulator point.
  8. Matched scientific computation. Produce a held-out result with total logical and classical resources, then compare against the best controlled classical workflow for the identical task.

The last rung supports a scientific computation at the reported regulator and uncertainty. Whether it supports a comparative advantage is a dated Research assessment requiring current baselines and implementation evidence. The logical possibility of efficient universal quantum simulation Lloyd 1996 and the polynomial scalar-QFT algorithm of Jordan, Lee, and Preskill 2012 motivate the ladder; neither substitutes for passing it.

Every rung reports:

  • target theory, physical parameters, regulator and boundary conditions;
  • encoding, local dimension, constraints, state preparation, evolution, and observable;
  • preregistered accuracy and confidence, raw and accepted samples, and covariance;
  • regulator, preparation, algorithmic, synthesis, noise, mitigation, measurement, and statistical errors;
  • logical qubits, compiled logical operations, depth or analog time, success probability, classical preprocessing, and analysis cost;
  • exact or analytic fixture, negative control, held-out prediction, matched baseline, shared inputs, and evidence ceiling.

These fields instantiate the chapter’s canonical claim–resource–evidence record. A compact publication may link to the complete machine-readable record, but it should not omit which observable and uncertainty the resources purchased.

The figure organizes the last three rungs. Inspect why continuum certification and matched advantage are sequential gates rather than labels attached to circuit size.

Benchmark evidence progresses from finite-regulator verification through independent continuum scaling and matched scientific baselines; executable suites require reproducible records and changing advantage status belongs to dated Research.

The benchmark ladder separates durable evidence classes. This Volume defines target accuracy, symbolic resources, verification, and continuum requirements; executable suites require frozen evidence; Research assesses current hardware capability, scientific utility, and advantage against dated matched baselines. The schematic boundary prevents a finite circuit or qubit count from skipping a scientific rung.

Use the periodic scalar Hamiltonian

H=xa[πx22+m02ϕx22+(ϕx+aϕx)22a2+λ0ϕx44!].H=\sum_xa\left[ \frac{\pi_x^2}{2}+\frac{m_0^2\phi_x^2}{2} +\frac{(\phi_{x+a}-\phi_x)^2}{2a^2} +\frac{\lambda_0\phi_x^4}{4!}\right].

Take [ϕx,πy]=iδxy/a[\phi_x,\pi_y]=i\delta_{xy}/a and a unitary site Fourier transform. At rung 1, test the projected canonical commutator and matrix elements of ϕ\phi, π\pi, HH, and the measured operator. At rung 2, set λ0=0\lambda_0=0 and reproduce

ωk2=m02+4a2sin2(ka/2),C(t,x)=1Nkeikxiωkt2aωk.\omega_k^2=m_0^2+4a^{-2}\sin^2(ka/2), \qquad C(t,x)=\frac1N\sum_k\frac{e^{ikx-i\omega_kt}}{2a\omega_k}.

At rung 3, enable a small λ0\lambda_0 and compare the lowest gaps and two correlators with exact diagonalization; reserve a third correlator. At rung 4, compare early times and predict later pre-reflection times. Rung 6 uses separated wave packets and a vacuum-subtracted outgoing flux. Rung 7 follows tuned masses across at least several aa values while varying LL and dlocd_{\rm loc} independently. Rung 8 freezes one observable and lets quantum and classical workflows predict it with the same tolerance and total-cost definition.

The rung is the minimum passed requirement. If the free dynamics passes but the interacting held-out correlator fails, the result remains at rung 2 even if a longer circuit was executed.

A matched comparison uses the same:

(theory,P,R,ρ,O,ϵ,1αstat).(\text{theory},\mathcal P,R,\rho,O,\epsilon,1-\alpha_{\rm stat}).

It includes preprocessing, tuning, rejected runs, mitigation, extrapolation, and uncertainty analysis on both sides. Classical baselines may include exact diagonalization, tensor networks, Hamiltonian truncation, Euclidean methods, perturbation theory, or specialized algorithms; their own truncation and convergence errors must be included. Shared regulator data and renormalization factors create covariance.

A durable claim has the form:

For the specified finite regulator and observable, the encoded protocol agrees with the named exact or controlled baseline within the preregistered tolerance over the stated state, time, and parameter domain, using the reported logical and classical resources. Continuum or comparative conclusions are limited to the separately demonstrated evidence.

Replace every placeholder with quantities. An advantage statement additionally names the dated classical baseline, physical implementation resources, total uncertainty, and comparison date in Research.

Harder circuit, easier observable. Circuit depth grows, but the measured local density is classically trivial. Benchmark the scientific task, not circuit complexity.

Held-out after looking. A time point is called blind only after other points reveal the answer. Freeze the observable and analysis before execution or use an external custodian.

One diagonal regulator path. aa decreases while LL and dlocd_{\rm loc} grow together. Apparent convergence cannot identify which cutoff controls the shift. Vary axes independently.

Classical baseline deliberately weak. The comparison omits symmetry, tensor structure, or a known specialized method. Record baseline selection and invite independent classical review.

Mitigation cost omitted. Extrapolation improves the central value, but shot overhead and model uncertainty are excluded. Compare at equal final error and confidence.

  • Freeze the rung, target, parameters, regulator, state, observable, tolerance, confidence, and negative control.
  • Require all lower rungs, not only the visually most advanced demonstration.
  • Separate calibration quantities from held-out observables, times, sectors, and regulator points.
  • Match quantum and classical workflows in task, accuracy, confidence, preprocessing, success overhead, and continuum treatment.
  • Vary local dimension, volume, time window, lattice spacing, algorithmic step, and mitigation scale independently.
  • State the exact claim ceiling and route changing hardware, utility, and advantage judgments to a dated Research dossier.

A gauge simulator preserves Gauss’s law, matches the free spectrum, and produces an interacting real-time curve, but no exact or controlled interacting point was reserved. What is the highest supported rung?

Solution

Rung 2 is secure. The gauge constraint test is valuable but does not supply the missing interacting equilibrium and held-out real-time baselines required by rungs 3–5. Rungs are cumulative.

The quantum calculation targets absolute error 0.020.02 at fixed aa, while the classical result quotes continuum error 0.010.01 after extrapolation. Can their runtimes establish an advantage?

Solution

No. They solve different tasks and include different regulator errors. Compare both at the same finite regulator or both through the same continuum target, with equal total tolerance and confidence and all preprocessing included.

After working this page, you should be able to:

  • Place a scalar or Abelian gauge simulator on a cumulative benchmark ladder using exact limits, constraints, held-out observables, independent classical baselines, regulator scaling, and complete resources.
  • Write a claim whose theory, regulator, state, time, observable, accuracy, confidence, resource scope, and evidence ceiling exactly match what the benchmark demonstrated.

The chapter closes here. Executable scalar and gauge benchmark suites and cross-formulation comparisons require independent, reproducible implementations. Any statement about current hardware capability, scientific utility, or quantum advantage requires a dated Research dossier with a matched contemporary classical baseline.