Skip to content

Anatomy and Severity of a Sign Problem

A sign problem occurs when the exact weight cannot be used as a nonnegative probability measure in the chosen representation. Its practical severity is not one number: phase cancellation controls the denominator of reweighting, overlap controls whether important target configurations are sampled, and correctness controls whether a replacement process computes the intended integral. Confusing these mechanisms produces confident but invalid error estimates.

Required background. Chemical potential on the Euclidean lattice supplies the complex determinant and charge convention.

Helpful background. Probability, random variables, and conditional expectation supplies variance and importance sampling. Complete lattice error budgets supplies the distinction between statistical and method bias.

Convention and regulator card. Work at finite lattice spacing and finite spatial volume VsV_s, with inverse temperature βT\beta_T. Assume both Z=∫dx w(x)Z=\int dx\,w(x) and Zpq=∫dx ∣w(x)∣Z_{\rm pq}=\int dx\,|w(x)| are finite and nonzero. On the support of ∣w∣|w|, write w(x)=∣w(x)∣u(x)w(x)=|w(x)|u(x) with u=eiϕu=e^{i\phi} and ∣u∣=1|u|=1. For a real weight, u=sgn⁡w∈{−1,+1}u=\operatorname{sgn}w\in\{-1,+1\}, so the average sign is the real special case of the average phase. The subscript “pq” means this precisely defined phase-quenched measure; it is not automatically the same physical theory with one parameter changed.

Define phase-quenched expectation by

⟨F⟩pq=1Zpq∫dx ∣w(x)∣F(x).\langle F\rangle_{\rm pq} =\frac{1}{Z_{\rm pq}}\int dx\,|w(x)|F(x).

In plain language: sample with the positive magnitude ∣w∣|w|, then restore the discarded phase in both the observable numerator and the normalization denominator. For any integrable observable,

⟨O⟩=∫dx ∣w∣eiϕO∫dx ∣w∣eiϕ=⟨Oeiϕ⟩pq⟨eiϕ⟩pq.\langle O\rangle =\frac{\int dx\,|w|e^{i\phi}O}{\int dx\,|w|e^{i\phi}} =\frac{\langle Oe^{i\phi}\rangle_{\rm pq}} {\langle e^{i\phi}\rangle_{\rm pq}}.

The average phase is the partition-function ratio

m≡⟨eiϕ⟩pq=ZZpq.m\equiv\langle e^{i\phi}\rangle_{\rm pq} =\frac{Z}{Z_{\rm pq}}.

At finite VsV_s, define the nonnegative free-energy-density difference

ΔfVs≡−1βTVslog⁡∣Z∣Zpq≥0.\Delta f_{V_s} \equiv-\frac{1}{\beta_TV_s}\log\frac{|Z|}{Z_{\rm pq}} \ge0.

Then the severity identity is exact:

∣m∣=exp⁡[−βTVsΔfVs].|m|=\exp[-\beta_TV_s\Delta f_{V_s}].

The inequality follows from ∣Z∣≤Zpq|Z|\le Z_{\rm pq}. If, at fixed βT\beta_T, ΔfVs→Δf∞>0\Delta f_{V_s}\to\Delta f_\infty>0, then

log⁡∣m∣=−βTVsΔf∞+o(Vs),\log|m|=-\beta_TV_s\Delta f_\infty+o(V_s),

and direct phase reweighting has an exponentially small signal. A vanishing thermodynamic limit or subextensive −log⁡∣m∣-\log|m| does not establish exponential severity. Both the finite-volume difference and its limit are representation dependent because ZpqZ_{\rm pq} changes with the magnitude–phase decomposition. See Alexandru et al. 2022, § I.C, pp. 015006-3–015006-4 for the reweighting and free-energy argument.

For NN independent phase samples uk=eiϕku_k=e^{i\phi_k}, the sample mean uˉ\bar u has

E∣uˉ−m∣2=1−∣m∣2N.\mathbb E|\bar u-m|^2=\frac{1-|m|^2}{N}.

Its relative mean-square error is therefore

E∣uˉ−m∣2∣m∣2=1−∣m∣2N∣m∣2.\frac{\mathbb E|\bar u-m|^2}{|m|^2} =\frac{1-|m|^2}{N|m|^2}.

If ∣m∣∼e−βTVsΔf∞|m|\sim e^{-\beta_TV_s\Delta f_\infty}, keeping this relative error fixed requires N∼e2βTVsΔf∞N\sim e^{2\beta_TV_s\Delta f_\infty} independent samples. This is the denominator baseline, not yet the error of a physical observable.

For the actual ratio estimator, set X=OuX=Ou, Y=uY=u, and θ=⟨O⟩=EpqX/EpqY\theta=\langle O\rangle=\mathbb E_{\rm pq}X/\mathbb E_{\rm pq}Y. With θ^=Xˉ/Yˉ\widehat\theta=\bar X/\bar Y on samples for which Yˉ≠0\bar Y\ne0, m=EpqY≠0m=\mathbb E_{\rm pq}Y\ne0, independent draws, and a finite covariance matrix for the real and imaginary parts of (X,Y)(X,Y), the multivariate central limit theorem and delta method give

N(θ^−θ)=1mN∑k=1Nuk(Ok−θ)+op(1),\sqrt N\bigl(\widehat\theta-\theta\bigr) =\frac{1}{m\sqrt N}\sum_{k=1}^{N} u_k\bigl(O_k-\theta\bigr)+o_p(1),

The resulting delta-method asymptotic mean-square error is

AMSE⁡(θ^)=Epq∣u(O−θ)∣2N∣m∣2.\operatorname{AMSE}(\widehat\theta) =\frac{\mathbb E_{\rm pq}|u(O-\theta)|^2} {N|m|^2}.

The centered combination X−θY=u(O−θ)X-\theta Y=u(O-\theta) is where numerator–denominator covariance enters. A finite sample can have Yˉ=0\bar Y=0 or a dangerously small denominator; that is a diagnosed estimator failure, not a value to silently report. Under stronger moment and denominator-regularity conditions, the ratio also has a finite-NN bias of order N−1N^{-1}. For a Markov chain, replace the independent-sample variance by the integrated autocovariance matrix of XX and YY, or equivalently of the influence variable u(O−θ)u(O-\theta); the autocorrelation time of the phase alone is generally insufficient. Resampling can propagate the covariance of configurations that were visited, but it cannot repair absent support.

The figure separates this cancellation law from overlap, ordinary instability, and worst-case complexity. Inspect the arrows: they share a complex measure, but none is logically interchangeable with another.

A complex measure yields separate diagnostics for phase-estimator variance, observable overlap, representation cost, and explicitly quantified asymptotic complexity; none is a universal severity score.

The sign problem has distinct diagnostics. When ΔfVs→Δf∞>0\Delta f_{V_s}\to\Delta f_\infty>0 at fixed βT\beta_T, the average phase scales as e−βTVsΔf∞+o(Vs)e^{-\beta_TV_s\Delta f_\infty+o(V_s)} and fixes direct phase-estimator signal-to-noise; overlap depends on the target observable and proposal measure; a variable change can trade phase for nonlocality or hard observables; and worst-case complexity requires a separately specified problem family. The map is schematic, not a quantitative performance comparison.

Text equivalent: the four diagnostic branches

Phase-estimator variance. Measure it with RelMSE⁡(uˉ)=(1−∣m∣2)/(N∣m∣2)\operatorname{RelMSE}(\bar u)=(1-\lvert m\rvert^2)/(N\lvert m\rvert^2) and the relevant autocorrelation. It does not establish the cost of every correlated ratio estimator.

Observable overlap. Measure it with proposal probability of the target support, log-weight tails, and sector occupancy. It cannot be inferred from mm alone.

Representation cost. Measure residual phase, transformed locality, observable complexity, constraints, and mixing. Nonnegative transformed weights do not by themselves make the computation cheap.

Asymptotic complexity. Specify the input family, allowed transformations, precision, resource, and worst- or average-case quantifier. One exponential fit does not establish NP-hardness.

Let p(x)p(x) be the proposal distribution and let AA be a region important for the target numerator. If p(A)=ε≪1p(A)=\varepsilon\ll1, then the probability of seeing no sample in AA after NN independent draws is

(1−ε)N≃e−Nε.(1-\varepsilon)^N\simeq e^{-N\varepsilon}.

Therefore, to visit AA at least once with probability 1−δ1-\delta, one needs

N≥log⁡δlog⁡(1−ε)≃log⁡(1/δ)ε.N\ge\frac{\log\delta}{\log(1-\varepsilon)} \simeq\frac{\log(1/\delta)}{\varepsilon}.

One visit is only a support check, not a precise tail estimate. A useful diagnostic is the distribution of log-weight differences, not merely the mean phase.

For a normalized positive target density p∗p_* with p∗≪pp_*\ll p, define r=p∗/pr=p_*/p and assume Epr2<∞\mathbb E_pr^2<\infty. The familiar empirical weight diagnostic obeys

N^effN=(∑k=1Nrk)2N∑k=1Nrk2⟶(Epr)2Epr2=e−D2(p∗∥p),\frac{\widehat N_{\rm eff}}{N} =\frac{(\sum_{k=1}^N r_k)^2}{N\sum_{k=1}^N r_k^2} \longrightarrow \frac{(\mathbb E_p r)^2}{\mathbb E_p r^2} =e^{-D_2(p_*\Vert p)},

where

D2(p∗∥p)=log⁡∫dx p∗(x)2p(x)D_2(p_*\Vert p)=\log\int dx\,\frac{p_*(x)^2}{p(x)}

is the order-two Rényi divergence and Epr=1\mathbb E_pr=1. This limit measures population weight degeneracy; it is not a unique finite-sample or observable-specific effective sample size, even for positive weights. For a complex ratio, report the phase statistic and inspect positive overlap envelopes for the denominator and the observable numerator separately. Agapiou et al. 2017, §§ 2.3.2–2.3.3, Eq. (2.4) derive the population diagnostic, while Elvira, Martino, and Robert 2022, §§ 3.3–3.5 explain why the usual empirical ESS can miss observable and rare-event failures.

An observable can be supported in a tail invisible to a global average phase. Conversely, a small average phase need not make every symmetry-protected ratio impossible if numerator and denominator correlations are exploited with a justified estimator. The sign, overlap, and Silver Blaze problems are separated explicitly in Gattringer and Langfeld 2016, § 1.2, pp. 1643007-3–1643007-5. The reliability question is always: which configurations dominate this observable, and how were they sampled?

Use the dimensionless one-angle toy weight

w(θ;μ)=ecos⁡θ(1+heμ+iθ)(1+he−μ−iθ),h=0.2,w(\theta;\mu)=e^{\cos\theta} \bigl(1+he^{\mu+i\theta}\bigr) \bigl(1+he^{-\mu-i\theta}\bigr), \qquad h=0.2,

and define z=∫−ππw dθz=\int_{-\pi}^{\pi}w\,d\theta, zpq=∫−ππ∣w∣ dθz_{\rm pq}=\int_{-\pi}^{\pi}|w|\,d\theta, and r=z/zpqr=z/z_{\rm pq}. Fourier orthogonality gives the independent analytic result

z(μ)=2π[(1+h2)I0(1)+2hcosh⁡μ I1(1)].z(\mu)=2\pi\left[(1+h^2)I_0(1)+2h\cosh\mu\,I_1(1)\right].

The chemical-potential page derives this Bessel formula directly. For nn independent copies,

Zn=zn,Zpq,n=zpqn,⟨eiΦ⟩pq,n=rn.Z_n=z^n, \qquad Z_{{\rm pq},n}=z_{\rm pq}^n, \qquad \langle e^{i\Phi}\rangle_{{\rm pq},n}=r^n.

This factorized microbenchmark demonstrates the cancellation algebra exactly:

−1nlog⁡∣⟨eiΦ⟩pq,n∣=−log⁡∣r∣.-\frac{1}{n}\log\left|\langle e^{i\Phi}\rangle_{{\rm pq},n}\right| =-\log|r|.

It does not demonstrate locality, continuum scaling, or a field-theory thermodynamic limit. Direct quadrature of zpqz_{\rm pq} and the Bessel formula for zz are independent calculations. At imaginary μ=iμI\mu=i\mu_I, the bracket is

∣1+hei(θ+μI)∣2=1.04+0.4cos⁡(θ+μI)>0,\bigl|1+he^{i(\theta+\mu_I)}\bigr|^2 =1.04+0.4\cos(\theta+\mu_I)>0,

so r=1r=1. Sampling at those points is sign-free. A Taylor series about any chosen point is limited by the nearest singularity of the normalized observable or log⁡z\log z; stepwise analytic continuation may follow another path only while it remains inside the connected analytic domain.

At the prescribed real points, direct quadrature gives:

μ\muOne-copy average phase r(μ)r(\mu)
001.000000001.00000000
0.50.50.993018490.99301849
110.968235630.96823563
1.51.50.914852480.91485248

At μ=1.5\mu=1.5, the independent-copy values are r4=0.70049r^4=0.70049, r16=0.24078r^{16}=0.24078, and r64=0.0033610r^{64}=0.0033610. These numbers reproduce the severity measure without fitting a free-energy slope; a Monte Carlo implementation should recover both the one-copy ratios and their exact powers.

To expose overlap independently, fix μ=1.5\mu=1.5 and sample from the full-support von Mises proposal

pκ(θ)=eκcos⁡θ2πI0(κ),κ=8,p_\kappa(\theta)=\frac{e^{\kappa\cos\theta}}{2\pi I_0(\kappa)}, \qquad \kappa=8,

For draws θi∼pκ\theta_i\sim p_\kappa, form the explicit partition-function and average-phase estimators

z^=1N∑i=1Nw(θi;1.5)pκ(θi),z^pq=1N∑i=1N∣w(θi;1.5)∣pκ(θi),r^=z^z^pq.\widehat z=\frac1N\sum_{i=1}^N\frac{w(\theta_i;1.5)}{p_\kappa(\theta_i)}, \qquad \widehat z_{\rm pq}=\frac1N\sum_{i=1}^N\frac{|w(\theta_i;1.5)|}{p_\kappa(\theta_i)}, \qquad \widehat r=\frac{\widehat z}{\widehat z_{\rm pq}}.

Then estimate the normalized complex-weight expectation of the circular rare-support observable

Oc(θ)=1dS1(θ,π)<0.05=1∣θ∣>π−0.05.O_c(\theta)=\mathbf 1_{d_{S^1}(\theta,\pi)<0.05} =\mathbf 1_{|\theta|>\pi-0.05}.

Its estimator is

⟨Oc⟩^=∑iw(θi;1.5)Oc(θi)/pκ(θi)∑iw(θi;1.5)/pκ(θi).\widehat{\langle O_c\rangle} =\frac{\sum_i w(\theta_i;1.5)O_c(\theta_i)/p_\kappa(\theta_i)} {\sum_i w(\theta_i;1.5)/p_\kappa(\theta_i)}.

The proposal is positive everywhere, so importance sampling is asymptotically valid, but finite runs are extremely unlikely to visit the endpoints. Such a run can agree with the independently quadrature-checked zz, zpqz_{\rm pq}, and rr within quoted, conditionally estimated uncertainty while giving no information about ⟨Oc⟩\langle O_c\rangle. If proposal support were exactly zero there, the estimator would instead be invalid at every sample size. This deliberate stress test distinguishes partition-function cancellation from observable-specific overlap.

Use the following classification before choosing a method.

  • Negative or complex weight: the integrand is not a probability in the selected variables. This is an algebraic property of the representation.
  • Cancellation severity: ∣m∣|m|, its volume slope, cumulants of ϕ\phi, and the covariance of the actual ratio estimator quantify statistical loss in reweighting.
  • Overlap failure: important target regions are rare or absent in the proposal ensemble. Tail probabilities, weight distributions, sector occupancy, and forward/reverse comparisons diagnose it.
  • Ordinary numerical instability: overflow, ill-conditioned linear solves, long autocorrelation, and optimizer failure can occur with positive weights. These require numerical remedies, not sign-problem rhetoric.
  • Correctness failure: a complex stochastic or contour method converges to the wrong integral because an integration-by-parts boundary term, missing sector, Jacobian, or reconstruction tail was omitted. More samples do not remove this bias. A worst-case computational obstruction is a further, hypothesis-dependent statement, as the explicit reduction of Troyer and Wiese 2005, pp. 170201-1–170201-3 illustrates.

A tiny error bar around zero phase. Suppose all samples come from one phase mode and give a precise uˉ\bar u. If a second mode with exponentially small proposal probability contributes comparably to ZZ, the quoted variance is conditional on the wrong support. Multiple starts, sector-resolved weights, and an exact small-volume result are required.

A stable result from a modified measure. Clipping large reweighting factors, discarding configurations near determinant zeros, or replacing a complex Jacobian by its magnitude may stabilize estimates. Each operation changes the target unless its correction is included and controlled.

Exponential fit from too little range. A straight line in −log⁡∣m∣-\log|m| over two volumes does not establish an asymptotic free-energy difference. Add volumes, check finite-size terms, and compare at fixed physical temperature and parameters.

  • Report the exact phase-quenched measure, ∣m∣|m|, its uncertainty, autocorrelation, and volume dependence.
  • For the actual observable, inspect numerator–denominator covariance, reweighting-factor tails, sector occupancy, and at least one overlap stress test.
  • Compare an exact or sign-free limit and a deliberately hostile observable or parameter point.
  • Separate statistical variance, autocorrelation, overlap bias, truncation, and method-correctness uncertainty.
  • Repeat the analysis after a legitimate representation change; changed severity is expected, while the exact observable must agree.

The chapter-wide finite-density method reliability matrix turns these diagnostics into a common evidence contract for every method family.

Assume ∣m∣=e−cV|m|=e^{-cV} and unit-modulus phase samples from a stationary Markov chain. Define the real normalized autocorrelation

ρ(t)=Re⁡E[(u0−m)∗(ut−m)]1−∣m∣2,τint=12+∑t≥1ρ(t),\rho(t)=\frac{\operatorname{Re}\mathbb E[(u_0-m)^*(u_t-m)]}{1-|m|^2}, \qquad \tau_{\rm int}=\frac12+\sum_{t\ge1}\rho(t),

and assume the series is summable. How must NN scale to keep the relative root-mean-square error of uˉ\bar u below ϵ\epsilon?

Solution

The variance of the correlated mean is asymptotically 2τint2\tau_{\rm int} times the independent-sample variance. Hence

N≳2τint(1−e−2cV)ϵ2e−2cV=2τintϵ−2(e2cV−1).N\gtrsim\frac{2\tau_{\rm int}(1-e^{-2cV})} {\epsilon^2e^{-2cV}} =2\tau_{\rm int}\epsilon^{-2}(e^{2cV}-1).

For large VV, N∼2τintϵ−2e2cVN\sim2\tau_{\rm int}\epsilon^{-2}e^{2cV}. If τint\tau_{\rm int} itself grows with volume, that additional scaling must be reported rather than hidden in an “effective sample” count.

A proposal assigns probability 10−510^{-5} to the region that supplies half of an observable’s numerator. How many independent samples are needed for at least 95% probability of visiting it once?

Solution

Require 1−(1−10−5)N≥0.951-(1-10^{-5})^N\ge0.95. Thus

Nmin⁡=⌈log⁡0.05log⁡(1−10−5)⌉=299,572.N_{\min}=\left\lceil\frac{\log0.05}{\log(1-10^{-5})}\right\rceil =299{,}572.

One visit is not enough for a precise contribution; this is only a necessary support test.

Classify each case as cancellation, overlap, ordinary numerical instability, or correctness failure: (a) ∣m∣|m| falls exponentially but the exact identity and support tests pass; (b) a positive-weight proposal never enters the sector that dominates the numerator; (c) a linear solver overflows in a sign-free ensemble; (d) a stationary complex-Langevin process has a nonvanishing integration-by-parts boundary term.

Solution

(a) is cancellation: the direct ratio has exponentially poor signal-to-noise. (b) is overlap: the relevant target support is absent from the observed sample. (c) is ordinary numerical instability, because the probability measure is still nonnegative. (d) is a correctness failure: stationarity does not identify the desired integral when the boundary term survives, and more samples do not remove that bias.

After working this page, you should be able to:

  • Derive the relative variance of the average-phase estimator and extract its measured volume-cost exponent with autocorrelation included.
  • Given sampling evidence for an observable, classify the limiting failure as cancellation, overlap, ordinary numerical instability, or correctness bias and name a discriminating test.

Basis dependence and computational complexity asks which of these diagnostics survive a representation change. The method pages then replace the generic warning with explicit correctness contracts.

  • Agapiou, Sergios, Omiros Papaspiliopoulos, Daniel Sanz-Alonso, and Andrew M. Stuart. “Importance Sampling: Intrinsic Dimension and Computational Cost.” Statistical Science 32 (2017): 405–431. doi:10.1214/17-STS611.
  • Alexandru, Andrei, Gökçe Başar, Paulo F. Bedaque, and Neill C. Warrington. “Complex Paths around the Sign Problem.” Reviews of Modern Physics 94 (2022): 015006. doi:10.1103/RevModPhys.94.015006.
  • Elvira, Víctor, Luca Martino, and Christian P. Robert. “Rethinking the Effective Sample Size.” International Statistical Review 90 (2022): 525–550. doi:10.1111/insr.12500.
  • Gattringer, Christof, and Kurt Langfeld. “Approaches to the Sign Problem in Lattice Field Theory.” International Journal of Modern Physics A 31 (2016): 1643007. doi:10.1142/S0217751X16430077. Open PDF: arXiv:1603.09517.
  • Troyer, Matthias, and Uwe-Jens Wiese. “Computational Complexity and Fundamental Limitations to Fermionic Quantum Monte Carlo Simulations.” Physical Review Letters 94 (2005): 170201. doi:10.1103/PhysRevLett.94.170201.

Original QFT.org content:CC BY 4.0, unless an item supplies different terms. Third-party material retains its own terms.