Skip to content

Uncertainty, disagreement, and negative results

Two results disagree only after they refer to the same object under compatible conditions. Before interpreting a tension, align the observable, kinematics, conventions, approximation order, inputs, and inferential question. Then keep shared uncertainties and genuinely independent checks visible.

This guide ends with a comparison another researcher can inspect: what agrees, what differs, how uncertainties are related, which conclusion is supported, and which new result would change the comparison.

Helpful background. Bring the claim sheet from Scope, conventions, and status and the protocol from Reproduce and validate a result. If covariance, effective sample size, or conditional inference is unfamiliar, use the statistics preparation check first.

Match the objects before comparing the numbers

Section titled “Match the objects before comparing the numbers”

Write one row for each result:

FieldResult AResult B
Quantity and units
External state, geometry, or dataset
Kinematic or parameter domain
Conventions and scheme
Approximation or truncation order
Shared inputs
Method-specific inputs
Central result
Uncertainty components and correlations
Strongest supported conclusion

Do not translate the central values without translating their uncertainty and scope. A coefficient quoted at two renormalization scales, a Euclidean correlator and a real-time response, or a partonic and a fiducial cross section are different objects until the required evolution or forward map is supplied.

Three lists make the comparison legible:

  • Common ground: identities, inputs, limits, or qualitative behavior both calculations share.
  • True differences: assumptions, data, approximation order, algorithm, prior, observable, or physical interpretation that remains different after translation.
  • Apparent differences: notation, units, basis, parameterization, or scale choices that disappear under a checked transformation.

Only a true difference about a common target is a scientific disagreement.

Running case: two truncations are not a contradiction

Section titled “Running case: two truncations are not a contradiction”

Continue the heavy-exchange example from the reproduction step. Normalize one exchange channel as

F(x)=11x,x=q2M2,x<1.F(x)=\frac{1}{1-x}, \qquad x=\frac{q^2}{M^2}, \qquad |x|<1.

The local expansion through order NN is

FN(x)=n=0Nxn,RN(x)=F(x)FN(x)=xN+11x.F_N(x)=\sum_{n=0}^{N}x^n, \qquad R_N(x)=F(x)-F_N(x)=\frac{x^{N+1}}{1-x}.

Relative to the exact result,

RN(x)F(x)=xN+1.\frac{\lvert R_N(x)\rvert}{\lvert F(x)\rvert} =\lvert x\rvert^{N+1}.

Suppose one group uses leading order, F0=1F_0=1, while another uses next order, F1=1+xF_1=1+x. At x=0.10x=0.10, their values are 1.001.00 and 1.101.10, whereas the exact value is 1.1111.111\ldots. Their absolute difference is 0.100.10, or ten percent relative to F0F_0, but they do not make contradictory claims: their relative truncation errors are ten percent and one percent, respectively.

The correct comparison is

ApproximationValue at x=0.10x=0.10Relative truncation errorSupported statement
F0F_01.001.0010%10\%Leading contact term captures the scale and sign, not percent precision.
F1F_11.101.101%1\%The first momentum correction reaches percent accuracy at this point.
FF1.1111.111\ldotsReference value for this one-channel toy calculation.

If both groups instead claim to have computed F1F_1 at the same xx and obtain different values after convention translation, there is a genuine calculation-level disagreement. The one-channel series illustrates controlled truncation; it does not establish the full decoupling theorem, loop matching, crossing relations, or the behavior of nondecoupling couplings.

Do not begin with a single error bar. Record each contribution at the stage where it arises:

  • definition: ambiguity in the observable, event class, operator, or phase criterion;
  • input: calibration, external parameters, boundary data, ensemble, or dataset selection;
  • theory: omitted perturbative orders, EFT powers, nonperturbative matrix elements, or model discrepancy;
  • numerical: discretization, finite volume, sampling, solver tolerance, or unstable inverse problems;
  • translation: matching, basis conversion, scale evolution, units, or detector response; and
  • inference: likelihood, priors, nuisance parameters, regularization, selection, and model comparison.

For every contribution, state how it was estimated, whether it is signed or bounded, what it correlates with, and which check supports it. The JCGM 100:2008, §§4–5 gives a general framework for propagating input uncertainties and covariance for a well-defined quantity. An EFT truncation estimate or model discrepancy still requires subject-specific reasoning; it should not be relabeled as sampling variation.

In the running case, the first omitted geometric-series term gives a useful truncation scale because the remainder is known exactly. In a real EFT, unknown coefficients, logarithms, thresholds, and multiple expansion parameters can weaken that estimate. Order variation is evidence only when the power counting and observed convergence support it.

Shared inputs change the uncertainty of a difference

Section titled “Shared inputs change the uncertainty of a difference”

Let estimates AA and BB have standard uncertainties σA,σB\sigma_A,\sigma_B and correlation ρ\rho. The variance of their difference is

Var(AB)=σA2+σB22ρσAσB.\operatorname{Var}(A-B) =\sigma_A^2+\sigma_B^2-2\rho\sigma_A\sigma_B.

Positive correlation can make a difference more precise because a common shift cancels. It can also make agreement less independent: two pipelines using the same data, calibration, perturbative coefficients, code library, or fit model may share the dominant failure.

Record dependencies explicitly. Agreement between correlated implementations is a valuable verification, but it is not the same as confirmation by a method with different failure modes. The distinction between computational reproducibility and broader scientific replication is developed in the National Academies 2019 report, chs. 3–4.

After matching objects and propagating shared uncertainty, classify the relation conservatively:

  • compatible: both claims can hold in their recorded scopes;
  • tension: their preferred values or regions pull apart, but the evidence does not justify contradiction;
  • contradiction: they make incompatible statements about the same target under the same assumptions;
  • underdetermined: available evidence does not discriminate among the relevant explanations;
  • non-comparable: the target objects or scopes remain different; or
  • correction: one calculation contains an identified error and the revised result replaces it for the present comparison.

A quoted “number of sigma” does not choose the category on its own. The sampling model, nuisance treatment, selection procedure, look-elsewhere effects, parameter boundaries, and asymptotic regime are part of its meaning. Likelihood-based asymptotic formulae and their conditions are presented in Cowan et al. 2011, §§2–3.

Read a negative result within its sensitivity

Section titled “Read a negative result within its sensitivity”

A negative result says what a test did not find in the region where it had power. Record:

  1. the target signal and operational signature;
  2. the parameter or function space tested;
  3. expected outcomes under the relevant alternatives;
  4. observed result and uncertainty;
  5. inference procedure and coverage or calibration checks;
  6. blind spots and failed assumptions; and
  7. the broader question left open.

For the heavy-scalar example, define the leading local operator by

ΔL=C8ϕ4,\Delta\mathcal L=\frac{C}{8}\phi^4,

so that its coefficient in this normalization is

C=g2M2.C=\frac{g^2}{M^2}.

An interval consistent with C=0C=0 does not separately determine gg and MM, exclude heavy fields with other couplings, or test energies near the heavy pole. It constrains the specified operator combination within the measurement, matching, and EFT domain. A failed numerical solver is even narrower: it is a negative result about that method and setup, not evidence that the physical solution does not exist.

A useful comparison points toward a result that could change it:

common target:
decision-relevant difference:
candidate explanations:
observable or theorem that separates them:
expected outcome under each explanation:
required precision or domain:
method with a different dominant failure:
shared dependencies that remain:
blind spots:
proceed / revise / stop rule:

In the running example, evaluating at several xx values separates a wrong normalization from a missing higher-order term. A normalization error remains roughly constant in xx; an omitted first correction produces a residual proportional to xx at leading order. Near x=1|x|=1, neither pattern validates the local expansion because the scale hierarchy itself is failing.

One paper writes the leading relative EFT error as O(q2/M2)O(q^2/M^2). Another writes O((E/M)2)O((E/M)^2) with q2=E2q^2=E^2. Do the statements disagree?

Solution

No. Substituting q2=E2q^2=E^2 gives

O ⁣(q2M2)=O ⁣(E2M2)=O ⁣[(EM)2].O\!\left(\frac{q^2}{M^2}\right) =O\!\left(\frac{E^2}{M^2}\right) =O\!\left[\left(\frac{E}{M}\right)^2\right].

They use different variables for the same power suppression. A disagreement would remain only if their q2q^2 and E2E^2 referred to different kinematics or if one statement included a coefficient or logarithm the other omitted.

Two estimates are A=1.00±0.05A=1.00\pm0.05 and B=1.10±0.06B=1.10\pm0.06. Compute AB/σAB|A-B|/\sigma_{A-B} first for ρ=0\rho=0 and then for ρ=0.8\rho=0.8.

Solution

For independent uncertainties,

σAB=0.052+0.0620.0781,ABσAB1.28.\sigma_{A-B} =\sqrt{0.05^2+0.06^2} \simeq0.0781, \qquad \frac{|A-B|}{\sigma_{A-B}}\simeq1.28.

For ρ=0.8\rho=0.8,

σAB=0.052+0.0622(0.8)(0.05)(0.06)0.0361,\sigma_{A-B} =\sqrt{0.05^2+0.06^2-2(0.8)(0.05)(0.06)} \simeq0.0361,

so the standardized difference is about 2.772.77. The positive shared component cancels in ABA-B. This calculation does not by itself establish a scientific tension; the Gaussian model, correlation estimate, selection procedure, and common target still need checking.

3. Interpret a null contact-interaction result

Section titled “3. Interpret a null contact-interaction result”

A fit finds a contact coefficient compatible with zero within its stated interval. Which conclusions are supported?

Solution

The fit constrains that coefficient—or a stated combination of coefficients—under the chosen data, likelihood, operator basis, matching, truncation, and kinematic cuts. If the interval has validated coverage, one may report the corresponding allowed region.

It does not prove that no heavy particle exists, determine gg and MM separately when only g2/M2g^2/M^2 enters, exclude interactions outside the chosen basis, or establish validity near q2M2q^2\sim M^2. Those are the nearest stronger nonclaims.

  • Cowan, Glen, Kyle Cranmer, Eilam Gross, and Ofer Vitells. 2011. “Asymptotic Formulae for Likelihood-Based Tests of New Physics.” European Physical Journal C 71: 1554. DOI and open article.
  • Joint Committee for Guides in Metrology. 2008. Evaluation of Measurement Data—Guide to the Expression of Uncertainty in Measurement. JCGM 100:2008. DOI. Official PDF.
  • National Academies of Sciences, Engineering, and Medicine. 2019. Reproducibility and Replicability in Science. Washington, DC: National Academies Press. DOI and open book.

Carry the matched comparison, shared dependencies, and discriminating test into Scope a first research project.