Correlated Standard Model Fits and Consistency Tests
A correlated Standard Model fit is a specified joint probability model for exact input releases, shared parameters, shared nuisances, and any dataset overlap. It is not a sum of independently quoted pulls. Reliable consistency statements require a frozen validity domain, a calibrated test statistic, and versioned data and theory objects. This page develops those durable operations and exact synthetic checks, without giving a current global-fit result.
Required background. Collider measurements, fiducial predictions, and likelihood provenance supplies release identity, overlap, and nuisance semantics. Electroweak precision observables supplies input-scheme and covariance construction for a correlated sector.
Helpful background. Higgs precision and coupling inference supplies production–decay, width, response, and EFT-validity assumptions for Higgs inputs.
Evergreen prediction methods and changing fit inputs have different lifecycles. The map below shows how an observable contract connects them while corrections, supersession, and dated interpretation remain outside the canonical derivation.
A new release or correction can supersede a Standard Model fit without changing the theory derivation. Every numerical conclusion therefore names the frozen evidence and likelihood versions; the diagram is schematic.
The joint likelihood is the fitted object
Section titled “The joint likelihood is the fitted object”Let denote the data from input block , the parameters of interest, block-specific nuisances, and shared nuisances. If the primary observations are conditionally independent after common causes are represented, the joint likelihood can be written
This product is justified by conditional independence, not by typography. If two inputs share events, an unfolded ancestor, control-region counts, or an auxiliary calibration, then one must build the relevant joint model, remove the overlap according to a predeclared rule, or omit one input. Duplicating a common auxiliary factor artificially tightens the nuisance and the parameters it controls.
The data block is meaningful only with its theory definition. A prediction may depend on pole and input schemes, perturbative scales, PDFs, hadronic matrix elements, matching coefficients, or EFT truncation. Sources representing the same missing-order or parametric direction must be correlated across blocks; sources with different physical causes should not be correlated merely because their quoted sizes are similar.
Generalized least squares as an exact check
Section titled “Generalized least squares as an exact check”For a linear Gaussian model
with fixed positive-definite , minimizing
gives
provided has full column rank in the metric defined by . This is a valuable analytic benchmark for a likelihood implementation.
Consider the exact two-bin fixture
Its determinant is , and
The numerator and denominator of the estimator are
so
The estimate is not the arithmetic mean: the correlated, unequal covariance changes the metric. Adding the two “pull squares” does not reproduce because it drops the off-diagonal term. A Cholesky factor supplies whitened residuals whose Euclidean norm does reproduce the quadratic form.
Profiling and marginalization are different operations
Section titled “Profiling and marginalization are different operations”Profiling maximizes the likelihood over nuisances at each parameter value:
Marginalization integrates them with a declared measure or prior density:
They answer different inferential questions and generally have different shapes. An exact Gaussian fixture shows a special coincidence. Let
Completing the square gives
Therefore
With and a flat integration measure in ,
The profiled likelihood has the same dependence in this constant-width Gaussian example. That equality is not general: parameter-dependent curvature, boundaries, non-Gaussian auxiliary data, or a different integration measure change the marginalization factor. A published result must identify which operation and nuisance measure were used. Profile-likelihood asymptotics also require regularity and adequate sample size; small samples and boundaries call for simulation-based calibration Cowan et al. 2011, §§2–3, pp. 4–12.
Goodness of fit, pulls, and trials
Section titled “Goodness of fit, pulls, and trials”In the fixed-covariance linear Gaussian model, if the mean model is correct and has rank , then follows a chi-square distribution with degrees of freedom. That statement can fail when:
- the covariance depends on fitted parameters or was estimated with material uncertainty;
- likelihood terms are non-Gaussian or counts are sparse;
- parameters lie on boundaries or are not identifiable under the null;
- the model is nonlinear over the supported region;
- data-dependent masks or regularization alter the statistic; or
- nuisance constraints are nonregular or duplicated.
In those cases, define the statistic and calibrate its sampling distribution with suitable pseudoexperiments or another justified method. The asymptotic formulas in Cowan et al. 2011, §§3–4, pp. 9–17 are approximations with explicit hypotheses, not labels that automatically turn a likelihood ratio into a significance.
A residual divided by its marginal standard deviation is not an independent test when entries are correlated. Useful diagnostics include whitened residuals, conditional residuals, nuisance pulls relative to their auxiliary constraints, and impacts obtained under a stated refit. Their collection remains correlated and should not be scanned as independent local significances.
Leave-one-block-out fits can reveal leverage or an inconsistent interface: refit after removing a predeclared block and compare predictions in that block. They do not by themselves assign a discovery probability, and repeatedly choosing the most discrepant omitted block introduces a trials problem.
For a search over masses, channels, operators, bins, or alternative masks, distinguish a local tail probability at a fixed hypothesis from a global probability for the predeclared search family. The family, scan resolution, and correlation structure must be fixed before calibration. A posterior selection of “interesting” trials cannot be repaired by quoting the local tail alone.
Freeze the fit contract
Section titled “Freeze the fit contract”| Component | Freeze before fitting | Check after fitting |
|---|---|---|
| Input identity | Dataset/table/workspace version, DOI or stable identifier, checksum and corrections | Every loaded byte and bin order matches the record |
| Observable and theory | Pole/fiducial definition, input and renormalization schemes, prediction version and accuracy | Benchmarks, dimensions, limits, and cross-scheme translation |
| Correlations | Covariance components or shared nuisance graph; auxiliary data included once | Symmetry/positivity, nuisance response, no duplicate constraint |
| Overlap | Event, control-region, unfolded-ancestor, and theory-input overlap rule | Product factorization is justified for retained blocks |
| Validity | Kinematic/EFT masks and parameter domains | Boundary hits and truncation stress tests are reported |
| Inference | Profiling or marginalization, statistic, degrees of freedom or calibration, trial family | Optimizer coverage, toy calibration, independent implementation |
| Lifecycle | Correction and supersession policy; output version | A corrected input creates a new result rather than silently replacing one |
Publishing a statistical model rather than only its final contour preserves more of this structure and makes combinations and alternative parameterizations testable Cranmer et al. 2022, §§2–4, pp. 4–18.
Independent checks and failure modes
Section titled “Independent checks and failure modes”Positive definiteness. For an ordinary invertible Gaussian covariance, test symmetry, eigenvalues, and Cholesky factorization. A singular covariance may be legitimate for normalized data, but then the constrained subspace and generalized-inverse prescription must be explicit.
Duplicate input. Set up a table of event samples, controls, auxiliary measurements, and theory sources for every block. A repeated item needs one joint representation or a documented removal rule.
Analytic benchmark. Reproduce and for the two-bin fixture. This tests matrix ordering, inversion, residual sign, and optimizer output independently.
Profile benchmark. Reproduce and the profiled curve above. Numerical profiling should agree over a grid, including points far enough from the optimum to exercise the implementation.
Coverage fixture. If and the interval is predeclared as , its exact coverage is
A simulation or quadrature check should reproduce this value within its numerical tolerance. Choosing the interval or mask after observing invalidates that coverage statement.
Stability. Repeat the fit under predeclared leave-one-block-out, scheme, correlation, and validity variations. Explain the physical source of motion; do not convert the largest variation into an independent Gaussian error without a model.
Common pitfalls
Section titled “Common pitfalls”Adding uncorrelated pulls. Off-diagonal covariance terms can raise or lower the joint discrepancy. Use the inverse covariance or full likelihood.
Profiling a nuisance twice. A covariance distilled from a profiled likelihood and the original auxiliary constraint are not independent inputs. Choose one faithful representation.
Treating profiling and marginalization as synonyms. They coincide in shape only in special Gaussian cases. State the operation, parameterization, and integration measure or auxiliary model.
Changing the data version or mask mid-fit. The resulting likelihood is not the one whose sampling properties were defined. Freeze versions and masks, then issue a new fit when an input changes.
Quoting a significance without a reference distribution. State the statistic, null, degrees of freedom or calibration, boundaries, and local/global trial definition.
Informal self-check
Section titled “Informal self-check”Derive the generalized least-squares estimate and minimum chi-square for the two-bin fixture above without numerical minimization.
Solution
Using the displayed inverse,
Thus and , giving . The residual is ; substitution into gives .
Handoffs
Section titled “Handoffs”- Return to collider measurements and likelihood provenance if any input lacks exact identity, nuisance semantics, or an overlap rule.
- Return to electroweak precision observables for input-scheme and sector-covariance construction.
- Return to Higgs precision and coupling inference for width, response, and EFT assumptions in Higgs blocks.
- Reproduce the analytic covariance, nuisance, and coverage fixtures before applying the workflow to released inputs.
References
Section titled “References”- Cowan, Glen, Kyle Cranmer, Eilam Gross, and Ofer Vitells. “Asymptotic Formulae for Likelihood-Based Tests of New Physics.” European Physical Journal C 71 (2011) 1554. DOI · Open PDF
- Cranmer, Kyle, et al. “Publishing Statistical Models: Getting the Most out of Particle Physics Experiments.” SciPost Physics 12 (2022) 037. DOI · Open PDF