Skip to content

Multistage Bayesian Global Inference and Evidence

A heavy-ion global analysis calibrates a computational model that maps fluctuating nuclear initial states through pre-equilibrium, viscous hydrodynamics, particlization, hadronic transport, and optionally rare probes to measured observables. Bayesian machinery makes the conditional uncertainty explicit; it cannot rescue an incomplete covariance, an invalid stage, a biased emulator, or a non-identifiable parameterization. The primary product is therefore a versioned posterior and its model/data provenance, not a best-fit curve.

Required background. Transport inverse problems and error budgets supplies identifiability and covariance, initial conditions supplies the first stage, and bulk evolution supplies the required downstream chain. Helpful background. Charge fluctuations, jets, electromagnetic probes, open heavy flavor, and quarkonium add typed optional probes.

Evidence status on this page was checked through 10 August 2026.

Let θ\theta denote model parameters, MM the complete model/version, and yy the selected measurements. Bayes’ theorem gives

p(θy,M)=p(yθ,M)p(θM)dθp(yθ,M)p(θM).p(\theta\mid y,M) =\frac{p(y\mid\theta,M)p(\theta\mid M)} {\int d\theta\,p(y\mid\theta,M)p(\theta\mid M)}.

For a Gaussian residual, a useful likelihood is

2lnp(yθ,M)=r(θ)TΣtot1r(θ)+lndetΣtot+constant,-2\ln p(y\mid\theta,M) =r(\theta)^T\Sigma_{\rm tot}^{-1}r(\theta) +\ln\det\Sigma_{\rm tot}+\text{constant},

where r=yexpyM(θ)δMr=y_{\rm exp}-y_M(\theta)-\delta_M and

Σtot=Σexp+ΣMC+Σemu+Σdisc.\Sigma_{\rm tot} =\Sigma_{\rm exp} +\Sigma_{\rm MC} +\Sigma_{\rm emu} +\Sigma_{\rm disc}.

The terms represent experimental covariance, finite-event simulation noise, emulator predictive covariance, and model discrepancy. They cannot generally be replaced by diagonal error bars. Correlated normalization systematics are often better represented by nuisance parameters than by adding independent variances.

The posterior is conditional on MM. A narrow posterior can result from an overly rigid parameterization or omitted discrepancy; precision is not automatically accuracy.

Full event-by-event simulations are expensive, so an analysis evaluates a space-filling design {θd}\{\theta_d\}, reduces correlated outputs if useful, and trains an emulator y^(θ)\widehat y(\theta). A Gaussian-process or neural emulator must return predictive uncertainty, not only a mean.

Validation uses held-out full-model points. Standardized residuals,

zdi=yM,i(θd)y^i(θd)Varemu,i(θd)+VarMC,i(θd),z_{di} =\frac{y_{M,i}(\theta_d)-\widehat y_i(\theta_d)} {\sqrt{\operatorname{Var}_{\rm emu,i}(\theta_d) +\operatorname{Var}_{\rm MC,i}(\theta_d)}},

should have the declared coverage and reveal no parameter-dependent structure. Good interpolation error averaged over outputs can hide a bias in the few combinations that control a transport coefficient, so validation should be performed in the likelihood metric.

The design must cover the posterior support. If sampling accumulates at a prior boundary or outside the convex region validated by full simulations, the remedy is additional design/model work, not trusting extrapolation.

Closure, posterior prediction, and identifiability

Section titled “Closure, posterior prediction, and identifiability”

A closure test generates pseudodata at a hidden θ\theta_\star, passes it through the same covariance and emulator pipeline, and asks whether calibrated credible regions cover θ\theta_\star at the expected frequency. Closure on noiseless emulator output is too weak; the test should use independent full-model runs and realistic noise.

Posterior predictive checks draw

θ(s)p(θy,M),y~(s)p(y~θ(s),M),\theta^{(s)}\sim p(\theta\mid y,M),\qquad \widetilde y^{(s)}\sim p(\widetilde y\mid\theta^{(s)},M),

including nuisance and discrepancy terms. They test whether the fitted model reproduces held-out distributions and correlations, not merely means used in calibration.

Local identifiability is encoded by the sensitivity matrix Jiα=yi/θαJ_{i\alpha}=\partial y_i/\partial\theta_\alpha. Small singular values of the whitened matrix

Σtot1/2J\Sigma_{\rm tot}^{-1/2}J

identify parameter combinations that data cannot distinguish. Reporting marginal intervals without the correlated directions can make a degeneracy look like several independent measurements.

A joint calibration must distinguish parameters intended to be common properties of QCD matter from system-specific latent variables. A hierarchical model may share a transport function across collision systems while allowing nuclear geometry, longitudinal deposition, normalization, and detector nuisance parameters to vary. Simply multiplying likelihoods is valid only after cross-system experimental and theory correlations have been represented.

Beam-energy transfer is especially nontrivial. Lower energies introduce stronger baryon stopping, three-dimensional longitudinal structure, nonzero conserved-charge densities, diffusion, and a finite-density EOS; a boost-invariant zero-density stage graph cannot be carried over unchanged. A shared posterior is meaningful only on the intersection of the stages’ validity domains or after the energy-dependent extensions are parameterized and tested.

Small systems provide an even sharper boundary. Short lifetimes and gradients can enlarge nonhydrodynamic corrections, while initial-state correlations can mimic part of the final anisotropy. Including pppp or pApA data can improve discrimination, but it does not by itself prove that the same hydrodynamic truncation is valid there. The analysis must pass switching, constitutive-residual, and structural-alternative tests in each system before a common transport interpretation is claimed.

Heavy-ion global-inference provenance table

Section titled “Heavy-ion global-inference provenance table”

This canonical table defines the minimum machine-readable record for bulk and typed probe analyses. A record fails schema validation when a required item is absent; “not evaluated” must be explicit.

Record blockRequired contentReject or limit the claim when
Analysis identityImmutable release, code/container versions, random seeds, checksums, evidence cutoff, supersession statusResults cannot be reproduced or the evidence date is unknown
Stage graphRequired bulk chain: ordered initial-state, pre-equilibrium, hydrodynamic, EOS, particlization, afterburner, detector-analysis versions and switch maps; optional probe branches explicitly markedA stage or interface is unnamed, conservation is untested, or an optional probe is mistaken for a required bulk stage
Parameter schemaNames, units, transforms, domains, functional bases, shared/fixed parametersA reported parameter changes meaning across stages or systems
PriorJoint density including correlations, bounds, hyperpriors, and rationalePosterior support is prior-boundary dominated without sensitivity analysis
Observable schemaDefinition, units, cuts, centrality, species/feed-down, estimator, theory analysis codeTheory and experiment implement different observables
Dataset identityCollaboration, collision system/energy, publication/data release, table/bin locator, revisionPoints cannot be traced to a primary dataset
Experimental covarianceStatistical and correlated systematic covariance or nuisance construction, including cross-observable/system informationDiagonal approximation materially changes the posterior
Design and simulator noiseDesign algorithm/points, event counts, convergence, ΣMC\Sigma_{\rm MC}, failed-run policyDesign misses posterior support or simulation noise is ignored
EmulatorType, preprocessing, hyperparameters, predictive covariance, held-out points, coverage in likelihood metricBias or undercoverage is unresolved
Model discrepancyFunctional form/kernel, domain knowledge, hyperpriors, identifiability with physics parametersOmitted discrepancy drives implausibly narrow or stage-dependent constraints
ClosureBlinded generator point/model, independent simulation, noise, coverage criterion and resultThe pipeline cannot recover known parameters
Posterior productsWeighted samples, log likelihood/prior, nuisance draws, convergence diagnostics, evidence when claimedOnly a best fit or corner plot is released
Posterior predictionCalibrated and held-out observables, full uncertainty, residual diagnosticsKey data are systematically missed
Alternative modelsInitial states, δf\delta f, afterburners, kernels or model averaging, with common data/covarianceClaim is sensitive to one untested structural choice
Sensitivity and claim statusWhitened sensitivity/identifiable combinations; label as fit, constraint, preference, exclusion, or discoveryLanguage exceeds the tested discriminator
Typed probe discriminatorOne of “bulk,” “jet,” “electromagnetic,” “open-heavy-flavor,” “quarkonium,” or a separately versioned extensionA probe record uses generic fields that erase its kernel
Jet extensionHard pppp/nuclear baseline, shower and medium kernels, recoil/response, reconstruction, covariance, validity domain, evidence cutoffA transport claim lacks a factorized baseline, response model, or common observable definition
Electromagnetic extensionCurrent/emission-rate version, equilibrium/viscous domain, prompt/decay/pre-equilibrium baselines, acceptance, covariance, validity domain, evidence cutoffSource attribution or temperature claim lacks a competing-source baseline
Open-heavy-flavor extensionProduction/CNM baseline, Boltzmann/Langevin kernel and coefficients, radiative terms, fragmentation/coalescence, feed-down, covariance, validity domain, evidence cutoffA diffusion claim lacks kernel definition or hadronization alternatives
Quarkonium extensionState/feed-down network, EFT hierarchy, open-system/rate kernel, formation, dissociation/regeneration, CNM baseline, covariance, validity domain, evidence cutoffA potential/rate claim mixes incompatible hierarchies or omits regeneration

The typed discriminator prevents physically different inverse problems from being represented as a generic “probe likelihood.” Each extension carries its own validity domain and the same evidence cutoff as the parent analysis.

Multisystem JETSCAPE calibration showed that soft observables across RHIC and LHC can constrain parameterized shear and bulk viscosities while exposing sensitivity to particlization models Everett et al. 2021. An IP-Glasma-based analysis added a distinct initial-state family and explored transfer learning and model averaging Heffernan et al. 2024. A 2025 jet analysis extended Bayesian calibration to inclusive hadron and jet suppression Ehlers et al. 2025.

Recent methodology explicitly models theory discrepancy; controlled examples show that doing so can broaden or reconcile physics-parameter posteriors that otherwise appear inconsistent Jaiswal et al. 2025. This is not a license for an arbitrarily flexible discrepancy, which could absorb all parameter sensitivity. Its kernel and hyperpriors need physical justification and closure.

As of the cutoff, global analyses provide quantitative, reproducible, model-conditional constraints and comparisons. They do not deliver a unique reconstruction of the QGP’s microscopic dynamics, a model-independent viscosity curve, or Bayesian proof that one complete collision paradigm is true.

  1. Freeze the stage graph, datasets, observable code, covariance, and priors.
  2. Run design simulations and numerical convergence checks.
  3. Validate emulator coverage on held-out full simulations.
  4. Pass blinded closure before looking at data posteriors.
  5. Sample with convergence diagnostics and retain weighted posterior products.
  6. Run posterior prediction, leave-one-system/observable-out tests, and prior/discrepancy sensitivity.
  7. Compare structural alternatives with the same dataset and covariance.
  8. Release the table record, data transformations, samples, and strongest supported claim.

The multistage graph is the forward map whose uncertainty model must accompany every global posterior.

A heavy-ion parameter vector passes through initial conditions, pre-equilibrium, viscous hydrodynamics, particlization and hadronic evolution, and detector-level analysis; emulation, correlated likelihoods, priors, discrepancy, and posterior predictive checks determine what the data identify.

Each solid edge is a physical or statistical transformation that can add parameters, numerical error, and structural alternatives. The likelihood must compare detector-matched observables with the full experimental covariance and a validated emulator, while posterior predictive checks test withheld or replicated data. The dashed branch marks the key limitation: a narrow posterior is a conditional constraint unless alternative model structures are also resolved. The diagram is schematic and not to scale.

The text equivalent is to version every stage, validate the emulator and closure tests, retain correlated systematics and model discrepancy, state priors, test structural alternatives on the same data, and report the strongest conclusion that remains stable under those changes.

Global inference can combine probes only by retaining the six distinct physical branches shown here.

One QCD temperature, flow, and charge history feeds six separate branches: bulk stress response, jet energy loss and broadening, photon and dilepton emission, open-heavy-flavor diffusion and hadronization, quarkonium dissociation and regeneration, and charge susceptibilities with critical response.

The common medium history creates useful cross-probe constraints and shared uncertainties, but each branch has a different operator, baseline, evolution kernel, acceptance, and model discrepancy. Quarkonium and charge cumulants are explicitly separate, as are quarkonium and open heavy flavor. A valid joint likelihood therefore preserves typed branch models and their block covariance rather than collapsing them into a generic probe term. The diagram is schematic and not to scale.

In text, share only physically common parameters and nuisance sources, retain branch-specific validity domains and evidence cutoffs, and test whether posterior constraints survive alternative kernels and medium histories. Apparent agreement across probes is not independent evidence if the same model error drives every branch.

1. Correlated normalization. Two data points share a 5%5\% normalization uncertainty. Why is adding 5%5\% independently to each diagonal element wrong?

Solution

The uncertainty moves both points coherently, so it contributes an outer-product covariance Σijnorm=σN2yiyj\Sigma_{ij}^{\rm norm}=\sigma_N^2y_iy_j, or an equivalent common nuisance parameter. A diagonal treatment permits opposite fluctuations and falsely increases shape uncertainty while decreasing correlation information.

2. Prior-bound posterior. A viscosity parameter’s posterior piles up at its lower prior bound. What can be concluded?

Solution

Only that, within the model and prior, likelihood support extends toward or beyond that boundary. One should expand a physically admissible prior, inspect model validity, emulator coverage, and degeneracies, and repeat sensitivity tests. Quoting the boundary as a measured value is unjustified.

Continue to the chapter overview to select the equilibrium, collision-stage, or probe-specific branch to which this inference record applies.

  • Ehlers, Raymond, et al. (JETSCAPE Collaboration). “Bayesian Inference Analysis of Jet Quenching Using Inclusive Jet and Hadron Suppression Measurements.” Physical Review C 111, no. 5 (2025): 054913. DOI.
  • Everett, D., et al. (JETSCAPE Collaboration). “Multisystem Bayesian Constraints on the Transport Coefficients of QCD Matter.” Physical Review C 103, no. 5 (2021): 054904. DOI.
  • Heffernan, Matthew, Charles Gale, Sangyong Jeon, and Jean-François Paquet. “Bayesian Quantification of Strongly Interacting Matter with Color Glass Condensate Initial Conditions.” Physical Review C 109, no. 6 (2024): 065207. DOI.
  • Jaiswal, Sunil, Chun Shen, Richard J. Furnstahl, Ulrich Heinz, and Matthew T. Pratola. “Bayesian Model–Data Comparison Incorporating Theoretical Uncertainties.” Physics Letters B 870 (2025): 139946. DOI.
  • Paquet, Jean-François. “Applications of Emulation and Bayesian Methods in Heavy-Ion Physics.” Journal of Physics G: Nuclear and Particle Physics 51, no. 10 (2024): 103001. DOI.