Skip to content

Evidence Triangulation and Reproducibility

Evidence triangulation combines analytic constraints, numerical calculations, and experiments that fail in different ways. Its force comes from conditional independence and successful cross-prediction, not from counting plots. Two results that use the same dataset, Hamiltonian, calibration, prior, or preprocessing may be valuable consistency checks while contributing far less independent evidence than their number suggests.

Required background. Model selection and parameter inference supplies the joint likelihood. QMC phase inference, tensor-network diagnostics, and exact-diagonalization evidence supply the computational limits being combined.

Evidence cutoff. This method and evidence account covers primary and official sources available through 10 August 2026. Later replications, corrections, retractions, datasets, and changing assessments belong in the dated Quantum Matter and Emergence Research synthesis.

Represent each result as a directed graph from inputs to claim. Nodes include samples or ensembles, raw data, calibration, preprocessing, Hamiltonian, numerical representation, priors, fit code, and derived observables. Shared ancestors reveal dependence. If two observations D1,D2D_1,D_2 are conditionally independent only after shared nuisance η\eta is specified, the correct combination is

p(D1,D2θ)=dηp(D1θ,η)p(D2θ,η)p(η),p(D_1,D_2\mid\theta) =\int d\eta\, p(D_1\mid\theta,\eta) p(D_2\mid\theta,\eta) p(\eta),

not the product of separately marginalized likelihoods. The latter uses p(η)p(\eta) twice and can be overconfident.

Independence can come from a different operator, specimen, laboratory, numerical representation, boundary condition, or theoretical assumption. None is absolute. ARPES and STM use different matrix elements but may share the same surface. QMC and tensor networks use different approximations but may share one truncated Hamiltonian. Multiple analysis codes applied to one preprocessed dataset test software sensitivity, not experimental replication.

The chapter validity map makes these branches explicit. Inspect where evidence streams merge and which failure nodes remain shared.

Spectroscopy, local probes, cold atoms, QMC, tensor networks, and exact diagonalization converge through model comparison and held-out prediction, while shared samples, Hamiltonians, calibration, priors, finite limits, and hidden preprocessing reduce independence.

Triangulation is dependence-aware. Agreement is strongest when distinct operators and failure modes make a common quantitative prediction; shared ancestors, null results, and unresolved alternatives remain visible. Schematic.

Reproduction, replication, and falsification

Section titled “Reproduction, replication, and falsification”

Use the terms operationally, following the distinctions synthesized by the National Academies 2019, ch. 3:

  • computational reproduction: the same data and specified workflow regenerate the reported numbers;
  • robustness: reasonable changes of analysis, calibration, priors, or representation preserve the conclusion;
  • replication: independent data, preferably from a different sample or group, test the same claim;
  • generalization: the claim predicts a new regime, material, observable, or perturbation.

Reproduction is necessary for checking transformations, but replication and generalization add more independent scientific evidence. A decisive negative test is often more informative than another fit in the training window. Record null results and failed predictions with the same care as positive features. Peng 2011 explains why code and data are integral evidence for computational conclusions.

The FAIR principles make data and workflows findable, accessible, interoperable, and reusable Wilkinson et al. 2016; the FAIR4RS principles extend that structure to research software Barker et al. 2022. FAIR does not guarantee correctness. Independent checks still need tests, tolerances, and physically meaningful benchmarks.

For experiment, preserve sample identity and preparation, raw or minimally processed records, calibration versions, masks, geometry, environmental logs, background and resolution models, and covariance. For computation, preserve Hamiltonian and convention, source revision, dependency lock, compiler or runtime, seeds, input hashes, symmetry sectors, convergence criteria, and outputs. For inference, preserve candidate models, priors, nuisance variables, discrepancy, train–test split, likelihood, diagnostics, and posterior or profile samples.

Attach each claim to an exact evidence version and a supersession state. A later correction can weaken one branch without silently rewriting the historical record. The scientific page should state the present conclusion and date; the Research synthesis maintains the changing evidentiary status.

The probe and computation claim test matrix is the semantic comparison for this chapter. It is useful only when populated with actual versions, limits, and negative tests rather than checked mechanically.

Shared calibration. Two probes measure y1=θ+η+ϵ1y_1=\theta+\eta+\epsilon_1 and y2=2θ+η+ϵ2y_2=2\theta+\eta+\epsilon_2, with independent ϵi\epsilon_i but one uncertain offset η\eta. Why is combining their already marginalized one-dimensional likelihoods incorrect?

Solution

Marginalizing each likelihood separately integrates over an independent copy of η\eta. The physical measurements share the same offset, so their residuals are correlated after marginalization. The correct joint likelihood conditions both data on one η\eta and integrates it once. In this toy model, y2y1=θ+ϵ2ϵ1y_2-y_1=\theta+\epsilon_2-\epsilon_1 cancels the shared offset exactly and therefore constrains θ\theta without η\eta.

  • Michelle Barker, Neil P. Chue Hong, Daniel S. Katz, Anna-Lena Lamprecht, Carlos Martinez-Ortiz, Fotis Psomopoulos, Jennifer Harrow, Leyla Jael Castro, Morane Gruenpeter, Paula Andrea Martinez, and Tom Honeyman, “Introducing the FAIR Principles for Research Software,” Scientific Data 9 (2022) 622. DOI
  • National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, National Academies Press, 2019. DOI
  • Roger D. Peng, “Reproducible Research in Computational Science,” Science 334 (2011) 1226–1227. DOI
  • Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Mercè Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra Gonzalez-Beltran, Alasdair J. G. Gray, Paul Groth, Carole Goble, Jeffrey S. Grethe, Jaap Heringa, Peter A. C. ’t Hoen, Rob Hooft, Tobias Kuhn, Ruben Kok, Joost Kok, Scott J. Lusher, Maryann E. Martone, Albert Mons, Abel L. Packer, Bengt Persson, Philippe Rocca-Serra, Marco Roos, Rene van Schaik, Susanna-Assunta Sansone, Erik Schultes, Thierry Sengstag, Ted Slater, George Strawn, Morris A. Swertz, Mark Thompson, Johan van der Lei, Erik van Mulligen, Jan Velterop, Andra Waagmeester, Peter Wittenburg, Katherine Wolstencroft, Jun Zhao, and Barend Mons, “The FAIR Guiding Principles for Scientific Data Management and Stewardship,” Scientific Data 3 (2016) 160018. DOI