Evidence Triangulation and Reproducibility
Evidence triangulation combines analytic constraints, numerical calculations, and experiments that fail in different ways. Its force comes from conditional independence and successful cross-prediction, not from counting plots. Two results that use the same dataset, Hamiltonian, calibration, prior, or preprocessing may be valuable consistency checks while contributing far less independent evidence than their number suggests.
Required background. Model selection and parameter inference supplies the joint likelihood. QMC phase inference, tensor-network diagnostics, and exact-diagonalization evidence supply the computational limits being combined.
Evidence cutoff. This method and evidence account covers primary and official sources available through 10 August 2026. Later replications, corrections, retractions, datasets, and changing assessments belong in the dated Quantum Matter and Emergence Research synthesis.
Map dependence before combining evidence
Section titled “Map dependence before combining evidence”Represent each result as a directed graph from inputs to claim. Nodes include samples or ensembles, raw data, calibration, preprocessing, Hamiltonian, numerical representation, priors, fit code, and derived observables. Shared ancestors reveal dependence. If two observations are conditionally independent only after shared nuisance is specified, the correct combination is
not the product of separately marginalized likelihoods. The latter uses twice and can be overconfident.
Independence can come from a different operator, specimen, laboratory, numerical representation, boundary condition, or theoretical assumption. None is absolute. ARPES and STM use different matrix elements but may share the same surface. QMC and tensor networks use different approximations but may share one truncated Hamiltonian. Multiple analysis codes applied to one preprocessed dataset test software sensitivity, not experimental replication.
The chapter validity map makes these branches explicit. Inspect where evidence streams merge and which failure nodes remain shared.
Triangulation is dependence-aware. Agreement is strongest when distinct operators and failure modes make a common quantitative prediction; shared ancestors, null results, and unresolved alternatives remain visible. Schematic.
Reproduction, replication, and falsification
Section titled “Reproduction, replication, and falsification”Use the terms operationally, following the distinctions synthesized by the National Academies 2019, ch. 3:
- computational reproduction: the same data and specified workflow regenerate the reported numbers;
- robustness: reasonable changes of analysis, calibration, priors, or representation preserve the conclusion;
- replication: independent data, preferably from a different sample or group, test the same claim;
- generalization: the claim predicts a new regime, material, observable, or perturbation.
Reproduction is necessary for checking transformations, but replication and generalization add more independent scientific evidence. A decisive negative test is often more informative than another fit in the training window. Record null results and failed predictions with the same care as positive features. Peng 2011 explains why code and data are integral evidence for computational conclusions.
The FAIR principles make data and workflows findable, accessible, interoperable, and reusable Wilkinson et al. 2016; the FAIR4RS principles extend that structure to research software Barker et al. 2022. FAIR does not guarantee correctness. Independent checks still need tests, tolerances, and physically meaningful benchmarks.
Minimum reproducible object
Section titled “Minimum reproducible object”For experiment, preserve sample identity and preparation, raw or minimally processed records, calibration versions, masks, geometry, environmental logs, background and resolution models, and covariance. For computation, preserve Hamiltonian and convention, source revision, dependency lock, compiler or runtime, seeds, input hashes, symmetry sectors, convergence criteria, and outputs. For inference, preserve candidate models, priors, nuisance variables, discrepancy, train–test split, likelihood, diagnostics, and posterior or profile samples.
Attach each claim to an exact evidence version and a supersession state. A later correction can weaken one branch without silently rewriting the historical record. The scientific page should state the present conclusion and date; the Research synthesis maintains the changing evidentiary status.
The probe and computation claim test matrix is the semantic comparison for this chapter. It is useful only when populated with actual versions, limits, and negative tests rather than checked mechanically.
Exercise
Section titled “Exercise”Shared calibration. Two probes measure and , with independent but one uncertain offset . Why is combining their already marginalized one-dimensional likelihoods incorrect?
Solution
Marginalizing each likelihood separately integrates over an independent copy of . The physical measurements share the same offset, so their residuals are correlated after marginalization. The correct joint likelihood conditions both data on one and integrates it once. In this toy model, cancels the shared offset exactly and therefore constrains without .
References
Section titled “References”- Michelle Barker, Neil P. Chue Hong, Daniel S. Katz, Anna-Lena Lamprecht, Carlos Martinez-Ortiz, Fotis Psomopoulos, Jennifer Harrow, Leyla Jael Castro, Morane Gruenpeter, Paula Andrea Martinez, and Tom Honeyman, “Introducing the FAIR Principles for Research Software,” Scientific Data 9 (2022) 622. DOI
- National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, National Academies Press, 2019. DOI
- Roger D. Peng, “Reproducible Research in Computational Science,” Science 334 (2011) 1226–1227. DOI
- Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Appleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Mercè Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra Gonzalez-Beltran, Alasdair J. G. Gray, Paul Groth, Carole Goble, Jeffrey S. Grethe, Jaap Heringa, Peter A. C. ’t Hoen, Rob Hooft, Tobias Kuhn, Ruben Kok, Joost Kok, Scott J. Lusher, Maryann E. Martone, Albert Mons, Abel L. Packer, Bengt Persson, Philippe Rocca-Serra, Marco Roos, Rene van Schaik, Susanna-Assunta Sansone, Erik Schultes, Thierry Sengstag, Ted Slater, George Strawn, Morris A. Swertz, Mark Thompson, Johan van der Lei, Erik van Mulligen, Jan Velterop, Andra Waagmeester, Peter Wittenburg, Katherine Wolstencroft, Jun Zhao, and Barend Mons, “The FAIR Guiding Principles for Scientific Data Management and Stewardship,” Scientific Data 3 (2016) 160018. DOI