Cross-Program Evidence and Falsifier Matrix
Comparing quantum-gravity programs requires more than listing which observations each can accommodate. A fair comparison asks whether a signal was derived before inspection, whether its parameter pattern is distinctive, whether the evidence is independent, and what result would make the program’s relevant low-energy realization untenable. Generic compatibility with successful EFT and general relativity is necessary but not positive discrimination.
Required background. Quantum-Gravity Observables and Test Taxonomy fixes the claim hierarchy, and Model Selection and Parameter Inference supplies statistical comparison. Helpful background. Correlated Evidence, Independence, and Triangulation and Evidence Triangulation and Reproducibility prevent double counting; Analogue and Phenomenological Evidence Ceilings fixes the ceiling of mechanism-level evidence.
The program-to-observable comparison object
Section titled “The program-to-observable comparison object”For program , observable class , and dataset , record
The first arrow must be a controlled derivation. If a program admits many vacua, states, discretizations, truncations, or phenomenological ansätze, those choices belong in and carry prior or complexity cost. A generic statement that “some realization may produce the effect” gives no usable likelihood.
The comparison record needs at least:
- the measured observable and data release;
- the derived mechanism, coefficient correlations, and validity regime;
- parameters fixed before the observation and parameters adjusted afterward;
- conventional and cross-program alternatives;
- shared datasets, calibrations, EFT operators, and source models;
- a quantitative or logical falsifier;
- the maximum conclusion supported by a positive or null result.
If two programs descend to the same unconstrained EFT operator, the experiment tests that operator, not the programs. Discrimination requires different coefficient relations, symmetries, scaling laws, or accompanying observables.
First application: four probes across five frameworks
Section titled “First application: four probes across five frameworks”Consider string constructions, loop-based approaches, asymptotic safety, discrete causal or combinatorial approaches, and program-neutral EFT. Map four representative probes:
Lorentz test. A photon dispersion coefficient is directly testable. Some realizations of several programs preserve Lorentz symmetry exactly; others can produce preferred-frame or deformed-kinematics effects. Unless a framework and state derive the coefficient’s sign, species, helicity, and scale, a null result constrains the EFT coefficient and those specific realizations—not the whole program.
Tabletop witness. Gravity-mediated entanglement tests classes of mediator models under locality and subsystem assumptions. Standard quantized low-energy gravity predicts an entangling phase, but so do many UV completions because they share that limit. A positive witness could reject a specified classical-channel class; it would not distinguish string, loop, safety, or discrete UV structure without further correlated effects.
Gravitational-wave probe. Dispersion, birefringence, and extra polarizations constrain low-energy propagation operators. A UV program whose controlled infrared limit predicts exactly general relativity expects a null result. Compatibility is not a program-specific success because all viable programs must recover the same tested regime.
Cosmological constraint. An initial-state oscillation or primordial tensor spectrum becomes discriminating only if the state, phase, amplitude, and correlated higher-point functions are derived. Reconstructing a template after seeing a feature is a postdiction shared by many early-universe models.
This exercise yields a substantive result: in the empirical sources reviewed here through 10 August 2026, most proposed phenomenological channels constrain program-neutral EFTs or broad mediator classes, and none reports a confirmed, unique selection among these quantum-gravity programs.
Independence and Bayesian comparison
Section titled “Independence and Bayesian comparison”For programs and datasets , one may write
The tempting factorization fails when datasets share calibration, population models, foregrounds, or the same EFT matching assumption. A hierarchical model should introduce shared nuisances explicitly. Without it, counting a photon timing bound, a threshold bound, and a polarization bound as three independent successes may exaggerate evidence if all rely on one coefficient restriction.
Prior-predictive precision matters. A model that predicted a narrow coefficient pattern before measurement earns more from agreement than one with flexible parameters spanning all outcomes. Conversely, a failed sharp prediction is more informative than an unconstrained postdiction.
Adversarial control: remove borrowed support
Section titled “Adversarial control: remove borrowed support”Repeat the comparison after deleting:
- effects inferred only after the relevant data were known;
- multiple analyses of the same events or sky maps;
- tests that constrain the same freely adjustable EFT coefficient;
- generic recovery of general relativity or ordinary QFT;
- predictions outside a controlled continuum or semiclassical regime.
Then ask whether any likelihood difference remains. If not, the correct conclusion is empirical non-discrimination, even if the programs have unequal theoretical development. Theoretical consistency, mathematical definition, semiclassical recovery, explanatory reach, and empirical evidence are separate axes and should not be collapsed into one score.
Falsifiers should target defined claims
Section titled “Falsifiers should target defined claims”A whole research program is rarely falsified by one low-energy null result because programs contain multiple regimes and realizations. But individual claims can be: a derived nonzero dispersion coefficient can be excluded; a universal bounce spectrum can fail; a specific compact-object model can contradict ringdown and imaging jointly; a mediator model can violate a witnessed correlation bound.
A useful falsifier therefore names the construction, state, approximation, parameter range, observable, and decision rule. Flexibility introduced only after failure counts against predictive force.
The low-energy EFT viewpoint explains why shared infrared coefficients do not automatically discriminate ultraviolet programs Donoghue 1994; the observable-first comparison used here follows the broader phenomenology framework of Amelino-Camelia 2013.
The chapter overview contains the structure diagram and validity and failure diagram. They are embedded there once so that their shared chapter-level context is not repeated on every article.
For the chapter-wide comparison of assumptions, counterevidence, falsifiers, and claim ceilings, see the claim-domain table.
References
Section titled “References”- Amelino-Camelia, G. “Quantum-Spacetime Phenomenology.” Living Reviews in Relativity 16, 5 (2013). DOI.
- Donoghue, J. F. “General Relativity as an Effective Field Theory: The Leading Quantum Corrections.” Physical Review D 50, 3874–3888 (1994). DOI.
- Liberati, S. “Tests of Lorentz Invariance: A 2013 Update.” Classical and Quantum Gravity 30, 133001 (2013). DOI.