Probability Spaces, Random Variables, and Conditional Expectation
A probability model is a normalized positive measure together with measurable quantities defined on it. Once a probability space and measurable random variables have been specified, expectation is Lebesgue integration, independence is factorization of the relevant event sigma-algebras, and conditioning on available information is the almost-surely unique -measurable average that preserves integrals over every event in .
That last description is the definition of conditional expectation, not just an analogy. For , the random variable exists by the Radon–Nikodym theorem. It becomes an orthogonal projection only when . These distinctions matter in QFT: a finite Euclidean regulator with a positive, normalizable weight can define a genuine probability measure, whereas a Lorentzian factor , a sign-changing weight, or a merely formal continuum path integral does not.
Required background. Measures and Measurable Functions, supplies sigma-algebras, measurable maps, null sets, and pushforward measures; Lebesgue Integration and Convergence Theorems supplies integrability and the Radon–Nikodym theorem used to construct conditional expectations.
Probability spaces and conditional averages
Section titled “Probability spaces and conditional averages”Readers who want a finite-model refresher can use the statistical ensembles and probability repair.
All random variables below are real-valued unless a state space is displayed. The notation abbreviates . Equalities between random variables and conditional expectations are understood almost surely. The QFT-facing calculation is Euclidean and finite-dimensional; its coordinate measure is ordinary Lebesgue measure, not a continuum “flat measure” on fields.
Probability spaces: outcomes, events, and weights
Section titled “Probability spaces: outcomes, events, and weights”A probability space is a triple
with the following parts:
| Symbol | Mathematical type | Role |
|---|---|---|
| set | possible outcomes | |
| sigma-algebra of subsets of | events to which probabilities may be assigned | |
| measure on | normalized nonnegative weight, with |
Countable additivity means that for pairwise disjoint ,
The sigma-algebra is part of the model. A subset of that is not in is not an event, and is not required to assign it a number. This is essential once uncountable spaces are used; taking “all subsets” is generally neither necessary nor possible for the standard measures of analysis.
An event with is a null event. A statement holds almost surely when it fails only on a null event. Thus probability theory naturally identifies measurable functions that differ only on a null set. Whether has been completed by adding all subsets of null events is a separate choice and should be stated when it affects measurability.
The axioms above are the probability specialization of measure theory; see Durrett 2019, §§ 1.1–1.3, pp. 1–18, PDF for the probability-space, law, and measurable-map construction.
Random variables and their laws
Section titled “Random variables and their laws”Let be a measurable state space. A random element is a measurable map
meaning
When with its Borel sigma-algebra, is a random variable; when , it is a random vector. The adjective “random” does not mean that the map itself changes unpredictably. It means that the input is distributed according to , while measurability makes questions about the output into events.
The law or distribution of is the pushforward probability measure
For every nonnegative measurable , and for every for which either side is integrable, the pushforward identity gives
Consequently, the law is enough to calculate every one-variable expectation. It does not determine how is coupled to another variable on the same sample space; that information lives in the joint law.
A probability density is not the same object as a probability law. Relative to a reference measure , a density exists only when , in which case
Discrete, singular, and mixed laws need not have a density with respect to Lebesgue measure. Writing as though it were a density is therefore unsafe: for an absolutely continuous law every singleton can have probability zero even where the density is positive.
Expectation and integrability
Section titled “Expectation and integrability”For a nonnegative random variable , the expectation
is defined in . For a signed random variable, write , where and . Its expectation is defined when the expression
does not have the indeterminate form . The clean hypothesis used throughout this page is
Then is finite and linear. If , its variance is
The second equality requires the displayed second moments to be finite. Expectation is therefore an integral with a domain condition, not an automatic averaging operation for every formula called an observable.
Independence is factorization of information
Section titled “Independence is factorization of information”Two sub-sigma-algebras are independent when
Random elements and are independent when and are independent. Equivalently, their joint law factorizes:
For bounded measurable test functions and , this implies
Conversely, factorization for all bounded measurable test functions characterizes independence. A finite or infinite family is independent when the corresponding factorization holds for every finite subfamily; pairwise independence alone is weaker. The sigma-algebra formulation and its equivalent tests are developed in Durrett 2019, § 2.1, pp. 43–49, PDF.
Zero covariance is not independence
Section titled “Zero covariance is not independence”Let be uniform on and set . Then
Nevertheless is completely determined by . For example,
The joint law does not factorize, so the variables are dependent. Independence implies zero covariance when the relevant products are integrable, but the converse fails outside special families such as jointly Gaussian variables.
Conditional expectation from observable events
Section titled “Conditional expectation from observable events”Fix a sub-sigma-algebra . It represents the events whose occurrence can be decided from the information currently available. For , a conditional expectation of given is a random variable satisfying
We write
The first two lines say that the answer is an integrable random variable using only the information in . The last line says that it preserves the average of on every event that this information can distinguish. Equivalently,
for every bounded -measurable random variable . Indicator tests give the event identity, and approximation by simple functions gives the bounded-test identity.
Why it exists and why it is only almost-surely unique
Section titled “Why it exists and why it is only almost-surely unique”For , define a measure on by
It is absolutely continuous with respect to the restriction of to . The Radon–Nikodym theorem supplies a -measurable derivative , which obeys the defining integral identity. For signed , apply the same argument to and and subtract the two derivatives.
If and both satisfy the definition, then their integrals agree on every . In particular, for
one has
Thus every is null. Reversing and proves almost surely. Conditional expectation is therefore an equivalence class of versions, not a pointwise value fixed on null outcomes. The definition, Radon–Nikodym construction, and uniqueness argument are given in Durrett 2019, § 4.1, pp. 205–207, PDF; MIT OpenCourseWare 2014, 18.175 Lecture 25, slides 4–6, PDF provides a compact teaching route through the same definition.
Finite information: conditioning on a partition
Section titled “Finite information: conditioning on a partition”Suppose form a measurable partition of and . For ,
The right-hand side is constant on each information cell . Its integral over is exactly , so sums of cells verify the definition for every event in .
For a fair die, take to be the face value and reveal only its parity. Let
Then
For example,
The even cell gives on both sides, and the empty set and whole space then follow. The conditional expectation is not the realized die value; it is the average compatible with the revealed information.
For a single event with , the scalar conditional mean is
This quotient is undefined when . Conditioning on a sigma-algebra does not require dividing by the probability of each fiber and therefore remains meaningful in continuous models.
The conditional-expectation calculus
Section titled “The conditional-expectation calculus”Let and . Each identity below is an equality of almost-sure classes.
Linearity and order. For constants ,
Total expectation and the tower law. Taking in the definition gives
Conditioning in stages gives
In either order, the coarser sigma-algebra determines the final answer.
Pull out what is known. If is bounded and -measurable, then
Boundedness is a convenient sufficient condition. More general versions require the products on both sides to be integrable.
Independent information changes nothing. If is independent of , then
More generally, for bounded measurable ,
The statement for every bounded captures independence. The single identity for does not: it can hold for dependent variables.
These formulas follow by checking measurability and the defining integral identity. That “guess and verify” method is safer than manipulating the conditioning bar as though it were an ordinary fraction. Durrett 2019, § 4.1, pp. 207–213, PDF supplies the partition example and these properties.
The L² projection and variance split
Section titled “The L² projection and variance split”Now assume and set
Conditional Jensen’s inequality gives . For every ,
Thus is orthogonal to the closed subspace . For any , the Pythagorean identity is
Consequently, is the unique almost-sure minimizer of mean squared error among -measurable predictors. This is the precise regime in which conditional expectation is the “best prediction.” A different loss function can select a different conditional summary, and an variable need not admit this Hilbert-space interpretation.
Define the conditional variance by
Taking expectations and using the same orthogonal decomposition yields the law of total variance:
The projection theorem and its mean-square consequence are checked in Durrett 2019, § 4.1, p. 213, PDF.
Conditioning on a random variable
Section titled “Conditioning on a random variable”Conditioning on a random element means conditioning on the information it generates:
For real- or Euclidean-valued , the Doob–Dynkin factorization gives a Borel function such that
Only the equivalence class of under the law is fixed. Values of on a -null set can be changed without altering the conditional expectation.
If has a joint density , let
For an integrable , on the set where and
set
and define arbitrarily elsewhere. Then almost surely. The density and the ratio are only almost-everywhere objects; in particular, a fiber with does not determine . Thus is not obtained from by an event quotient when a continuous has .
There is a stronger object. A regular conditional distribution of given is a probability kernel such that
and is a version of . For standard Borel state spaces such a kernel can be chosen. On arbitrary measurable spaces it need not exist, and even in the standard case its values on -null fibers are not fixed by the joint law. Likewise,
defines an almost-sure class for each fixed ; without a regular version it is not automatically a pointwise probability measure in . These existence and null-fiber qualifications are treated in Durrett 2019, § 4.1.3, pp. 214–215, PDF.
Finite-dimensional Euclidean random fields
Section titled “Finite-dimensional Euclidean random fields”Consider a finite lattice or a finite mode truncation whose real coordinates are split into retained variables and hidden variables . Let be a measurable, dimensionless Euclidean action and assume
Then
is a probability measure on the Borel sets of . Choosing a positive normalizable Euclidean weight is the model input; the probability identities that follow do not establish that input physically.
Let and , and define the partial weight
For an integrable observable , set
where and the numerator is absolutely finite; define arbitrarily on the remaining marginal-null set. Tonelli’s and Fubini’s theorems show that these conditions hold almost everywhere relevant to the marginal law and that, for every bounded Borel ,
Therefore
This is the precise probabilistic content of “integrating out” the hidden finite-dimensional variables. A finite lattice turns field integration into finite-dimensional integration, while removing the regulator is a separate, nontrivial limit; see Münster and Walzl 2000, §§ 2.3–2.4, pp. 9–16.
Two coupled Gaussian modes
Section titled “Two coupled Gaussian modes”Take and
with
These conditions make the precision matrix
positive definite. Completing the square gives
The normalization and marginal precision are therefore
At fixed , the conditional density of is
It follows that
An independent check comes from
Indeed,
which agrees with the conditional mean. When , the density factorizes and the modes are independent. As at fixed , the variance and diverge, exposing the lost normalization hypothesis.
What this example does not establish
Section titled “What this example does not establish”- It uses a finite-dimensional Borel probability measure. It does not construct a continuum field measure or justify a symbol such as .
- A Lorentzian factor and a complex or sign-changing Euclidean weight are not probability densities, so the conditioning formulas do not apply to them as written.
- The calculation does not prove reflection positivity, analytic continuation, a continuum limit, or the existence of an interacting QFT.
- Wick contractions, Gaussian random distributions, Schwinger functions, and numerical sampling each require additional structure developed at their own destinations.
Common pitfalls
Section titled “Common pitfalls”Treating a density as the probability law. A density is a Radon–Nikodym derivative relative to a named reference measure. The law is the measure itself and may have no Lebesgue density.
Replacing independence by zero covariance. Covariance tests one product of centered variables. Independence requires factorization for all events, or equivalently for a separating class of test functions.
Conditioning on a null event by division. The quotient requires . Conditioning on a continuous observation is defined through and, when available, a regular conditional kernel whose null-fiber values remain version dependent.
Calling every conditional mean the best prediction. The minimizer claim is specifically an statement for squared-error loss. Under absolute loss, a conditional median is the relevant optimizer instead.
Reading every Euclidean weight probabilistically. Positivity and finite, nonzero normalization are indispensable. A formal measure symbol, a complex action, or an unconstructed continuum limit does not supply them.
Check your understanding
Section titled “Check your understanding”1. Separate a law from a density
Section titled “1. Separate a law from a density”Let be uniform on . Write its law, compute , and decide whether the law has a density with respect to Lebesgue measure.
Solution
The law is
Thus
The measure is concentrated on three points, a Lebesgue-null set, so it is not absolutely continuous with respect to Lebesgue measure and has no Lebesgue density.
2. Verify a partition conditional expectation
Section titled “2. Verify a partition conditional expectation”Let be a finite positive-probability partition. Verify the partition formula for an arbitrary event in .
Solution
Every event in the generated sigma-algebra is a union of some cells, say . For
finite additivity over the disjoint cells gives
The variable is constant on every cell and hence measurable with respect to the generated sigma-algebra. It is also integrable because
Thus all three defining conditions hold.
3. Derive total variance
Section titled “3. Derive total variance”For , set . Derive the law of total variance from .
Solution
The second term is -measurable, so orthogonality gives
Squaring the decomposition and taking expectations yields
Finally, total expectation applied to the conditional variance gives
which is the stated formula.
4. Check the Euclidean Gaussian boundary
Section titled “4. Check the Euclidean Gaussian boundary”For the two-mode action, integrate over and show why is needed. What happens as ?
Solution
Completing the square gives
The remaining integral is finite exactly when . With , this is , and then
As at fixed , the retained-mode precision tends to zero and its variance diverges. The normalized probability measure is lost at the boundary.
Synthesis and continuations
Section titled “Synthesis and continuations”Probability is fixed by the normalized measure ; random variables are measurable maps and their laws are pushforwards; expectation is integration; independence is factorization of the generated information; and conditional expectation is the -measurable object that preserves all -event integrals. It is unique only almost surely, becomes an orthogonal projection only in , and requires a regular conditional kernel before pointwise conditioning can be treated as a family of probability measures.
For the next mathematical step—convergence in probability, almost sure convergence, laws of large numbers, and central limit theorems—continue to Probabilistic Convergence, Laws of Large Numbers, and Central Limit Theorems. For the developed physical role of positive Euclidean field measures and their correlators, continue to Euclidean Correlators and Schwinger Functions.
References
Section titled “References”-
Rick Durrett, Probability: Theory and Examples, fifth edition, PDF, Cambridge University Press, 2019. §§ 1.1–1.3 and 1.6, pp. 1–18 and 28–34, support probability spaces, random variables, laws, and expectation; § 2.1, pp. 43–49, supports independence; and § 4.1, pp. 205–215, supports conditional expectation, its properties, the projection, and regular conditional distributions.
-
Gernot Münster and Marco Walzl, “Lattice Gauge Theory—A Short Primer”, arXiv:hep-lat/0012005, 2000, §§ 2.3–2.4, pp. 9–16, supports the positive Euclidean weight, finite-lattice reduction to finite-dimensional integration, and the separate continuum-limit problem.
-
Scott Sheffield, 18.175 Theory of Probability, Lecture 25, PDF, MIT OpenCourseWare, Spring 2014, slides 4–6. This is the teaching source for the defining integral identity, almost-sure uniqueness, and the Radon–Nikodym existence route.