Skip to content

Probability Spaces, Random Variables, and Conditional Expectation

A probability model is a normalized positive measure together with measurable quantities defined on it. Once a probability space (Ω,F,P)(\Omega,\mathcal F,\mathbb P) and measurable random variables have been specified, expectation is Lebesgue integration, independence is factorization of the relevant event sigma-algebras, and conditioning on available information GF\mathcal G\subseteq\mathcal F is the almost-surely unique G\mathcal G-measurable average that preserves integrals over every event in G\mathcal G.

That last description is the definition of conditional expectation, not just an analogy. For XL1X\in L^1, the random variable E[XG]\mathbb E[X\mid\mathcal G] exists by the Radon–Nikodym theorem. It becomes an orthogonal projection only when XL2X\in L^2. These distinctions matter in QFT: a finite Euclidean regulator with a positive, normalizable weight can define a genuine probability measure, whereas a Lorentzian factor eiSe^{iS}, a sign-changing weight, or a merely formal continuum path integral does not.

Required background. Measures and Measurable Functions, supplies sigma-algebras, measurable maps, null sets, and pushforward measures; Lebesgue Integration and Convergence Theorems supplies integrability and the Radon–Nikodym theorem used to construct conditional expectations.

Probability spaces and conditional averages

Section titled “Probability spaces and conditional averages”

Readers who want a finite-model refresher can use the statistical ensembles and probability repair.

All random variables below are real-valued unless a state space is displayed. The notation Lp(F)L^p(\mathcal F) abbreviates Lp(Ω,F,P)L^p(\Omega,\mathcal F,\mathbb P). Equalities between random variables and conditional expectations are understood almost surely. The QFT-facing calculation is Euclidean and finite-dimensional; its coordinate measure is ordinary Lebesgue measure, not a continuum “flat measure” on fields.

Probability spaces: outcomes, events, and weights

Section titled “Probability spaces: outcomes, events, and weights”

A probability space is a triple

(Ω,F,P),(\Omega,\mathcal F,\mathbb P),

with the following parts:

SymbolMathematical typeRole
Ω\Omegasetpossible outcomes ω\omega
F\mathcal Fsigma-algebra of subsets of Ω\Omegaevents to which probabilities may be assigned
P\mathbb Pmeasure on (Ω,F)(\Omega,\mathcal F)normalized nonnegative weight, with P(Ω)=1\mathbb P(\Omega)=1

Countable additivity means that for pairwise disjoint A1,A2,FA_1,A_2,\ldots\in\mathcal F,

P ⁣(n=1An)=n=1P(An).\mathbb P\!\left(\bigcup_{n=1}^{\infty}A_n\right) =\sum_{n=1}^{\infty}\mathbb P(A_n).

The sigma-algebra is part of the model. A subset of Ω\Omega that is not in F\mathcal F is not an event, and P\mathbb P is not required to assign it a number. This is essential once uncountable spaces are used; taking “all subsets” is generally neither necessary nor possible for the standard measures of analysis.

An event NFN\in\mathcal F with P(N)=0\mathbb P(N)=0 is a null event. A statement holds almost surely when it fails only on a null event. Thus probability theory naturally identifies measurable functions that differ only on a null set. Whether F\mathcal F has been completed by adding all subsets of null events is a separate choice and should be stated when it affects measurability.

The axioms above are the probability specialization of measure theory; see Durrett 2019, §§ 1.1–1.3, pp. 1–18, PDF for the probability-space, law, and measurable-map construction.

Let (S,S)(S,\mathcal S) be a measurable state space. A random element is a measurable map

X:(Ω,F)(S,S),X:(\Omega,\mathcal F)\longrightarrow(S,\mathcal S),

meaning

X1(B)Ffor every BS.X^{-1}(B)\in\mathcal F \qquad \text{for every }B\in\mathcal S.

When S=RS=\mathbb R with its Borel sigma-algebra, XX is a random variable; when S=RnS=\mathbb R^n, it is a random vector. The adjective “random” does not mean that the map itself changes unpredictably. It means that the input ω\omega is distributed according to P\mathbb P, while measurability makes questions about the output into events.

The law or distribution of XX is the pushforward probability measure

μX=X#P,μX(B)=P(XB)=P ⁣(X1(B)).\begin{aligned} \mu_X &=X_{\#}\mathbb P,\\ \mu_X(B) &=\mathbb P(X\in B) =\mathbb P\!\left(X^{-1}(B)\right). \end{aligned}

For every nonnegative measurable ff, and for every ff for which either side is integrable, the pushforward identity gives

E[f(X)]=Ωf(X(ω))dP(ω)=Sf(x)dμX(x).\mathbb E[f(X)] =\int_{\Omega}f(X(\omega))\,\mathrm d\mathbb P(\omega) =\int_S f(x)\,\mathrm d\mu_X(x).

Consequently, the law is enough to calculate every one-variable expectation. It does not determine how XX is coupled to another variable on the same sample space; that information lives in the joint law.

A probability density is not the same object as a probability law. Relative to a reference measure λ\lambda, a density exists only when μXλ\mu_X\ll\lambda, in which case

pX=dμXdλ.p_X=\frac{\mathrm d\mu_X}{\mathrm d\lambda}.

Discrete, singular, and mixed laws need not have a density with respect to Lebesgue measure. Writing P(X=x)\mathbb P(X=x) as though it were a density is therefore unsafe: for an absolutely continuous law every singleton can have probability zero even where the density is positive.

For a nonnegative random variable XX, the expectation

E[X]=ΩXdP\mathbb E[X]=\int_\Omega X\,\mathrm d\mathbb P

is defined in [0,][0,\infty]. For a signed random variable, write X=X+XX=X^+-X^-, where X+=max(X,0)X^+=\max(X,0) and X=max(X,0)X^-=\max(-X,0). Its expectation is defined when the expression

E[X+]E[X]\mathbb E[X^+]-\mathbb E[X^-]

does not have the indeterminate form \infty-\infty. The clean hypothesis used throughout this page is

XL1E[X]<.X\in L^1 \quad\Longleftrightarrow\quad \mathbb E[|X|]<\infty.

Then E[X]\mathbb E[X] is finite and linear. If XL2X\in L^2, its variance is

Var(X)=E ⁣[(XE[X])2]=E[X2]E[X]2.\operatorname{Var}(X) =\mathbb E\!\left[(X-\mathbb E[X])^2\right] =\mathbb E[X^2]-\mathbb E[X]^2.

The second equality requires the displayed second moments to be finite. Expectation is therefore an integral with a domain condition, not an automatic averaging operation for every formula called an observable.

Independence is factorization of information

Section titled “Independence is factorization of information”

Two sub-sigma-algebras A,BF\mathcal A,\mathcal B\subseteq\mathcal F are independent when

P(AB)=P(A)P(B)for every AA, BB.\mathbb P(A\cap B)=\mathbb P(A)\mathbb P(B) \qquad \text{for every }A\in\mathcal A,\ B\in\mathcal B.

Random elements XX and YY are independent when σ(X)\sigma(X) and σ(Y)\sigma(Y) are independent. Equivalently, their joint law factorizes:

μ(X,Y)=μXμY.\mu_{(X,Y)}=\mu_X\otimes\mu_Y.

For bounded measurable test functions ff and gg, this implies

E[f(X)g(Y)]=E[f(X)]E[g(Y)].\mathbb E[f(X)g(Y)] =\mathbb E[f(X)]\,\mathbb E[g(Y)].

Conversely, factorization for all bounded measurable test functions characterizes independence. A finite or infinite family is independent when the corresponding factorization holds for every finite subfamily; pairwise independence alone is weaker. The sigma-algebra formulation and its equivalent tests are developed in Durrett 2019, § 2.1, pp. 43–49, PDF.

Let UU be uniform on {1,0,1}\{-1,0,1\} and set V=U2V=U^2. Then

E[U]=0,E[V]=23,E[UV]=E[U3]=0,Cov(U,V)=0.\begin{aligned} \mathbb E[U]&=0, &\mathbb E[V]&=\frac23,\\ \mathbb E[UV]&=\mathbb E[U^3]=0, &\operatorname{Cov}(U,V)&=0. \end{aligned}

Nevertheless VV is completely determined by UU. For example,

P(U=0,V=0)=13,P(U=0)P(V=0)=19.\begin{aligned} \mathbb P(U=0,V=0)&=\frac13,\\ \mathbb P(U=0)\mathbb P(V=0)&=\frac19. \end{aligned}

The joint law does not factorize, so the variables are dependent. Independence implies zero covariance when the relevant products are integrable, but the converse fails outside special families such as jointly Gaussian variables.

Conditional expectation from observable events

Section titled “Conditional expectation from observable events”

Fix a sub-sigma-algebra GF\mathcal G\subseteq\mathcal F. It represents the events whose occurrence can be decided from the information currently available. For XL1X\in L^1, a conditional expectation of XX given G\mathcal G is a random variable MM satisfying

M is G-measurable,E[M]<,AMdP=AXdPfor every AG.\begin{aligned} M&\text{ is }\mathcal G\text{-measurable},\\ \mathbb E[|M|]&<\infty,\\ \int_A M\,\mathrm d\mathbb P &=\int_A X\,\mathrm d\mathbb P \qquad \text{for every }A\in\mathcal G. \end{aligned}

We write

M=E[XG].M=\mathbb E[X\mid\mathcal G].

The first two lines say that the answer is an integrable random variable using only the information in G\mathcal G. The last line says that it preserves the average of XX on every event that this information can distinguish. Equivalently,

E[ZX]=E ⁣[ZE[XG]]\mathbb E[ZX] =\mathbb E\!\left[Z\,\mathbb E[X\mid\mathcal G]\right]

for every bounded G\mathcal G-measurable random variable ZZ. Indicator tests give the event identity, and approximation by simple functions gives the bounded-test identity.

Why it exists and why it is only almost-surely unique

Section titled “Why it exists and why it is only almost-surely unique”

For X0X\geq0, define a measure on (Ω,G)(\Omega,\mathcal G) by

ν(A)=AXdP.\nu(A)=\int_A X\,\mathrm d\mathbb P.

It is absolutely continuous with respect to the restriction of P\mathbb P to G\mathcal G. The Radon–Nikodym theorem supplies a G\mathcal G-measurable derivative M=dν/dPM=\mathrm d\nu/\mathrm d\mathbb P, which obeys the defining integral identity. For signed XL1X\in L^1, apply the same argument to X+X^+ and XX^- and subtract the two derivatives.

If MM and MM' both satisfy the definition, then their integrals agree on every AGA\in\mathcal G. In particular, for

An={MM+1/n}G,A_n=\{M\geq M'+1/n\}\in\mathcal G,

one has

0=An(MM)dP1nP(An).0=\int_{A_n}(M-M')\,\mathrm d\mathbb P \geq\frac1n\mathbb P(A_n).

Thus every AnA_n is null. Reversing MM and MM' proves M=MM=M' almost surely. Conditional expectation is therefore an equivalence class of versions, not a pointwise value fixed on null outcomes. The definition, Radon–Nikodym construction, and uniqueness argument are given in Durrett 2019, § 4.1, pp. 205–207, PDF; MIT OpenCourseWare 2014, 18.175 Lecture 25, slides 4–6, PDF provides a compact teaching route through the same definition.

Finite information: conditioning on a partition

Section titled “Finite information: conditioning on a partition”

Suppose A1,,AmA_1,\ldots,A_m form a measurable partition of Ω\Omega and P(Ai)>0\mathbb P(A_i)>0. For G=σ(A1,,Am)\mathcal G=\sigma(A_1,\ldots,A_m),

E[XG]=i=1mE[X1Ai]P(Ai)1Ai.\boxed{ \mathbb E[X\mid\mathcal G] =\sum_{i=1}^{m} \frac{\mathbb E[X\mathbf 1_{A_i}]}{\mathbb P(A_i)} \mathbf 1_{A_i}. }

The right-hand side is constant on each information cell AiA_i. Its integral over AiA_i is exactly E[X1Ai]\mathbb E[X\mathbf 1_{A_i}], so sums of cells verify the definition for every event in G\mathcal G.

For a fair die, take XX to be the face value and reveal only its parity. Let

O={1,3,5},E={2,4,6},G=σ(O,E).O=\{1,3,5\}, \qquad E=\{2,4,6\}, \qquad \mathcal G=\sigma(O,E).

Then

E[XG]=31O+41E.\mathbb E[X\mid\mathcal G] =3\mathbf 1_O+4\mathbf 1_E.

For example,

E[X1O]=1+3+56=32,E[31O]=3P(O)=32.\begin{aligned} \mathbb E[X\mathbf 1_O] &=\frac{1+3+5}{6}=\frac32,\\ \mathbb E[3\mathbf 1_O] &=3\,\mathbb P(O)=\frac32. \end{aligned}

The even cell gives 22 on both sides, and the empty set and whole space then follow. The conditional expectation is not the realized die value; it is the average compatible with the revealed information.

For a single event BB with P(B)>0\mathbb P(B)>0, the scalar conditional mean is

E[XB]=E[X1B]P(B).\mathbb E[X\mid B] =\frac{\mathbb E[X\mathbf 1_B]}{\mathbb P(B)}.

This quotient is undefined when P(B)=0\mathbb P(B)=0. Conditioning on a sigma-algebra does not require dividing by the probability of each fiber and therefore remains meaningful in continuous models.

Let X,YL1X,Y\in L^1 and HGF\mathcal H\subseteq\mathcal G\subseteq\mathcal F. Each identity below is an equality of almost-sure classes.

Linearity and order. For constants a,ba,b,

E[aX+bYG]=aE[XG]+bE[YG],XYE[XG]E[YG].\begin{aligned} \mathbb E[aX+bY\mid\mathcal G] &=a\mathbb E[X\mid\mathcal G] +b\mathbb E[Y\mid\mathcal G],\\ X\leq Y &\Longrightarrow \mathbb E[X\mid\mathcal G] \leq\mathbb E[Y\mid\mathcal G]. \end{aligned}

Total expectation and the tower law. Taking A=ΩA=\Omega in the definition gives

E ⁣[E[XG]]=E[X].\mathbb E\!\left[\mathbb E[X\mid\mathcal G]\right] =\mathbb E[X].

Conditioning in stages gives

E ⁣[E[XG]H]=E[XH],E ⁣[E[XH]G]=E[XH].\begin{aligned} \mathbb E\!\left[ \mathbb E[X\mid\mathcal G] \mid\mathcal H \right] &=\mathbb E[X\mid\mathcal H],\\ \mathbb E\!\left[ \mathbb E[X\mid\mathcal H] \mid\mathcal G \right] &=\mathbb E[X\mid\mathcal H]. \end{aligned}

In either order, the coarser sigma-algebra H\mathcal H determines the final answer.

Pull out what is known. If ZZ is bounded and G\mathcal G-measurable, then

E[ZXG]=ZE[XG].\mathbb E[ZX\mid\mathcal G] =Z\,\mathbb E[X\mid\mathcal G].

Boundedness is a convenient sufficient condition. More general versions require the products on both sides to be integrable.

Independent information changes nothing. If XX is independent of G\mathcal G, then

E[XG]=E[X].\mathbb E[X\mid\mathcal G]=\mathbb E[X].

More generally, for bounded measurable ff,

E[f(X)G]=E[f(X)].\mathbb E[f(X)\mid\mathcal G]=\mathbb E[f(X)].

The statement for every bounded ff captures independence. The single identity for f(x)=xf(x)=x does not: it can hold for dependent variables.

These formulas follow by checking measurability and the defining integral identity. That “guess and verify” method is safer than manipulating the conditioning bar as though it were an ordinary fraction. Durrett 2019, § 4.1, pp. 207–213, PDF supplies the partition example and these properties.

Now assume XL2X\in L^2 and set

M=E[XG].M=\mathbb E[X\mid\mathcal G].

Conditional Jensen’s inequality gives ML2(G)M\in L^2(\mathcal G). For every ZL2(G)Z\in L^2(\mathcal G),

E[(XM)Z]=0.\mathbb E[(X-M)Z]=0.

Thus XMX-M is orthogonal to the closed subspace L2(G)L2(F)L^2(\mathcal G)\subseteq L^2(\mathcal F). For any YL2(G)Y\in L^2(\mathcal G), the Pythagorean identity is

E[(XY)2]=E[(XM)2]+E[(MY)2].\begin{aligned} \mathbb E[(X-Y)^2] ={}&\mathbb E[(X-M)^2]\\ &+\mathbb E[(M-Y)^2]. \end{aligned}

Consequently, MM is the unique almost-sure minimizer of mean squared error among G\mathcal G-measurable L2L^2 predictors. This is the precise regime in which conditional expectation is the “best prediction.” A different loss function can select a different conditional summary, and an L1L^1 variable need not admit this Hilbert-space interpretation.

Define the conditional variance by

Var(XG)=E[(XM)2G]=E[X2G]M2.\begin{aligned} \operatorname{Var}(X\mid\mathcal G) &=\mathbb E[(X-M)^2\mid\mathcal G]\\ &=\mathbb E[X^2\mid\mathcal G]-M^2. \end{aligned}

Taking expectations and using the same orthogonal decomposition yields the law of total variance:

Var(X)=E[Var(XG)]+Var(E[XG]).\boxed{ \operatorname{Var}(X) =\mathbb E[\operatorname{Var}(X\mid\mathcal G)] +\operatorname{Var}(\mathbb E[X\mid\mathcal G]). }

The projection theorem and its mean-square consequence are checked in Durrett 2019, § 4.1, p. 213, PDF.

Conditioning on a random element YY means conditioning on the information it generates:

E[XY]E[Xσ(Y)].\mathbb E[X\mid Y] \equiv\mathbb E[X\mid\sigma(Y)].

For real- or Euclidean-valued YY, the Doob–Dynkin factorization gives a Borel function gg such that

E[XY]=g(Y)almost surely.\mathbb E[X\mid Y]=g(Y) \qquad\text{almost surely}.

Only the equivalence class of gg under the law μY\mu_Y is fixed. Values of g(y)g(y) on a μY\mu_Y-null set can be changed without altering the conditional expectation.

If (X,Y)(X,Y) has a joint density pX,Yp_{X,Y}, let

pY(y)=RpX,Y(x,y)dx.p_Y(y)=\int_{\mathbb R}p_{X,Y}(x,y)\,\mathrm dx.

For an integrable h(X)h(X), on the set where 0<pY(y)<0<p_Y(y)<\infty and

Rh(x)pX,Y(x,y)dx<,\int_{\mathbb R}|h(x)|p_{X,Y}(x,y)\,\mathrm dx<\infty,

set

g(y)=Rh(x)pX,Y(x,y)dxpY(y)g(y) =\frac{ \displaystyle\int_{\mathbb R}h(x)p_{X,Y}(x,y)\,\mathrm dx }{p_Y(y)}

and define gg arbitrarily elsewhere. Then E[h(X)Y]=g(Y)\mathbb E[h(X)\mid Y]=g(Y) almost surely. The density and the ratio are only almost-everywhere objects; in particular, a fiber with pY(y)=0p_Y(y)=0 does not determine g(y)g(y). Thus E[XY=y]\mathbb E[X\mid Y=y] is not obtained from P(Y=y)\mathbb P(Y=y) by an event quotient when a continuous YY has P(Y=y)=0\mathbb P(Y=y)=0.

There is a stronger object. A regular conditional distribution of XX given YY is a probability kernel KK such that

yK(y,B)is measurable for every B,BK(y,B)is a probability measure for every y,\begin{aligned} y&\longmapsto K(y,B) &&\text{is measurable for every }B,\\ B&\longmapsto K(y,B) &&\text{is a probability measure for every }y, \end{aligned}

and K(Y,B)K(Y,B) is a version of P(XBσ(Y))\mathbb P(X\in B\mid\sigma(Y)). For standard Borel state spaces such a kernel can be chosen. On arbitrary measurable spaces it need not exist, and even in the standard case its values on μY\mu_Y-null fibers are not fixed by the joint law. Likewise,

P(AG)=E[1AG]\mathbb P(A\mid\mathcal G) =\mathbb E[\mathbf 1_A\mid\mathcal G]

defines an almost-sure class for each fixed AA; without a regular version it is not automatically a pointwise probability measure in AA. These existence and null-fiber qualifications are treated in Durrett 2019, § 4.1.3, pp. 214–215, PDF.

Finite-dimensional Euclidean random fields

Section titled “Finite-dimensional Euclidean random fields”

Consider a finite lattice or a finite mode truncation whose real coordinates are split into retained variables ϕRn\phi\in\mathbb R^n and hidden variables χRm\chi\in\mathbb R^m. Let SE:Rn+mRS_E:\mathbb R^{n+m}\to\mathbb R be a measurable, dimensionless Euclidean action and assume

0<Z=Rn+meSE(ϕ,χ)dnϕdmχ<.0<Z =\int_{\mathbb R^{n+m}} e^{-S_E(\phi,\chi)} \,\mathrm d^n\phi\,\mathrm d^m\chi <\infty.

Then

dP(ϕ,χ)=1ZeSE(ϕ,χ)dnϕdmχ\mathrm d\mathbb P(\phi,\chi) =\frac1Z e^{-S_E(\phi,\chi)} \,\mathrm d^n\phi\,\mathrm d^m\chi

is a probability measure on the Borel sets of Rn+m\mathbb R^{n+m}. Choosing a positive normalizable Euclidean weight is the model input; the probability identities that follow do not establish that input physically.

Let Φ(ϕ,χ)=ϕ\Phi(\phi,\chi)=\phi and Ξ(ϕ,χ)=χ\Xi(\phi,\chi)=\chi, and define the partial weight

Zχ(ϕ)=RmeSE(ϕ,χ)dmχ.Z_\chi(\phi) =\int_{\mathbb R^m} e^{-S_E(\phi,\chi)}\,\mathrm d^m\chi.

For an integrable observable O(ϕ,χ)\mathcal O(\phi,\chi), set

g(ϕ)=RmO(ϕ,χ)eSE(ϕ,χ)dmχZχ(ϕ)g(\phi) =\frac{ \displaystyle\int_{\mathbb R^m} \mathcal O(\phi,\chi)e^{-S_E(\phi,\chi)} \,\mathrm d^m\chi }{Z_\chi(\phi)}

where 0<Zχ(ϕ)<0<Z_\chi(\phi)<\infty and the numerator is absolutely finite; define gg arbitrarily on the remaining marginal-null set. Tonelli’s and Fubini’s theorems show that these conditions hold almost everywhere relevant to the marginal law and that, for every bounded Borel hh,

E[h(Φ)O]=1ZRnh(ϕ)[RmO(ϕ,χ)eSEdmχ]dnϕ=E[h(Φ)g(Φ)].\begin{aligned} \mathbb E[h(\Phi)\mathcal O] &=\frac1Z\int_{\mathbb R^n} h(\phi) \left[ \int_{\mathbb R^m} \mathcal O(\phi,\chi)e^{-S_E} \,\mathrm d^m\chi \right] \mathrm d^n\phi\\ &=\mathbb E[h(\Phi)g(\Phi)]. \end{aligned}

Therefore

E[Oσ(Φ)]=g(Φ)almost surely.\boxed{ \mathbb E[\mathcal O\mid\sigma(\Phi)] =g(\Phi) \quad\text{almost surely}. }

This is the precise probabilistic content of “integrating out” the hidden finite-dimensional variables. A finite lattice turns field integration into finite-dimensional integration, while removing the regulator is a separate, nontrivial limit; see Münster and Walzl 2000, §§ 2.3–2.4, pp. 9–16.

Take n=m=1n=m=1 and

SE(ϕ,χ)=12(aϕ2+2bϕχ+cχ2),S_E(\phi,\chi) =\frac12\left( a\phi^2+2b\phi\chi+c\chi^2 \right),

with

c>0,Δ=acb2>0.c>0, \qquad \Delta=ac-b^2>0.

These conditions make the precision matrix

K=(abbc)K= \begin{pmatrix} a & b\\ b & c \end{pmatrix}

positive definite. Completing the square gives

aϕ2+2bϕχ+cχ2=c(χ+bcϕ)2+(ab2c)ϕ2.\begin{aligned} a\phi^2+2b\phi\chi+c\chi^2 ={}&c\left(\chi+\frac bc\phi\right)^2\\ &+\left(a-\frac{b^2}{c}\right)\phi^2. \end{aligned}

The normalization and marginal precision are therefore

Z=2πΔ,Var(Φ)=(ab2c)1=cΔ.Z=\frac{2\pi}{\sqrt\Delta}, \qquad \operatorname{Var}(\Phi) =\left(a-\frac{b^2}{c}\right)^{-1} =\frac c\Delta.

At fixed ϕ\phi, the conditional density of χ\chi is

p(χϕ)=c2πexp ⁣[c2(χ+bcϕ)2].p(\chi\mid\phi) =\sqrt{\frac{c}{2\pi}} \exp\!\left[ -\frac c2 \left(\chi+\frac bc\phi\right)^2 \right].

It follows that

E[ΞΦ]=bcΦ,E[Ξ2Φ]=1c+b2c2Φ2.\boxed{ \begin{aligned} \mathbb E[\Xi\mid\Phi] &=-\frac bc\Phi,\\ \mathbb E[\Xi^2\mid\Phi] &=\frac1c+\frac{b^2}{c^2}\Phi^2. \end{aligned} }

An independent check comes from

K1=1Δ(cbba).K^{-1} =\frac1\Delta \begin{pmatrix} c & -b\\ -b & a \end{pmatrix}.

Indeed,

Cov(Ξ,Φ)Var(Φ)=b/Δc/Δ=bc,\frac{\operatorname{Cov}(\Xi,\Phi)} {\operatorname{Var}(\Phi)} =\frac{-b/\Delta}{c/\Delta} =-\frac bc,

which agrees with the conditional mean. When b=0b=0, the density factorizes and the modes are independent. As Δ0\Delta\downarrow0 at fixed c>0c>0, the variance and ZZ diverge, exposing the lost normalization hypothesis.

  • It uses a finite-dimensional Borel probability measure. It does not construct a continuum field measure or justify a symbol such as xdϕ(x)\prod_x\mathrm d\phi(x).
  • A Lorentzian factor eiSe^{iS} and a complex or sign-changing Euclidean weight are not probability densities, so the conditioning formulas do not apply to them as written.
  • The calculation does not prove reflection positivity, analytic continuation, a continuum limit, or the existence of an interacting QFT.
  • Wick contractions, Gaussian random distributions, Schwinger functions, and numerical sampling each require additional structure developed at their own destinations.

Treating a density as the probability law. A density is a Radon–Nikodym derivative relative to a named reference measure. The law is the measure itself and may have no Lebesgue density.

Replacing independence by zero covariance. Covariance tests one product of centered variables. Independence requires factorization for all events, or equivalently for a separating class of test functions.

Conditioning on a null event by division. The quotient P(AB)/P(B)\mathbb P(A\cap B)/\mathbb P(B) requires P(B)>0\mathbb P(B)>0. Conditioning on a continuous observation is defined through σ(Y)\sigma(Y) and, when available, a regular conditional kernel whose null-fiber values remain version dependent.

Calling every conditional mean the best prediction. The minimizer claim is specifically an L2L^2 statement for squared-error loss. Under absolute loss, a conditional median is the relevant optimizer instead.

Reading every Euclidean weight probabilistically. Positivity and finite, nonzero normalization are indispensable. A formal measure symbol, a complex action, or an unconstructed continuum limit does not supply them.

Let UU be uniform on {1,0,1}\{-1,0,1\}. Write its law, compute E[U2]\mathbb E[U^2], and decide whether the law has a density with respect to Lebesgue measure.

Solution

The law is

μU=13δ1+13δ0+13δ1.\mu_U =\frac13\delta_{-1} +\frac13\delta_0 +\frac13\delta_1.

Thus

E[U2]=x2dμU(x)=23.\mathbb E[U^2] =\int x^2\,\mathrm d\mu_U(x) =\frac23.

The measure is concentrated on three points, a Lebesgue-null set, so it is not absolutely continuous with respect to Lebesgue measure and has no Lebesgue density.

2. Verify a partition conditional expectation

Section titled “2. Verify a partition conditional expectation”

Let A1,,AmA_1,\ldots,A_m be a finite positive-probability partition. Verify the partition formula for an arbitrary event in σ(A1,,Am)\sigma(A_1,\ldots,A_m).

Solution

Every event BB in the generated sigma-algebra is a union of some cells, say B=iIAiB=\bigcup_{i\in I}A_i. For

M=iE[X1Ai]P(Ai)1Ai,M=\sum_i \frac{\mathbb E[X\mathbf 1_{A_i}]}{\mathbb P(A_i)} \mathbf 1_{A_i},

finite additivity over the disjoint cells gives

BMdP=iIE[X1Ai]=E[X1B]=BXdP.\begin{aligned} \int_B M\,\mathrm d\mathbb P &=\sum_{i\in I}\mathbb E[X\mathbf 1_{A_i}]\\ &=\mathbb E[X\mathbf 1_B] =\int_B X\,\mathrm d\mathbb P. \end{aligned}

The variable MM is constant on every cell and hence measurable with respect to the generated sigma-algebra. It is also integrable because

E[M]=iE[X1Ai]iE[X1Ai]=E[X]<.\begin{aligned} \mathbb E[|M|] &=\sum_i \left|\mathbb E[X\mathbf 1_{A_i}]\right|\\ &\leq\sum_i\mathbb E[|X|\mathbf 1_{A_i}] =\mathbb E[|X|]<\infty. \end{aligned}

Thus all three defining conditions hold.

For XL2X\in L^2, set M=E[XG]M=\mathbb E[X\mid\mathcal G]. Derive the law of total variance from XE[X]=(XM)+(ME[X])X-\mathbb E[X]=(X-M)+(M-\mathbb E[X]).

Solution

The second term is G\mathcal G-measurable, so orthogonality gives

E ⁣[(XM)(ME[X])]=0.\mathbb E\!\left[ (X-M)(M-\mathbb E[X]) \right]=0.

Squaring the decomposition and taking expectations yields

Var(X)=E[(XM)2]+Var(M).\operatorname{Var}(X) =\mathbb E[(X-M)^2] +\operatorname{Var}(M).

Finally, total expectation applied to the conditional variance gives

E[(XM)2]=E[Var(XG)],\mathbb E[(X-M)^2] =\mathbb E[\operatorname{Var}(X\mid\mathcal G)],

which is the stated formula.

For the two-mode action, integrate over χ\chi and show why Δ=acb2>0\Delta=ac-b^2>0 is needed. What happens as Δ0\Delta\downarrow0?

Solution

Completing the square gives

ReSE(ϕ,χ)dχ=2πcexp ⁣[12Δcϕ2].\int_{\mathbb R}e^{-S_E(\phi,\chi)}\,\mathrm d\chi =\sqrt{\frac{2\pi}{c}} \exp\!\left[ -\frac12\frac{\Delta}{c}\phi^2 \right].

The remaining ϕ\phi integral is finite exactly when Δ/c>0\Delta/c>0. With c>0c>0, this is Δ>0\Delta>0, and then

Z=2πc2πcΔ=2πΔ.Z =\sqrt{\frac{2\pi}{c}} \sqrt{\frac{2\pi c}{\Delta}} =\frac{2\pi}{\sqrt\Delta}.

As Δ0\Delta\downarrow0 at fixed c>0c>0, the retained-mode precision tends to zero and its variance c/Δc/\Delta diverges. The normalized probability measure is lost at the boundary.

Probability is fixed by the normalized measure (Ω,F,P)(\Omega,\mathcal F,\mathbb P); random variables are measurable maps and their laws are pushforwards; expectation is integration; independence is factorization of the generated information; and conditional expectation is the G\mathcal G-measurable L1L^1 object that preserves all G\mathcal G-event integrals. It is unique only almost surely, becomes an orthogonal projection only in L2L^2, and requires a regular conditional kernel before pointwise conditioning can be treated as a family of probability measures.

For the next mathematical step—convergence in probability, almost sure convergence, laws of large numbers, and central limit theorems—continue to Probabilistic Convergence, Laws of Large Numbers, and Central Limit Theorems. For the developed physical role of positive Euclidean field measures and their correlators, continue to Euclidean Correlators and Schwinger Functions.

  • Rick Durrett, Probability: Theory and Examples, fifth edition, PDF, Cambridge University Press, 2019. §§ 1.1–1.3 and 1.6, pp. 1–18 and 28–34, support probability spaces, random variables, laws, and expectation; § 2.1, pp. 43–49, supports independence; and § 4.1, pp. 205–215, supports conditional expectation, its properties, the L2L^2 projection, and regular conditional distributions.

  • Gernot Münster and Marco Walzl, “Lattice Gauge Theory—A Short Primer”, arXiv:hep-lat/0012005, 2000, §§ 2.3–2.4, pp. 9–16, supports the positive Euclidean weight, finite-lattice reduction to finite-dimensional integration, and the separate continuum-limit problem.

  • Scott Sheffield, 18.175 Theory of Probability, Lecture 25, PDF, MIT OpenCourseWare, Spring 2014, slides 4–6. This is the teaching source for the defining integral identity, almost-sure uniqueness, and the Radon–Nikodym existence route.