Skip to content

Delta Distributions, Weak Derivatives, Pullbacks, and Pushforwards

The Dirac delta is not an infinitely high, infinitely narrow function. It is the continuous linear functional that evaluates a test function at a point. Once that definition is taken seriously, differentiation, localized sources, constraint surfaces, and changes of variables all follow from duality. The main lesson is equally important: every pullback or pushforward of a distribution comes with a geometric support or rank hypothesis.

Required background. Test-Function Spaces, Distributions, Support, and Convergence supplies test-function pairings, continuity, support, and distributional convergence.

We work first on open subsets of Euclidean space, pair distributions linearly with complex-valued test functions, and use absolute Jacobians because those pairings use Lebesgue measure. The first QFT application is the delta contact term in the free-scalar Green equation; the positive-energy mass shell is a second bounded application in the site’s (+)(+---) metric convention. The QFT interpretation of coincident operator products and contact terms belongs to Coincident Products and Contact Terms. Mathematical criteria for products and extensions remain in Products, Scaling Degree, and Distribution Extensions.

Let ΩRd\Omega\subseteq\mathbb R^d be open and let aΩa\in\Omega. The Dirac distribution at aa is

δaD(Ω),δa,φ=φ(a),φD(Ω).\delta_a\in\mathcal D'(\Omega), \qquad \langle\delta_a,\varphi\rangle=\varphi(a), \qquad \varphi\in\mathcal D(\Omega).

Evaluation is continuous in the test-function topology: if suppφKΩ\operatorname{supp}\varphi\subseteq K\Subset\Omega, then

δa,φsupKφ|\langle\delta_a,\varphi\rangle| \leq \sup_K|\varphi|

when aKa\in K, while the pairing vanishes when aKa\notin K. The notation

Rdδ(d)(xa)φ(x)ddx=φ(a)\int_{\mathbb R^d} \delta^{(d)}(x-a)\varphi(x)\,\mathrm d^d x = \varphi(a)

is therefore a pairing identity, not an instruction to assign a pointwise value to δ(d)\delta^{(d)}.

Translations and invertible linear changes of variables give the first useful calculus rule. If AGL(d,R)A\in GL(d,\mathbb R) and bRdb\in\mathbb R^d, then

δ(d)(Axb)=1detAδ(d)(xA1b)\delta^{(d)}(Ax-b) = \frac{1}{|\det A|} \delta^{(d)}(x-A^{-1}b)

as distributions in xx. Indeed, for every test function φ\varphi,

Rdδ(d)(Axb)φ(x)ddx=1detARdδ(d)(yb)φ(A1y)ddy=φ(A1b)detA.\begin{aligned} \int_{\mathbb R^d} \delta^{(d)}(Ax-b)\varphi(x)\,\mathrm d^d x &= \frac{1}{|\det A|} \int_{\mathbb R^d} \delta^{(d)}(y-b)\varphi(A^{-1}y)\,\mathrm d^d y \\ &= \frac{\varphi(A^{-1}b)}{|\det A|}. \end{aligned}

The absolute value is compulsory. Reversing orientation does not reverse a Lebesgue density. Differential forms carry orientation signs; the scalar distributions used here do not.

Distributional differentiation is integration by parts

Section titled “Distributional differentiation is integration by parts”

For a multi-index α\alpha, define the derivative of uD(Ω)u\in\mathcal D'(\Omega) by

αu,φ=(1)αu,αφ.\boxed{ \langle\partial^\alpha u,\varphi\rangle = (-1)^{|\alpha|} \langle u,\partial^\alpha\varphi\rangle. }

The right-hand side is again a continuous linear functional on D(Ω)\mathcal D(\Omega). Thus every distribution has derivatives of every order, the derivatives commute, and differentiation is sequentially continuous:

unu in Dαunαu.u_n\longrightarrow u\text{ in }\mathcal D' \quad\Longrightarrow\quad \partial^\alpha u_n\longrightarrow\partial^\alpha u.

For a regular distribution ufu_f with fC1f\in C^1, this definition agrees with the classical derivative after integration by parts. It also implies

supp(αu)suppu,\operatorname{supp}(\partial^\alpha u) \subseteq \operatorname{supp}u,

because a test function supported where uu vanishes has all its derivatives supported there as well. Dyatlov 2022, § 3.1, PDF proves these facts and works out the Heaviside and delta examples below.

In one dimension,

δa(n),φ=(1)nφ(n)(a).\langle\delta_a^{(n)},\varphi\rangle = (-1)^n\varphi^{(n)}(a).

More generally,

αδa,φ=(1)ααφ(a).\langle\partial^\alpha\delta_a,\varphi\rangle = (-1)^{|\alpha|}\partial^\alpha\varphi(a).

The delta probes the value of a test function; its derivatives probe the successive Taylor coefficients at the same point. They remain supported at {a}\{a\}.

If gC(Ω)g\in C^\infty(\Omega), multiplication and differentiation obey the usual Leibniz rule:

j(gu)=(jg)u+gju.\partial_j(gu) = (\partial_j g)u+g\,\partial_j u.

This is not a new pointwise product. It follows from the already defined smooth multiplication gu,φ=u,gφ\langle gu,\varphi\rangle=\langle u,g\varphi\rangle and the product rule for the smooth test function gφg\varphi.

Heaviside functions turn jumps into localized terms

Section titled “Heaviside functions turn jumps into localized terms”

Let HH be the locally integrable Heaviside function,

H(x)={0,x<0,1,x>0.H(x)= \begin{cases} 0,&x<0,\\ 1,&x>0. \end{cases}

Its value at the single point 00 is irrelevant to the associated regular distribution. For every φD(R)\varphi\in\mathcal D(\mathbb R),

H,φ=0φ(x)dx=φ(0)=δ0,φ.\begin{aligned} \langle H',\varphi\rangle &= -\int_0^\infty\varphi'(x)\,\mathrm dx \\ &= \varphi(0) = \langle\delta_0,\varphi\rangle. \end{aligned}

Therefore

H=δ0.H'=\delta_0.

Suppose the two one-sided pieces extend to C1C^1 functions ff_- and f+f_+ on a neighborhood of aa, and write

f(x)=H(ax)f(x)+H(xa)f+(x).f(x) = H(a-x)f_-(x)+H(x-a)f_+(x).

With [f]a=f+(a)f(a)[f]_a=f_+(a)-f_-(a), its distributional derivative is

xf=H(ax)f(x)+H(xa)f+(x)+[f]aδa.\partial_x f = H(a-x)f_-'(x) + H(x-a)f_+'(x) + [f]_a\delta_a.

The ordinary derivatives describe the two open regions, while the delta remembers the jump. A second derivative would also contain a [f]aδa[f]_a\delta_a' term, so higher derivatives retain progressively more interface data.

“Distributional derivative” and “weak derivative” are not synonyms

Section titled ““Distributional derivative” and “weak derivative” are not synonyms”

Every fLloc1(Ω)f\in L^1_{\mathrm{loc}}(\Omega) has a distributional derivative. One says that gg is the weak derivative of ff in a specified function space, for example gLlocpg\in L^p_{\mathrm{loc}}, only when

Ωgφddx=Ωfjφddx\int_\Omega g\varphi\,\mathrm d^d x = -\int_\Omega f\,\partial_j\varphi\,\mathrm d^d x

for every test function and gg has the stated regularity. A function with a nonzero jump has a perfectly good distributional derivative, but its delta term has no LlocpL^p_{\mathrm{loc}} representative for 1p1\leq p\leq\infty. The extra function-space claim therefore fails. This is the Sobolev-space usage in Evans 2010, Chapter 5, § 5.2.1.

Interfaces convert conservation laws into matching conditions

Section titled “Interfaces convert conservation laws into matching conditions”

Distributional differentiation makes surface sources explicit. Let g:ΩRg:\Omega\to\mathbb R be smooth with dg0\mathrm dg\neq0 on Σ=g1(0)\Sigma=g^{-1}(0), and let J+μJ^\mu_+ and JμJ^\mu_- be smooth currents on the two sides. Set

Jμ=H(g)J+μ+H(g)Jμ.J^\mu = H(g)J^\mu_+ + H(-g)J^\mu_-.

The chain rule for the regular-value pullback of HH gives

μJμ=H(g)μJ+μ+H(g)μJμ+δ(g)μg(J+μJμ).\partial_\mu J^\mu = H(g)\partial_\mu J^\mu_+ + H(-g)\partial_\mu J^\mu_- + \delta(g)\,\partial_\mu g \bigl(J^\mu_+-J^\mu_-\bigr).

If both bulk currents are conserved, the global current is conserved without an added surface source precisely when the normal flux is continuous:

μg(J+μJμ)Σ=0.\left. \partial_\mu g \bigl(J^\mu_+-J^\mu_-\bigr) \right|_\Sigma =0.

This calculation is independent of how gg is rescaled by a smooth positive factor: the transformation of δ(g)\delta(g) cancels the transformation of its normal covector on Σ\Sigma. It is the basic mechanism behind localized sources, junction conditions, and delta contact terms. The operator-product and time-ordering consequences are developed at the Coincident Products and Contact Terms page.

First QFT application: a free-scalar contact term

Section titled “First QFT application: a free-scalar contact term”

Let ϕ\phi be a free real scalar field in dd-dimensional Minkowski space, and define its time-ordered two-point distribution by

DF(xy)=0T{ϕ(x)ϕ(y)}0.D_F(x-y) = \langle0|T\{\phi(x)\phi(y)\}|0\rangle.

For bosonic operators, differentiating the step functions in the time-ordering symbol gives

x0T{A(x)B(y)}=T{0A(x)B(y)}+δ(x0y0)[A(x),B(y)].\begin{aligned} \partial_{x^0}T\{A(x)B(y)\} &= T\{\partial_0 A(x)B(y)\} \\ &\quad + \delta(x^0-y^0)[A(x),B(y)]. \end{aligned}

Use the equal-time canonical relations

[ϕ(t,x),ϕ(t,y)]=0,[π(t,x),ϕ(t,y)]=iδ(d1)(xy),π=0ϕ.\begin{aligned} [\phi(t,\mathbf x),\phi(t,\mathbf y)]&=0, \\ [\pi(t,\mathbf x),\phi(t,\mathbf y)] &= -i\delta^{(d-1)}(\mathbf x-\mathbf y), \qquad \pi=\partial_0\phi. \end{aligned}

The first time derivative of DFD_F has no contact term because the equal-time field commutator vanishes. Differentiating once more produces

x02DF(xy)=0T{02ϕ(x)ϕ(y)}0iδ(x0y0)δ(d1)(xy).\begin{aligned} \partial_{x^0}^2D_F(x-y) &= \langle0| T\{\partial_0^2\phi(x)\phi(y)\} |0\rangle \\ &\quad - i\delta(x^0-y^0) \delta^{(d-1)}(\mathbf x-\mathbf y). \end{aligned}

Spatial derivatives do not differentiate the time-ordering step functions. Combining this identity with the free equation (+m2)ϕ=0(\Box+m^2)\phi=0 gives

(x+m2)DF(xy)=iδ(d)(xy).(\Box_x+m^2)D_F(x-y) = -i\delta^{(d)}(x-y).

The field obeys the homogeneous equation away from coincidence, but its time-ordered product is an inhomogeneous Green distribution. The localized term is created by differentiating the ordering prescription and using the canonical commutator; it is not an ordinary pointwise source. Schwartz 2014, § 6.2, pp. 75–77, and § 7.1, Eqs. (7.9)–(7.10) gives this normalization and sign. The interpretation of such coincident-point terms continues at Coincident Products and Contact Terms.

Pullback requires a rank or singularity condition

Section titled “Pullback requires a rank or singularity condition”

Let Φ:UV\Phi:U\to V be a smooth map between open subsets of Euclidean spaces. For a smooth function ff there is always a classical pullback

Φf=fΦ.\Phi^*f=f\circ\Phi.

The same assertion is false for an arbitrary distribution. A map can send an entire region into the singular support of the object being pulled back, and a distribution has no pointwise values there to compose.

A diffeomorphism is safe and, more generally, so is any submersion for which dΦx\mathrm d\Phi_x is surjective at every xUx\in U.

For a diffeomorphism Φ:UV\Phi:U\to V, duality with change of variables fixes the formula

Φu,φ=u,detDΦ1φΦ1.\langle\Phi^*u,\varphi\rangle = \left\langle u,\, |\det D\Phi^{-1}|\, \varphi\circ\Phi^{-1} \right\rangle.

Here the determinant is evaluated at the argument of the test function on VV. If u=ufu=u_f is regular, this formula reduces to Φuf=ufΦ\Phi^*u_f=u_{f\circ\Phi}, exactly as required. Applied to a delta, it gives

Φδb=δΦ1(b)detDΦ(Φ1(b)).\Phi^*\delta_b = \frac{\delta_{\Phi^{-1}(b)}} {|\det D\Phi(\Phi^{-1}(b))|}.

For a projection

π:Rn+kRn,π(x,y)=x,\pi:\mathbb R^{n+k}\longrightarrow\mathbb R^n, \qquad \pi(x,y)=x,

the submersion pullback is characterized by

πu,φ=u,xRkφ(x,y)dky.\langle\pi^*u,\varphi\rangle = \left\langle u,\, x\longmapsto \int_{\mathbb R^k}\varphi(x,y)\,\mathrm d^k y \right\rangle.

Every submersion is locally a projection after a diffeomorphism. This factorization produces a unique sequentially continuous map

Φ:D(V)D(U)\Phi^*:\mathcal D'(V)\longrightarrow\mathcal D'(U)

that agrees with composition on locally integrable functions. Dyatlov 2022, Theorem 10.2, PDF gives the local submersion construction; Hörmander 2003, Theorem 6.1.1 supplies the structural treatment.

Submersion is a sufficient condition for pulling back every distribution, not a necessary condition for pulling back one particular distribution. The sharper later criterion compares the covectors normal to Φ\Phi with the distribution’s wavefront set. Singular Support and Wavefront Sets develops the mathematical criterion; the theorem-level criteria for products, pullbacks, and pushforwards in QFT belong to Wavefront-Set Products, Pullbacks, and Pushforwards.

Delta constraints are pullbacks to regular level sets

Section titled “Delta constraints are pullbacks to regular level sets”

Let g:ΩRg:\Omega\to\mathbb R be smooth and suppose 00 is a regular value:

g(x)0whenever g(x)=0.\nabla g(x)\neq0 \qquad \text{whenever }g(x)=0.

Then Σ=g1(0)\Sigma=g^{-1}(0) is a smooth hypersurface and the pullback gδ0g^*\delta_0, conventionally written δ(g)\delta(g), is well defined. To see why the regular-value hypothesis is enough, choose an open neighborhood WW of Σ\Sigma on which dg\mathrm dg never vanishes. The submersion theorem defines (gW)δ0(g|_W)^*\delta_0; it is supported on the closed set Σ\Sigma, so localization extends it uniquely by zero to Ω\Omega. The coarea formula gives

δ(g),φ=Σφ(x)g(x)dS(x).\boxed{ \langle\delta(g),\varphi\rangle = \int_\Sigma \frac{\varphi(x)}{|\nabla g(x)|}\, \mathrm dS(x). }

The norm and surface measure in this formula are Euclidean because the ambient pairing uses ddx\mathrm d^d x. The distribution depends on the normalization of the defining function: δ(ag)=a1δ(g)\delta(a g)=|a|^{-1}\delta(g) on the zero set for a smooth nowhere-zero factor aa. The geometric combination δ(g)dg\delta(g)\,\mathrm dg is unchanged when a>0a>0, while gδ(g)|\nabla g|\delta(g) is independent of the defining function altogether.

In one dimension, if the zeros xix_i of gg meeting suppφ\operatorname{supp}\varphi are simple, the formula becomes

δ(g(x))=g(xi)=0δ(xxi)g(xi).\delta(g(x)) = \sum_{g(x_i)=0} \frac{\delta(x-x_i)}{|g'(x_i)|}.

For a smooth map F:ΩRkF:\Omega\to\mathbb R^k of full rank kk along F1(0)F^{-1}(0), the codimension-kk version is

δ(k)(F),φ=F1(0)φ(x)JF(x)dHdk(x),\left\langle\delta^{(k)}(F),\varphi\right\rangle = \int_{F^{-1}(0)} \frac{\varphi(x)}{J_F(x)}\, \mathrm d\mathcal H^{d-k}(x),

where

JF(x)=det ⁣(DFxDFxT).J_F(x) = \sqrt{\det\!\bigl(DF_x\,DF_x^{\mathsf T}\bigr)}.

This is the precise meaning of imposing kk independent smooth constraints inside an ordinary integral. The same localization argument reduces the construction to a submersion near F1(0)F^{-1}(0). Federer 1969, § 3.2.22 supplies the codimension-kk coarea formula. Dyatlov 2022, Proposition 10.12, PDF proves the hypersurface formula.

The regular-value hypothesis cannot be deleted. For g(x)=x2g(x)=x^2, the only zero is critical and the simple-zero rule would demand division by zero. The expression δ(x2)\delta(x^2) is not a canonical pullback of δ0\delta_0. Assigning one requires additional extension data; it is not justified by the delta change-of-variables rule.

Pushforward requires properness on support

Section titled “Pushforward requires properness on support”

Pullback transports a generalized function against the direction of a map. Pushforward transports a localized distribution in the direction of the map. Let Φ:UV\Phi:U\to V be smooth and uD(U)u\in\mathcal D'(U). Assume that Φ\Phi is proper on suppu\operatorname{supp}u:

Φ1(K)suppuis compact for every compact KV.\Phi^{-1}(K)\cap\operatorname{supp}u \quad\text{is compact for every compact }K\subset V.

Then the pushforward is defined by

Φu,ψ=u,ψΦ.\langle\Phi_*u,\psi\rangle = \langle u,\psi\circ\Phi\rangle.

Although ψΦ\psi\circ\Phi need not have compact support in all of UU, its intersection with suppu\operatorname{supp}u is compact. Multiplying it by a cutoff equal to 11 near that intersection makes the right-hand side a legitimate distributional pairing, independent of the cutoff. A compactly supported uu therefore has a pushforward under every smooth Φ\Phi.

Three examples separate this operation from pullback:

Φδa=δΦ(a).\Phi_*\delta_a=\delta_{\Phi(a)}.

For the projection π(x,y)=x\pi(x,y)=x and a compactly supported regular distribution ufu_f,

πuf=uffib,ffib(x)=Rkf(x,y)dky.\pi_*u_f = u_{f_{\mathrm{fib}}}, \qquad f_{\mathrm{fib}}(x) = \int_{\mathbb R^k}f(x,y)\,\mathrm d^k y.

For a diffeomorphism and a regular distribution,

Φuf=uf~,f~(y)=f(Φ1(y))detDΦ1(y).\Phi_*u_f = u_{\widetilde f}, \qquad \widetilde f(y) = f(\Phi^{-1}(y)) |\det D\Phi^{-1}(y)|.

Thus a pushforward integrates a density along fibers and includes the Jacobian appropriate to that integration. A pullback of the regular function ff is instead just fΦf\circ\Phi. The proper-on-support definition is given explicitly in Dinh and Sibony 2005, § 2.2, PDF.

Without properness, ψΦ\psi\circ\Phi may probe a noncompact part of suppu\operatorname{supp}u, so an ordinary distribution need not be able to act on it. For example, pushing the constant regular distribution on R2\mathbb R^2 through the projection (x,y)x(x,y)\mapsto x would require the divergent fiber integral Rdy\int_{\mathbb R}\mathrm dy. A regulator or a different distribution space would be extra structure, not part of the definition above.

Secondary QFT application: the positive-energy mass shell

Section titled “Secondary QFT application: the positive-energy mass shell”

Use the site’s (+)(+---) convention in dd spacetime dimensions:

p2=(p0)2p2,Ep=p2+m2,m>0.p^2=(p^0)^2-|\mathbf p|^2, \qquad E_{\mathbf p} = \sqrt{|\mathbf p|^2+m^2}, \qquad m>0.

At fixed p\mathbf p, the mass-shell constraint has two simple roots:

(p0)2Ep2=(p0Ep)(p0+Ep).(p^0)^2-E_{\mathbf p}^2 = (p^0-E_{\mathbf p})(p^0+E_{\mathbf p}).

The one-dimensional regular-zero formula therefore gives

δ(p2m2)=δ(p0Ep)+δ(p0+Ep)2Ep.\delta(p^2-m^2) = \frac{ \delta(p^0-E_{\mathbf p}) + \delta(p^0+E_{\mathbf p}) } {2E_{\mathbf p}}.

The shell has two components separated by the open strip p0<m|p^0|<m. Choose χC(R)\chi\in C^\infty(\mathbb R) with

χ(p0)=0for p0m2,χ(p0)=1for p0m2,\chi(p^0)=0\quad\text{for }p^0\leq-\frac m2, \qquad \chi(p^0)=1\quad\text{for }p^0\geq\frac m2,

and define

δ+(p2m2)=χ(p0)δ(p2m2),δ(p2m2)=(1χ(p0))δ(p2m2).\delta_+(p^2-m^2) = \chi(p^0)\delta(p^2-m^2), \qquad \delta_-(p^2-m^2) = \bigl(1-\chi(p^0)\bigr)\delta(p^2-m^2).

These distributions do not depend on the chosen interpolation of χ\chi, because the shell does not meet that interpolation region. The conventional notation

δ±(p2m2)=θ(±p0)δ(p2m2)\delta_\pm(p^2-m^2) = \theta(\pm p^0)\delta(p^2-m^2)

is shorthand for this smooth localization, not an unrestricted product with a discontinuous function.

Selecting the future sheet now yields, for every test function Ψ\Psi,

Rdddpδ+(p2m2)Ψ(p)=Rd1dd1p2EpΨ(Ep,p).\begin{aligned} \int_{\mathbb R^d} \mathrm d^d p\, \delta_+(p^2-m^2)\Psi(p) &= \int_{\mathbb R^{d-1}} \frac{\mathrm d^{d-1}\mathbf p}{2E_{\mathbf p}}\, \Psi(E_{\mathbf p},\mathbf p). \end{aligned}

Equivalently, let

ι+:Rd1Rd,ι+(p)=(Ep,p).\iota_+:\mathbb R^{d-1}\longrightarrow\mathbb R^d, \qquad \iota_+(\mathbf p)=(E_{\mathbf p},\mathbf p).

Then the equality of measures on Rd\mathbb R^d is

δ+(p2m2)ddp=(ι+)(dd1p2Ep).\delta_+(p^2-m^2)\,\mathrm d^d p = (\iota_+)_* \left( \frac{\mathrm d^{d-1}\mathbf p}{2E_{\mathbf p}} \right).

This measure is invariant under proper orthochronous Lorentz transformations. Factors of (2π)(d1)(2\pi)^{-(d-1)} are conventional and have deliberately not been included. With the common scattering convention, one writes

dΠp=dd1p(2π)d12Ep.\mathrm d\Pi_p = \frac{\mathrm d^{d-1}\mathbf p} {(2\pi)^{d-1}2E_{\mathbf p}}.

This is the localized shell constraint underlying relativistic phase-space integrals. The complete multiparticle construction belongs to Lorentz-Invariant Phase Space. For d=4d=4, Schwartz 2014, § 5.1, p. 61, Eq. (5.21) gives d3p/[(2π)32Ep]\mathrm d^3\mathbf p/[(2\pi)^3 2E_{\mathbf p}]. The displayed dd-dimensional version follows from the same one-variable delta calculation.

Before manipulating a delta or another distribution, identify five pieces of data:

  1. Test space and meaning. Is the object in D\mathcal D', S\mathcal S', or a more specialized dual? A statement such as δ(0)=\delta(0)=\infty gives no pairing and therefore defines nothing.
  2. Derivative claim. Distributional derivatives always exist. A weak derivative in LpL^p is an additional representation claim and may fail when a jump produces a delta.
  3. Map and direction. A pullback needs a submersion or a valid singularity criterion; a pushforward needs properness on support. Keep the absolute Jacobian in pullback formulas.
  4. Product. Smooth functions may multiply distributions, but the rules on this page do not define δ2\delta^2 or an arbitrary product of singular distributions. The extension problem belongs to Products, Scaling Degree, and Distribution Extensions.
  5. Pairing test. After moving the proposed operation to the test function, is the resulting probe smooth and compactly supported where required?

Stop when one of these checks fails. Writing one line of duality usually exposes the missing hypothesis before formal delta notation can hide it.

  1. Let f(x)=H(xa)(xa+1)f(x)=H(x-a)(x-a+1). Compute its first distributional derivative.

    Check

    The smooth factor has value 11 at aa. The Leibniz rule gives

    xf=(xa+1)δa+H(xa)=δa+H(xa).\partial_x f = (x-a+1)\delta_a+H(x-a) = \delta_a+H(x-a).

    Pairing directly with a test function gives the same boundary term.

  2. Let a>0a>0. Determine δ(x2a2)\delta(x^2-a^2) and explain why the corresponding formula does not apply at a=0a=0.

    Check

    The zeros are x=±ax=\pm a, and (x2a2)=2a|(x^2-a^2)'|=2a at each one. Hence

    δ(x2a2)=δ(xa)+δ(x+a)2a.\delta(x^2-a^2) = \frac{\delta(x-a)+\delta(x+a)}{2a}.

    At a=0a=0, the sole zero is critical because the derivative of x2x^2 vanishes there. The regular-zero pullback theorem gives no distribution δ(x2)\delta(x^2).

  3. For Φ(x)=αx+β\Phi(x)=\alpha x+\beta with α0\alpha\neq0, compare Φδb\Phi^*\delta_b and Φδa\Phi_*\delta_a.

    Check

    Pullback solves the constraint Φ(x)=b\Phi(x)=b and includes its Jacobian:

    Φδb=1αδbβα.\Phi^*\delta_b = \frac{1}{|\alpha|} \delta_{\frac{b-\beta}{\alpha}}.

    Pushforward simply transports the point mass:

    Φδa=δαa+β.\Phi_*\delta_a = \delta_{\alpha a+\beta}.

    The two operations point in opposite directions and answer different questions.

  4. Integrate Ψ(p0,p)\Psi(p^0,\mathbf p) over the negative-energy mass shell.

    Check

    The negative-sheet localization δ\delta_- retains only the root p0=Epp^0=-E_{\mathbf p}:

    ddpδ(p2m2)Ψ(p)=dd1p2EpΨ(Ep,p).\int\mathrm d^d p\, \delta_-(p^2-m^2)\Psi(p) = \int \frac{\mathrm d^{d-1}\mathbf p}{2E_{\mathbf p}}\, \Psi(-E_{\mathbf p},\mathbf p).

    The Jacobian is positive on both sheets because it is 2p0=2Ep|2p^0|=2E_{\mathbf p}.

  • Tien-Cuong Dinh and Nessim Sibony, Introduction to the Theory of Currents, § 2.2, PDF, dated September 21, 2005. This supplies the explicit pushforward definition under properness on support.
  • Semyon Dyatlov, Lecture Notes for 18.155: Differential Analysis, PDF, §§ 3.1–3.2 and 10.1, MIT, 2022. These are the teaching sources for distributional derivatives, smooth multiplication, submersion pullbacks, the chain rule, and hypersurface deltas.
  • Lawrence C. Evans, Partial Differential Equations, 2nd ed., Chapter 5, § 5.2.1, American Mathematical Society, 2010. Book record. This fixes the function-space meaning of weak derivatives used here.
  • Herbert Federer, Geometric Measure Theory, § 3.2.22, Springer, 1969. Book record. This is the source for the coarea formula and its codimension-kk Jacobian.
  • Lars Hörmander, The Analysis of Linear Partial Differential Operators I: Distribution Theory and Fourier Analysis, 2nd ed., Chapters 3 and 6, Springer, 2003. Book record. This is the structural source for differentiation, composition with smooth maps, and the limits of unrestricted pullback.
  • Matthew D. Schwartz, Quantum Field Theory and the Standard Model, § 5.1, p. 61, Eq. (5.21); § 6.2, pp. 75–77; and § 7.1, Eqs. (7.9)–(7.10), Cambridge University Press, 2014. Book record. This is the QFT source for the free-scalar contact term and for the four-dimensional on-shell factor d3p/[(2π)32Ep]\mathrm d^3\mathbf p/[(2\pi)^3 2E_{\mathbf p}]; the page derives the latter’s dd-dimensional extension.