Skip to content

Vector Fields and Gauge Redundancy

Spin-zero and spin-one-half fields can be quantized without introducing an obvious redundancy in the field variables. Vector fields are different. A massive vector field has three physical polarizations, and the Proca equation enforces this by a constraint. A massless vector field, however, has only two physical helicities. If we insist on using a Lorentz four-vector Aμ(x)A_\mu(x) to describe those two helicities, we have necessarily introduced more variables than physical degrees of freedom.

This is the origin of gauge redundancy. The Maxwell field is not merely a vector field with m=0m=0. It is a vector field whose physically meaningful content is invariant under

Aμ(x)⟼Aμ(x)+∂μα(x).A_\mu(x)\longmapsto A_\mu(x)+\partial_\mu\alpha(x).

The redundancy is not an optional decoration. It is what makes a local Lorentz-covariant description of massless spin one possible, and it is the reason conserved currents, covariant derivatives, Ward identities, and gauge fixing all enter QFT together.

There are three statements to keep separate throughout the page. First, AμA_\mu is a Lorentz four-vector field. Second, a photon is a massless spin-one particle with two helicities. Third, a gauge choice is a representative of a redundant description, not an observable. Many mistakes in gauge theory come from treating these three statements as if they were the same.

Required background. Scalar and vector representations and mass shells supplies the Lorentz-representation and polarization language used below. Helpful background. Dirac fields and spinors supplies the on-shell spinor identities used when matter is introduced.

Massive vector fields and the Proca constraint

Section titled “Massive vector fields and the Proca constraint”

A Lorentz vector field transforms as

A′μ(x′)=ΛμνAν(x),x′=Λx.A'^\mu(x')=\Lambda^\mu{}_{\nu}A^\nu(x), \qquad x'=\Lambda x.

As a field representation this has four components. A massive spin-one particle, however, has only three spin states. The Proca Lagrangian implements the required reduction:

LProca=−14FμνFμν+12M2AμAμ,Fμν=∂μAν−∂νAμ.\mathcal L_{\mathrm{Proca}} =-{1\over4}F_{\mu\nu}F^{\mu\nu} +{1\over2}M^2A_\mu A^\mu, \qquad F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu.

Varying AνA_\nu gives

∂μFμν+M2Aν=0.\partial_\mu F^{\mu\nu}+M^2A^\nu=0.

Taking the divergence gives

∂ν∂μFμν+M2∂νAν=0.\partial_\nu\partial_\mu F^{\mu\nu}+M^2\partial_\nu A^\nu=0.

The first term vanishes because FμνF^{\mu\nu} is antisymmetric while ∂ν∂μ\partial_\nu\partial_\mu is symmetric. Therefore, for M≠0M\neq0,

∂μAμ=0.\partial_\mu A^\mu=0.

Substituting this back into the equation of motion gives

(∂2+M2)Aν=0.(\partial^2+M^2)A^\nu=0.

For a plane wave

Aμ(x)=ϵμ(p)e−ip⋅x,A^\mu(x)=\epsilon^\mu(p)e^{-ip\cdot x},

the equations become

p2=M2,pμϵμ(p)=0.p^2=M^2, \qquad p_\mu\epsilon^\mu(p)=0.

The second condition is the Proca transversality constraint. It removes one of the four components of ϵμ\epsilon^\mu, leaving three physical polarizations.

This condition is an equation-of-motion constraint, not a gauge choice. The mass term removes the Maxwell gauge redundancy: for M≠0M\neq0, two Proca configurations that differ by ∂μα\partial_\mu\alpha are generally physically different. In Hamiltonian language A0A_0 has no independent second-order evolution equation; eliminating it leaves three propagating canonical modes.

In the rest frame pμ=(M,0,0,0)p^\mu=(M,0,0,0), the condition p⋅ϵ=0p\cdot\epsilon=0 gives

Mϵ0=0,M\epsilon^0=0,

so ϵ0=0\epsilon^0=0. The remaining three components are an ordinary spatial vector. A convenient basis is

ϵ1μ=(0,1,0,0),ϵ2μ=(0,0,1,0),ϵ3μ=(0,0,0,1).\epsilon_1^\mu=(0,1,0,0), \qquad \epsilon_2^\mu=(0,0,1,0), \qquad \epsilon_3^\mu=(0,0,0,1).

Thus the Proca field contains precisely the spin-one representation of the massive little group SO(3)SO(3).

The massive polarization sum is

∑λ=13ϵλμ(p)ϵλ∗ν(p)=−ημν+pμpνM2,\sum_{\lambda=1}^{3} \epsilon_\lambda^\mu(p)\epsilon_\lambda^{*\nu}(p) =-\eta^{\mu\nu}+{p^\mu p^\nu\over M^2},

where the physical polarizations are normalized by

ϵλ∗(p)⋅ϵλ′(p)=−δλλ′.\epsilon_\lambda^*(p)\cdot\epsilon_{\lambda'}(p)=-\delta_{\lambda\lambda'}.

The polarization sum on the right is singular as M→0M\to0. As a mixed-index map CμνC^\mu{}_\nu, it satisfies C2=−CC^2=-C; the idempotent projection onto p⊥p^\perp is T=−CT=-C. This distinction and the conjugation needed for a complex basis were derived in lesson 32 and the canonical polarization treatment. The singularity warns that the massless limit requires separate control of the longitudinal mode.

The same polarization tensor appears in the free Proca propagator,

DμνP(p)=−ip2−M2+iϵ(ημν−pμpνM2).D^{\mathrm P}_{\mu\nu}(p) ={-i\over p^2-M^2+i\epsilon} \left(\eta_{\mu\nu}-{p_\mu p_\nu\over M^2}\right).

Unlike the photon propagator below, this is not part of a one-parameter family of gauges: the Proca kinetic operator is invertible for M≠0M\neq0. The pμpν/M2p_\mu p_\nu/M^2 term records a physical longitudinal polarization rather than a redundant gauge direction.

Massive and massless vector degrees of freedom

A massive vector has four components minus the Proca constraint p⋅ϵ=0p\cdot\epsilon=0, leaving three polarizations. A massless vector has the on-shell condition p2=0p^2=0, transversality, and the gauge equivalence ϵμ∼ϵμ+βpμ\epsilon^\mu\sim\epsilon^\mu+\beta p^\mu, leaving two helicities.

The longitudinal mode and conserved currents

Section titled “The longitudinal mode and conserved currents”

The apparent singularity in the massive polarization sum is associated with the longitudinal polarization. For momentum pμ=(E,0,0,∣p∣)p^\mu=(E,0,0,|\mathbf p|), one convenient longitudinal massive polarization is

ϵLμ(p)=1M(∣p∣,0,0,E).\epsilon_L^\mu(p)={1\over M}(|\mathbf p|,0,0,E).

It obeys

p⋅ϵL=E∣p∣M−∣p∣EM=0,p\cdot\epsilon_L=E{|\mathbf p|\over M}-|\mathbf p|{E\over M}=0,

and

ϵL2=−1.\epsilon_L^2=-1.

Take M→0M\to0 at fixed ∣p∣>0|\mathbf p|>0. Then E∼∣p∣E\sim|\mathbf p|, so

ϵLμ(p)=pμM+O(M).\epsilon_L^\mu(p)={p^\mu\over M}+O(M).

If the vector field couples to a current matrix element JμJ_\mu whose components stay bounded in this limit, the longitudinal contribution contains

ϵLμJμ=1MpμJμ+O(M).\epsilon_L^\mu J_\mu ={1\over M}p^\mu J_\mu+O(M).

For a conserved current,

pμJμ=0.p_\mu J^\mu=0.

Therefore the dangerous 1/M1/M part vanishes. More explicitly, conservation gives EJ0=∣p∣J3EJ^0=|\mathbf p|J^3, hence ϵL⋅J=−MJ3/E\epsilon_L\cdot J=-MJ^3/E. This vanishes for the bounded current at fixed nonzero momentum. Conservation alone does not bound a family of current matrix elements as a parameter changes; the stated regularity is essential to this decoupling argument.

This statement is often the cleanest way to remember what gauge invariance is doing. The massless photon has no longitudinal physical polarization. Coupling it consistently to matter requires amplitudes to be insensitive to adding a multiple of pμp^\mu to the polarization vector. Current conservation is the corresponding condition on the matter side.

Set the vector mass to zero and keep only the field-strength term:

LMaxwell=−14FμνFμν.\mathcal L_{\mathrm{Maxwell}} =-{1\over4}F_{\mu\nu}F^{\mu\nu}.

For redundancy statements, work on topologically trivial Minkowski space and restrict to proper gauge transformations, for example smooth gauge functions of compact support, or boundary falloff that gives zero boundary charge. Transformations with nontrivial boundary charges require a separate physical-symmetry analysis. The canonical Maxwell treatment explains this boundary distinction.

This Lagrangian is invariant under

Aμ↦Aμ+∂μα.A_\mu\mapsto A_\mu+\partial_\mu\alpha.

Indeed,

Fμν↦∂μ(Aν+∂να)−∂ν(Aμ+∂μα)=Fμν,F_{\mu\nu}\mapsto \partial_\mu(A_\nu+\partial_\nu\alpha) -\partial_\nu(A_\mu+\partial_\mu\alpha) =F_{\mu\nu},

because mixed partial derivatives commute. The field strength, not the potential itself, is gauge invariant.

Denote the Maxwell source current by JsourceμJ_{\mathrm{source}}^\mu and choose the source coupling

Lsource=−JsourceμAμ,\mathcal L_{\mathrm{source}}=-J_{\mathrm{source}}^\mu A_\mu,

the Euler–Lagrange equation is

∂μFμν=Jsourceν.\partial_\mu F^{\mu\nu}=J_{\mathrm{source}}^\nu.

The divergence of the left side vanishes identically:

∂ν∂μFμν=0.\partial_\nu\partial_\mu F^{\mu\nu}=0.

Therefore consistency requires

∂νJsourceν=0.\partial_\nu J_{\mathrm{source}}^\nu=0.

The same condition follows from gauge invariance of the coupling to matter. The source action is

Sint=−∫d4x JsourceμAμ.S_{\mathrm{int}}=-\int d^4x\,J_{\mathrm{source}}^\mu A_\mu.

Under Aμ↦Aμ+∂μαA_\mu\mapsto A_\mu+\partial_\mu\alpha,

δSint=−∫d4x Jsourceμ∂μα=∫d4x α ∂μJsourceμ,\delta S_{\mathrm{int}} =-\int d^4x\,J_{\mathrm{source}}^\mu\partial_\mu\alpha =\int d^4x\,\alpha\,\partial_\mu J_{\mathrm{source}}^\mu,

up to a boundary term. For arbitrary α(x)\alpha(x), this vanishes only if

∂μJsourceμ=0.\partial_\mu J_{\mathrm{source}}^\mu=0.

The logic can be read in either direction. If a theory has a conserved current, it can couple naturally to a massless vector field. If a massless vector field couples consistently to matter, gauge redundancy forces the current seen by the vector field to be conserved.

In vacuum, a plane wave obeys p2=0p^2=0 and may be represented by a polarization satisfying p⋅ϵ=0p\cdot\epsilon=0. The latter condition still leaves three components because a null vector is orthogonal to itself. The residual equivalence

ϵμ∼ϵμ+βpμ\epsilon_\mu\sim\epsilon_\mu+\beta p_\mu

removes one more component, leaving two helicities. This is different from Proca theory twice over: in Maxwell theory ∂⋅A=0\partial\cdot A=0 is imposed as a convenient Lorenz gauge condition rather than obtained as a Proca constraint, and the remaining longitudinal shift is a redundancy rather than a physical mode.

Why naive Lorentz-covariant quantization fails

Section titled “Why naive Lorentz-covariant quantization fails”

One might try to quantize the four components of AμA_\mu as if they were four scalar fields. This immediately produces trouble. A Lorentz-covariant oscillator algebra has the schematic form

[aμ(p),aν†(q)]=−(2π)32Ep ημνδ(3)(p−q).[a_\mu(\mathbf p),a_\nu^\dagger(\mathbf q)] =-(2\pi)^3 2E_{\mathbf p}\,\eta_{\mu\nu}\delta^{(3)}(\mathbf p-\mathbf q).

The spatial components then have positive norm because −ηij=δij-\eta_{ij}=\delta_{ij}, but the time component has

[a0(p),a0†(q)]=−(2π)32Epδ(3)(p−q).[a_0(\mathbf p),a_0^\dagger(\mathbf q)] =-(2\pi)^3 2E_{\mathbf p}\delta^{(3)}(\mathbf p-\mathbf q).

Thus the one-particle state a0†∣0⟩a_0^\dagger|0\rangle has negative norm:

⟨0∣a0a0†∣0⟩<0.\langle0|a_0a_0^\dagger|0\rangle<0.

A Hilbert space with negative-norm physical states is unacceptable. Gauge theory avoids this conclusion by telling us that not all components of AμA_\mu create physical states. The timelike and longitudinal modes are artifacts of the Lorentz-covariant description.

This is not just a minor bookkeeping problem. The Maxwell quadratic action can be written, after integrating by parts, as

S0=12∫d4x Aμ(ημν∂2−∂μ∂ν)Aν.S_0={1\over2}\int d^4x\, A_\mu\left(\eta^{\mu\nu}\partial^2-\partial^\mu\partial^\nu\right)A_\nu.

In momentum space the kinetic operator is proportional to

Kμν(p)=−p2ημν+pμpν,K^{\mu\nu}(p)=-p^2\eta^{\mu\nu}+p^\mu p^\nu,

where the sign follows from the course convention ∂μ↦−ipμ\partial_\mu\mapsto-ip_\mu. It has a zero mode:

Kμν(p)pν=0.K^{\mu\nu}(p)p_\nu=0.

The zero mode is exactly the gauge direction Aμ(p)∼Aμ(p)−ipμα(p)A_\mu(p)\sim A_\mu(p)-ip_\mu\alpha(p), obtained by Fourier transforming Aμ↦Aμ+∂μαA_\mu\mapsto A_\mu+\partial_\mu\alpha. Therefore the kinetic operator is not invertible until we fix a gauge or otherwise remove the redundant directions.

Gauge orbit, gauge slice, and physical configurations

The figure is a local schematic for the admitted proper gauge transformations. A gauge condition selects representatives, with boundary data and any residual transformations still to be fixed; it is not a theorem that an arbitrary gauge condition intersects every orbit exactly once. Potentials along such a proper orbit describe the same physical configuration.

To calculate with a photon propagator, we need an inverse kinetic operator with specified boundary data. A Lorenz condition still admits residual transformations satisfying □α=0\Box\alpha=0; the vacuum prescription and remaining boundary restrictions must handle them. One standard Lorentz-covariant choice adds the gauge-fixing term

Lgf=−12ξ(∂μAμ)2.\mathcal L_{\mathrm{gf}} =-{1\over2\xi}(\partial_\mu A^\mu)^2.

For ξ≠0\xi\neq0 the quadratic operator is

Kμν(ξ)(p)=−p2ημν+(1−1ξ)pμpν,K^{(\xi)}_{\mu\nu}(p) =-p^2\eta_{\mu\nu} +\left(1-{1\over\xi}\right)p_\mu p_\nu,

with the overall sign fixed by L=−F2/4−(∂⋅A)2/(2ξ)\mathcal L=-F^2/4-(\partial\cdot A)^2/(2\xi). Its inverse gives the momentum-space photon propagator

Dμν(p)=−ip2+iϵ[ημν−(1−ξ)pμpνp2].D_{\mu\nu}(p) ={-i\over p^2+i\epsilon} \left[ \eta_{\mu\nu}-(1-\xi){p_\mu p_\nu\over p^2} \right].

The pole prescription in the longitudinal term is understood as the corresponding limit of the gauge-fixed inverse; writing a second p2+iϵp^2+i\epsilon there is an equivalent common shorthand. In Feynman gauge, ξ=1\xi=1, this simplifies to

Dμν(p)=−iημνp2+iϵ.D_{\mu\nu}(p)={-i\eta_{\mu\nu}\over p^2+i\epsilon}.

This expression propagates four components, but only gauge-invariant combinations are observable. Between conserved currents, the gauge-dependent term is proportional to

(p⋅J1)(p⋅J2)=0,(p\cdot J_1)(p\cdot J_2)=0,

so the exchange amplitude is independent of ξ\xi. In Abelian gauge theory the Faddeev–Popov determinant for a linear covariant gauge is field independent, so its ghosts decouple. In non-Abelian gauge theory the determinant depends on the gauge field, and ghost loops are essential. For the present QFT I discussion, the main point is simpler: the propagator exists only after the redundant gauge directions are handled.

Different gauges emphasize different virtues. Lorenz-type gauges preserve manifest Lorentz covariance. Coulomb gauge, ∇⋅A=0\nabla\cdot\mathbf A=0, displays the two transverse photon polarizations more directly but hides manifest covariance. Temporal gauge, A0=0A_0=0, is sometimes intuitive but leaves residual gauge freedom. Physical answers must be independent of this choice.

The Maxwell kinetic operator has longitudinal zero modes

The Maxwell kinetic operator annihilates longitudinal modes proportional to pμp_\mu. This is why the free Maxwell quadratic form cannot be inverted before gauge fixing. Gauge fixing lifts the degeneracy of the quadratic form without changing gauge-invariant observables.

Local phase symmetry and the covariant derivative

Section titled “Local phase symmetry and the covariant derivative”

Now suppose a matter theory has a global U(1)U(1) symmetry. For a Dirac field,

L0=ψˉ(iγμ∂μ−m)ψ\mathcal L_0=\bar\psi(i\gamma^\mu\partial_\mu-m)\psi

is invariant under

ψ↦eiqαψ,ψˉ↦ψˉe−iqα,\psi\mapsto e^{iq\alpha}\psi, \qquad \bar\psi\mapsto \bar\psi e^{-iq\alpha},

when α\alpha is constant. If α\alpha depends on spacetime, then

∂μψ↦∂μ(eiqαψ)=eiqα(∂μψ+iq(∂μα)ψ).\partial_\mu\psi\mapsto \partial_\mu(e^{iq\alpha}\psi) =e^{iq\alpha}\left(\partial_\mu\psi+iq(\partial_\mu\alpha)\psi\right).

The extra term prevents ∂μψ\partial_\mu\psi from transforming like ψ\psi. Introduce a vector field AμA_\mu and define

Dμψ=(∂μ−iqAμ)ψ.D_\mu\psi=(\partial_\mu-iqA_\mu)\psi.

If

Aμ↦Aμ+∂μα,A_\mu\mapsto A_\mu+\partial_\mu\alpha,

then

Dμψ↦eiqαDμψ.D_\mu\psi\mapsto e^{iq\alpha}D_\mu\psi.

Thus a locally invariant Dirac Lagrangian is

L=ψˉ(iγμDμ−m)ψ−14FμνFμν.\mathcal L=\bar\psi(i\gamma^\mu D_\mu-m)\psi -{1\over4}F_{\mu\nu}F^{\mu\nu}.

Expanding the covariant derivative gives

ψˉiγμDμψ=ψˉiγμ∂μψ+qψˉγμψAμ.\bar\psi i\gamma^\mu D_\mu\psi =\bar\psi i\gamma^\mu\partial_\mu\psi +q\bar\psi\gamma^\mu\psi A_\mu.

Begin with the conserved Dirac bilinear

jbilμ=ψˉγμψ.j_{\mathrm{bil}}^\mu=\bar\psi\gamma^\mu\psi.

The interaction and source conventions define two related charge-weighted currents. The covariant derivative above gives

Lint=qAμjbilμ=Aμjqμ,jqμ=qjbilμ.\mathcal L_{\mathrm{int}} =qA_\mu j_{\mathrm{bil}}^\mu =A_\mu j_q^\mu, \qquad j_q^\mu=qj_{\mathrm{bil}}^\mu.

Relative to the earlier convention Lsource=−Jsource ⁣⋅A\mathcal L_{\mathrm{source}}=-J_{\mathrm{source}}\!\cdot A, the current on the right-hand side of Maxwell’s equation is therefore

Jsourceμ=−jqμ=−qjbilμ.J_{\mathrm{source}}^\mu=-j_q^\mu =-qj_{\mathrm{bil}}^\mu.

Thus a positive qq gives the Lorentzian interaction +qAμjbilμ+qA_\mu j_{\mathrm{bil}}^\mu. For an electron, q=−eq=-e, so the interaction is −eAμjbilμ-eA_\mu j_{\mathrm{bil}}^\mu and the vertex is −ieγμ-ie\gamma^\mu. Keeping the labels distinct prevents the sign in the vertex coefficient from being mistaken for the sign in the convention −Jsource ⁣⋅A-J_{\mathrm{source}}\!\cdot A.

The scalar version is similar. For a complex scalar field of charge qq,

L=(Dμϕ)∗(Dμϕ)−m2ϕ∗ϕ−14FμνFμν.\mathcal L=(D_\mu\phi)^*(D^\mu\phi)-m^2\phi^*\phi -{1\over4}F_{\mu\nu}F^{\mu\nu}.

Here the star matters:

(Dμϕ)∗=(∂μ+iqAμ)ϕ∗.(D_\mu\phi)^* =(\partial_\mu+iqA_\mu)\phi^*.

Thus the conjugate field carries the opposite charge. A common alternative is to absorb the charge into the gauge potential. Define

Aμ=qAμ,Dμ=∂μ−iAμ.\mathcal A_\mu=qA_\mu, \qquad \mathcal D_\mu=\partial_\mu-i\mathcal A_\mu.

Then

(Dμϕ)∗(Dμϕ)=(∂μ+iAμ)ϕ∗(∂μ−iAμ)ϕ.(\mathcal D_\mu\phi)^*(\mathcal D^\mu\phi) =(\partial_\mu+i\mathcal A_\mu)\phi^* (\partial^\mu-i\mathcal A^\mu)\phi.

If the kinetic term for the rescaled gauge field is written with the coupling outside, the Lagrangian becomes schematically

L=(Dμϕ)∗(Dμϕ)−m2∣ϕ∣2−14q2FμνFμν.\mathcal L=(\mathcal D_\mu\phi)^*(\mathcal D^\mu\phi) -m^2|\phi|^2-{1\over4q^2}\mathcal F_{\mu\nu}\mathcal F^{\mu\nu}.

The placement of the coupling is a convention; the opposite signs on the two conjugate scalar factors are not. They are required because ϕ∗\phi^* has charge −q-q when ϕ\phi has charge qq.

Covariant derivative and field strength as connection and curvature

A local phase rotation makes ordinary derivatives transform inhomogeneously. The gauge field AμA_\mu supplies the connection term in DμD_\mu. The field strength FμνF_{\mu\nu} is the curvature: it measures the phase accumulated around an infinitesimal loop.

The field strength can be characterized without guessing its formula. For the Abelian covariant derivative

Dμ=∂μ−iqAμ,D_\mu=\partial_\mu-iqA_\mu,

the commutator is

[Dμ,Dν]=−iq(∂μAν−∂νAμ)=−iqFμν.[D_\mu,D_\nu] =-iq(\partial_\mu A_\nu-\partial_\nu A_\mu) =-iqF_{\mu\nu}.

Thus FμνF_{\mu\nu} is the obstruction to commuting two covariant derivatives. Geometrically, AμA_\mu is a connection and FμνF_{\mu\nu} is its curvature.

This viewpoint generalizes immediately, but the sign bookkeeping is slightly less forgiving. With Hermitian generators and Dμ=∂μ−igAμaTaD_\mu=\partial_\mu-igA_\mu^aT^a, the non-Abelian field strength contains a plus sign in components, Fμνa=∂μAνa−∂νAμa+gfabcAμbAνcF_{\mu\nu}^a=\partial_\mu A_\nu^a-\partial_\nu A_\mu^a+gf^{abc}A_\mu^bA_\nu^c. In matrix notation this same statement is cleaner. Let U(x)=eigαa(x)TaU(x)=e^{ig\alpha^a(x)T^a} and let ψ\psi transform in a representation of a non-Abelian group GG:

ψ(x)↦U(x)ψ(x),U(x)∈G.\psi(x)\mapsto U(x)\psi(x), \qquad U(x)\in G.

Write the gauge field as a matrix

Aμ=AμaTa,[Ta,Tb]=ifabcTc,A_\mu=A_\mu^aT^a, \qquad [T^a,T^b]=if^{abc}T^c,

and define

Dμ=∂μ−igAμ.D_\mu=\partial_\mu-igA_\mu.

The gauge field must transform so that

Dμψ↦U(Dμψ).D_\mu\psi\mapsto U(D_\mu\psi).

Equivalently,

Dμ↦UDμU−1.D_\mu\mapsto UD_\mu U^{-1}.

This gives

Aμ↦UAμU−1−ig(∂μU)U−1.A_\mu\mapsto U A_\mu U^{-1}-{i\over g}(\partial_\mu U)U^{-1}.

The field strength is defined by

[Dμ,Dν]=−igFμν,[D_\mu,D_\nu]=-igF_{\mu\nu},

so

Fμν=∂μAν−∂νAμ−ig[Aμ,Aν].F_{\mu\nu} =\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].

The commutator term is the new feature of Yang–Mills theory. Since [Ta,Tb]=ifabcTc[T^a,T^b]=if^{abc}T^c, the displayed matrix formula gives

Fμνa=∂μAνa−∂νAμa+gfabcAμbAνc.F_{\mu\nu}^a =\partial_\mu A_\nu^a-\partial_\nu A_\mu^a+gf^{abc}A_\mu^bA_\nu^c.

References that instead use Dμ=∂μ+igAμD_\mu=\partial_\mu+igA_\mu have the opposite sign in the commutator term. The defining equation [Dμ,Dν]=−igFμν[D_\mu,D_\nu]=-igF_{\mu\nu} fixes the signs used here.

Under a gauge transformation,

Fμν↦UFμνU−1.F_{\mu\nu}\mapsto UF_{\mu\nu}U^{-1}.

Therefore tr⁡(FμνFμν)\operatorname{tr}(F_{\mu\nu}F^{\mu\nu}) is gauge invariant, and the Yang–Mills action is

LYM=−12tr⁡(FμνFμν).\mathcal L_{\mathrm{YM}} =-{1\over2}\operatorname{tr}(F_{\mu\nu}F^{\mu\nu}).

Here the trace is normalized by tr⁡(TaTb)=12δab\operatorname{tr}(T^aT^b)=\tfrac12\delta^{ab}. For a trace in a representation with tr⁡R(TRaTRb)=T(R)δab\operatorname{tr}_R(T_R^aT_R^b)=T(R)\delta^{ab}, the coefficient is −1/[4T(R)]-1/[4T(R)].

For U(1)U(1) the commutator vanishes and this reduces to Maxwell theory. For a non-Abelian group, the gauge bosons carry the gauge charge themselves, and the field strength contains cubic and quartic self-interactions after the action is expanded.

Ward identities from replacing a polarization by momentum

Section titled “Ward identities from replacing a polarization by momentum”

For an external photon with momentum qq, a gauge transformation shifts a plane-wave polarization by

ϵμ(q)⟼ϵμ(q)+βqμ.\epsilon_\mu(q)\longmapsto \epsilon_\mu(q)+\beta q_\mu.

A physical scattering amplitude cannot change under this replacement. If the amplitude with the external photon removed is written as Mμ\mathcal M^\mu, then

M=ϵμMμ.\mathcal M=\epsilon_\mu\mathcal M^\mu.

Gauge invariance requires

qμMμ=0.q_\mu\mathcal M^\mu=0.

This is the simplest form of a Ward identity.

For example, the scalar QED three-point vertex for a scalar of charge QQ is proportional to

+iQ(p1+p2)μ,+iQ(p_1+p_2)_\mu,

where the incoming photon momentum is k=p2−p1k=p_2-p_1 with both scalar legs on shell. Replacing the photon polarization by kμk_\mu gives the factor

k⋅(p1+p2)=(p2−p1)⋅(p1+p2)=p22−p12.k\cdot(p_1+p_2) =(p_2-p_1)\cdot(p_1+p_2) =p_2^2-p_1^2.

If the scalar particles have equal mass and are on shell,

p12=p22=m2,p_1^2=p_2^2=m^2,

so

k⋅(p1+p2)=0.k\cdot(p_1+p_2)=0.

This small calculation contains the whole moral: the longitudinal piece of a photon polarization does not contribute to a physical amplitude. In the next page, this becomes a practical rule for QED vertices and tree amplitudes.

Gauge redundancy is not an ordinary global symmetry

Section titled “Gauge redundancy is not an ordinary global symmetry”

It is tempting to speak of gauge invariance as a symmetry, and this language is standard. But one should remember that it is a special kind of symmetry. A global symmetry can act on physical states. Proper gauge redundancy maps one description of a physical configuration to another description of the same physical configuration; a boundary transformation carrying a nonzero charge is not included in this identification.

The difference is visible already in Maxwell theory. The field strength FμνF_{\mu\nu} is unchanged by

Aμ↦Aμ+∂μα.A_\mu\mapsto A_\mu+\partial_\mu\alpha.

Thus AμA_\mu and Aμ+∂μαA_\mu+\partial_\mu\alpha cannot represent distinct measured electromagnetic fields. The physical observables must be gauge invariant, such as FμνF_{\mu\nu}, Wilson loops, scattering amplitudes, or properly dressed charged operators.

This does not make gauge theory empty. Quite the opposite: the redundancy imposes strong constraints. It forbids a photon mass term M2AμAμ/2M^2A_\mu A^\mu/2, requires matter to enter through covariant derivatives, enforces current conservation, constrains counterterms, and leads to Ward identities. In non-Abelian gauge theories, the same principle generates the self-interactions of gluons and underlies the structure of the Standard Model.

A massive vector field is described by the Proca equation. Its divergence gives the dynamical constraint ∂μAμ=0\partial_\mu A^\mu=0, so a four-component Lorentz vector carries three physical spin-one polarizations. At fixed nonzero momentum, a conserved current matrix element that remains bounded as M→0M\to0 decouples from the longitudinal polarization.

A massless vector field is described by Maxwell theory with the gauge redundancy Aμ∼Aμ+∂μαA_\mu\sim A_\mu+\partial_\mu\alpha. The field strength FμνF_{\mu\nu} is invariant, and the free Maxwell kinetic operator has zero modes precisely along gauge directions. Gauge fixing is required to define a propagator, but physical quantities must be independent of the gauge choice.

Local phase invariance replaces ordinary derivatives by covariant derivatives. The field strength is the commutator of covariant derivatives. For Abelian gauge theory this gives Fμν=∂μAν−∂νAμF_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu; for non-Abelian gauge theory it adds a commutator term and hence gauge-boson self-interactions. The first practical consequence is the Ward identity: amplitudes vanish when an external photon polarization is replaced by its momentum.

Components are not polarizations. The covariant field AμA_\mu has four components, but a photon has two physical helicities. Transversality and the gauge equivalence must both be used in the massless count.

The Proca constraint is not a gauge condition. For M≠0M\neq0, ∂⋅A=0\partial\cdot A=0 follows from the equations of motion and the longitudinal polarization is physical. In Maxwell theory a Lorenz condition selects representatives of gauge orbits and does not create a third photon polarization.

Proper gauge transformations are redundancies. They identify descriptions of the same physical configuration. Transformations carrying boundary charges can instead act as physical symmetries; equality of the local field strength alone does not decide their status.

The Maxwell operator has no inverse before gauge fixing. Its zero modes are the gauge directions, not an algebraic accident. A propagator becomes well defined only after restricting or lifting those directions.

A bare vector mass breaks Maxwell gauge invariance. The term M2AμAμ/2M^2A_\mu A^\mu/2 is not invariant by itself. Gauge-invariant massive-vector theories require additional structure, such as a Stueckelberg or Higgs field.

A sign convention is a package. With Dμ=∂μ−iqAμD_\mu=\partial_\mu-iqA_\mu and ψ↦eiqαψ\psi\mapsto e^{iq\alpha}\psi, the transformation of AμA_\mu, the matter vertex, and the conjugate derivative are fixed. In particular, (Dμϕ)∗=(∂μ+iqAμ)ϕ∗(D_\mu\phi)^*=(\partial_\mu+iqA_\mu)\phi^*.

Derive the Proca constraint. Starting from

∂μFμν+M2Aν=0,\partial_\mu F^{\mu\nu}+M^2A^\nu=0,

show that ∂μAμ=0\partial_\mu A^\mu=0 for M≠0M\neq0, and then show that each component of AμA^\mu obeys the Klein–Gordon equation.

Solution

Take ∂ν\partial_\nu of the Proca equation:

∂ν∂μFμν+M2∂νAν=0.\partial_\nu\partial_\mu F^{\mu\nu}+M^2\partial_\nu A^\nu=0.

Since Fμν=−FνμF^{\mu\nu}=-F^{\nu\mu} and ∂ν∂μ\partial_\nu\partial_\mu is symmetric in μ,ν\mu,\nu,

∂ν∂μFμν=0.\partial_\nu\partial_\mu F^{\mu\nu}=0.

Therefore

M2∂νAν=0.M^2\partial_\nu A^\nu=0.

For M≠0M\neq0,

∂νAν=0.\partial_\nu A^\nu=0.

Now expand

∂μFμν=∂μ(∂μAν−∂νAμ)=∂2Aν−∂ν(∂μAμ).\partial_\mu F^{\mu\nu} =\partial_\mu(\partial^\mu A^\nu-\partial^\nu A^\mu) =\partial^2A^\nu-\partial^\nu(\partial_\mu A^\mu).

Using ∂μAμ=0\partial_\mu A^\mu=0, the equation becomes

(∂2+M2)Aν=0.(\partial^2+M^2)A^\nu=0.

Thus the Proca field is a massive Klein–Gordon field in each component, supplemented by the transversality constraint that removes one component.

Show explicitly that the longitudinal Proca polarization decouples from a conserved current in the massless limit at fixed ∣p∣>0|\mathbf p|>0, assuming the current components remain bounded. Use

ϵLμ(p)=1M(∣p∣,0,0,E),pμ=(E,0,0,∣p∣),\epsilon_L^\mu(p)={1\over M}(|\mathbf p|,0,0,E), \qquad p^\mu=(E,0,0,|\mathbf p|),

with E2−p2=M2E^2-\mathbf p^2=M^2.

Solution

First check transversality:

p⋅ϵL=E∣p∣M−∣p∣EM=0.p\cdot\epsilon_L =E{|\mathbf p|\over M}-|\mathbf p|{E\over M}=0.

The norm is

ϵL2=∣p∣2−E2M2=−1.\epsilon_L^2={|\mathbf p|^2-E^2\over M^2}=-1.

For small MM,

E=∣p∣+M22∣p∣+O(M4),E=|\mathbf p|+{M^2\over2|\mathbf p|}+O(M^4),

so

ϵLμ(p)=pμM+O(M).\epsilon_L^\mu(p)={p^\mu\over M}+O(M).

The coupling to a current is

ϵLμJμ=1MpμJμ+O(M).\epsilon_L^\mu J_\mu={1\over M}p^\mu J_\mu+O(M).

If the current is conserved, then in momentum space

pμJμ=0.p_\mu J^\mu=0.

The singular piece vanishes. In fact conservation gives J0=∣p∣J3/EJ^0=|\mathbf p|J^3/E, so the exact contraction is ϵL⋅J=−MJ3/E\epsilon_L\cdot J=-MJ^3/E. Bounded J3J^3 therefore makes it vanish. As a negative control, Jμ=ϵLμ/MJ^\mu=\epsilon_L^\mu/M is also transverse, but ϵL⋅J=−1/M\epsilon_L\cdot J=-1/M diverges. This family violates the bounded-current hypothesis and shows why conservation alone is insufficient.

Let

Kμν(p)=−p2ημν+pμpν.K^{\mu\nu}(p)=-p^2\eta^{\mu\nu}+p^\mu p^\nu.

Show that Kμνpν=0K^{\mu\nu}p_\nu=0. Then add Lgf=−(∂⋅A)2/(2ξ)\mathcal L_{\mathrm{gf}}=-(\partial\cdot A)^2/(2\xi), invert the resulting operator, and explain why its ξ\xi-dependent part vanishes between conserved currents.

Solution

Compute

Kμνpν=−p2ημνpν+pμpνpν.K^{\mu\nu}p_\nu =-p^2\eta^{\mu\nu}p_\nu+p^\mu p^\nu p_\nu.

Since ημνpν=pμ\eta^{\mu\nu}p_\nu=p^\mu and pνpν=p2p^\nu p_\nu=p^2,

Kμνpν=−p2pμ+pμp2=0.K^{\mu\nu}p_\nu =-p^2p^\mu+p^\mu p^2=0.

Thus pνp_\nu is a null eigenvector of the kinetic operator. A matrix or differential operator with a null eigenvector has no inverse on the full vector space. Since a propagator is the inverse of the quadratic operator, the Maxwell propagator is undefined until one removes or lifts the gauge-zero directions by a gauge condition.

Introduce the transverse and longitudinal projectors

PμνT=ημν−pμpνp2,PμνL=pμpνp2.P^{\mathrm T}_{\mu\nu}=\eta_{\mu\nu}-{p_\mu p_\nu\over p^2}, \qquad P^{\mathrm L}_{\mu\nu}={p_\mu p_\nu\over p^2}.

The gauge-fixed operator is proportional to

−p2(PT+1ξPL),-p^2\left(P^{\mathrm T}+{1\over\xi}P^{\mathrm L}\right),

so its Feynman inverse is

Dμν(p)=−ip2+iϵ(PμνT+ξPμνL)=−ip2+iϵ[ημν−(1−ξ)pμpνp2].D_{\mu\nu}(p) ={-i\over p^2+i\epsilon} \left(P^{\mathrm T}_{\mu\nu}+\xi P^{\mathrm L}_{\mu\nu}\right) ={-i\over p^2+i\epsilon} \left[\eta_{\mu\nu}-(1-\xi){p_\mu p_\nu\over p^2}\right].

Contracting the last term with two currents gives a factor (p⋅J1)(p⋅J2)(p\cdot J_1)(p\cdot J_2). It vanishes when both currents are conserved, proving that the exchange amplitude is independent of ξ\xi.

For scalar QED, the one-photon vertex for an incoming scalar momentum p1p_1 and outgoing scalar momentum p2p_2 is proportional to

(p1+p2)μ.(p_1+p_2)_\mu.

Let the photon momentum be q=p1−p2q=p_1-p_2. Show that replacing the photon polarization by qμq_\mu gives zero for equal-mass on-shell scalar particles.

Solution

The replacement gives the contraction

q⋅(p1+p2).q\cdot(p_1+p_2).

Using q=p1−p2q=p_1-p_2,

q⋅(p1+p2)=(p1−p2)⋅(p1+p2)=p12−p22.q\cdot(p_1+p_2) =(p_1-p_2)\cdot(p_1+p_2) =p_1^2-p_2^2.

For equal-mass on-shell scalar particles,

p12=p22=m2.p_1^2=p_2^2=m^2.

Therefore

q⋅(p1+p2)=0.q\cdot(p_1+p_2)=0.

This is the tree-level Ward identity for the scalar three-point vertex.

With

Dμ=∂μ−igAμ,Aμ=AμaTa,D_\mu=\partial_\mu-igA_\mu, \qquad A_\mu=A_\mu^aT^a,

define FμνF_{\mu\nu} by

[Dμ,Dν]=−igFμν.[D_\mu,D_\nu]=-igF_{\mu\nu}.

Derive

Fμν=∂μAν−∂νAμ−ig[Aμ,Aν].F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].
Solution

Let DμD_\mu act on a test field ψ\psi. Then

DμDνψ=(∂μ−igAμ)(∂νψ−igAνψ).D_\mu D_\nu\psi =(\partial_\mu-igA_\mu)(\partial_\nu\psi-igA_\nu\psi).

Expanding,

DμDνψ=∂μ∂νψ−ig(∂μAν)ψ−igAν∂μψ−igAμ∂νψ−g2AμAνψ.D_\mu D_\nu\psi =\partial_\mu\partial_\nu\psi -ig(\partial_\mu A_\nu)\psi -igA_\nu\partial_\mu\psi -igA_\mu\partial_\nu\psi -g^2A_\mu A_\nu\psi.

Subtract the same expression with μ\mu and ν\nu exchanged. The second-derivative terms cancel, and so do the terms where a gauge field multiplies a derivative of ψ\psi. The result is

[Dμ,Dν]ψ=−ig(∂μAν−∂νAμ)ψ−g2(AμAν−AνAμ)ψ.[D_\mu,D_\nu]\psi =-ig(\partial_\mu A_\nu-\partial_\nu A_\mu)\psi -g^2(A_\mu A_\nu-A_\nu A_\mu)\psi.

This can be written as

[Dμ,Dν]ψ=−ig(∂μAν−∂νAμ−ig[Aμ,Aν])ψ.[D_\mu,D_\nu]\psi =-ig\left(\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu]\right)\psi.

Therefore

Fμν=∂μAν−∂νAμ−ig[Aμ,Aν].F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].

For an Abelian group the commutator vanishes, recovering the Maxwell field strength.

  • Coleman, Sidney. Lectures of Sidney Coleman on Quantum Field Theory. Edited by Bryan Gin-ge Chen, David Derbes, David Griffiths, Brian Hill, Richard Sohn, and Yuan-Sen Ting. World Scientific, 2019. Chapters 26–27 discuss massive vectors, gauge invariance, and gauge-field quantization.
  • Polyakov, A. M. Gauge Fields and Strings. Harwood Academic Publishers, 1987. Chapter 1 develops the gauge-field viewpoint from long-distance physics.
  • Srednicki, Mark. Quantum Field Theory. Cambridge University Press, 2007. The chapters on spin-one fields and electrodynamics give a compact route from the Maxwell operator to gauge-fixed propagators and Ward identities.
  • Weinberg, Steven. The Quantum Theory of Fields, Volume I: Foundations. Cambridge University Press, 1995. Sections 5.3 and 8.1–8.6 treat vector fields, constraints, gauge invariance, and QED.
  • Zee, A. Quantum Field Theory in a Nutshell. 2nd ed., Princeton University Press, 2010. Chapters III.4 and IV.5 discuss gauge redundancy, Maxwell zero modes, and non-Abelian curvature.

Original QFT.org content:CC BY 4.0, unless an item supplies different terms. Third-party material retains its own terms.