Skip to content

Vector Fields and Gauge Redundancy

Spin-zero and spin-one-half fields can be quantized without introducing an obvious redundancy in the field variables. Vector fields are different. A massive vector field has three physical polarizations, and the Proca equation enforces this by a constraint. A massless vector field, however, has only two physical helicities. If we insist on using a Lorentz four-vector Aμ(x)A_\mu(x) to describe those two helicities, we have necessarily introduced more variables than physical degrees of freedom.

This is the origin of gauge redundancy. The Maxwell field is not merely a vector field with m=0m=0. It is a vector field whose physically meaningful content is invariant under

Aμ(x)Aμ(x)+μα(x).A_\mu(x)\longmapsto A_\mu(x)+\partial_\mu\alpha(x).

The redundancy is not an optional decoration. It is what makes a local Lorentz-covariant description of massless spin one possible, and it is the reason conserved currents, covariant derivatives, Ward identities, and gauge fixing all enter QFT together.

There are three statements to keep separate throughout the page. First, AμA_\mu is a Lorentz four-vector field. Second, a photon is a massless spin-one particle with two helicities. Third, a gauge choice is a representative of a redundant description, not an observable. Many mistakes in gauge theory come from treating these three statements as if they were the same.

Required background. Scalar and vector representations and mass shells supplies the Lorentz-representation and polarization language used below. Helpful background. Dirac fields and spinors supplies the on-shell spinor identities used when matter is introduced.

Massive vector fields and the Proca constraint

Section titled “Massive vector fields and the Proca constraint”

A Lorentz vector field transforms as

Aμ(x)=ΛμνAν(x),x=Λx.A'^\mu(x')=\Lambda^\mu{}_{\nu}A^\nu(x), \qquad x'=\Lambda x.

As a field representation this has four components. A massive spin-one particle, however, has only three spin states. The Proca Lagrangian implements the required reduction:

LProca=14FμνFμν+12M2AμAμ,Fμν=μAννAμ.\mathcal L_{\mathrm{Proca}} =-{1\over4}F_{\mu\nu}F^{\mu\nu} +{1\over2}M^2A_\mu A^\mu, \qquad F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu.

Varying AνA_\nu gives

μFμν+M2Aν=0.\partial_\mu F^{\mu\nu}+M^2A^\nu=0.

Taking the divergence gives

νμFμν+M2νAν=0.\partial_\nu\partial_\mu F^{\mu\nu}+M^2\partial_\nu A^\nu=0.

The first term vanishes because FμνF^{\mu\nu} is antisymmetric while νμ\partial_\nu\partial_\mu is symmetric. Therefore, for M0M\neq0,

μAμ=0.\partial_\mu A^\mu=0.

Substituting this back into the equation of motion gives

(2+M2)Aν=0.(\partial^2+M^2)A^\nu=0.

For a plane wave

Aμ(x)=ϵμ(p)eipx,A^\mu(x)=\epsilon^\mu(p)e^{-ip\cdot x},

the equations become

p2=M2,pμϵμ(p)=0.p^2=M^2, \qquad p_\mu\epsilon^\mu(p)=0.

The second condition is the Proca transversality constraint. It removes one of the four components of ϵμ\epsilon^\mu, leaving three physical polarizations.

This condition is an equation-of-motion constraint, not a gauge choice. The mass term removes the Maxwell gauge redundancy: for M0M\neq0, two Proca configurations that differ by μα\partial_\mu\alpha are generally physically different. In Hamiltonian language A0A_0 has no independent second-order evolution equation; eliminating it leaves three propagating canonical modes.

In the rest frame pμ=(M,0,0,0)p^\mu=(M,0,0,0), the condition pϵ=0p\cdot\epsilon=0 gives

Mϵ0=0,M\epsilon^0=0,

so ϵ0=0\epsilon^0=0. The remaining three components are an ordinary spatial vector. A convenient basis is

ϵ1μ=(0,1,0,0),ϵ2μ=(0,0,1,0),ϵ3μ=(0,0,0,1).\epsilon_1^\mu=(0,1,0,0), \qquad \epsilon_2^\mu=(0,0,1,0), \qquad \epsilon_3^\mu=(0,0,0,1).

Thus the Proca field contains precisely the spin-one representation of the massive little group SO(3)SO(3).

The massive polarization sum is

λ=13ϵλμ(p)ϵλν(p)=ημν+pμpνM2,\sum_{\lambda=1}^{3} \epsilon_\lambda^\mu(p)\epsilon_\lambda^{*\nu}(p) =-\eta^{\mu\nu}+{p^\mu p^\nu\over M^2},

where the physical polarizations are normalized by

ϵλ(p)ϵλ(p)=δλλ.\epsilon_\lambda(p)\cdot\epsilon_{\lambda'}(p)=-\delta_{\lambda\lambda'}.

The projector on the right is singular as M0M\to0. This singularity is the first warning that the massless limit of a vector field is not just a smooth limit of the massive theory unless the longitudinal mode decouples.

The same projector appears in the free Proca propagator,

DμνP(p)=ip2M2+iϵ(ημνpμpνM2).D^{\mathrm P}_{\mu\nu}(p) ={-i\over p^2-M^2+i\epsilon} \left(\eta_{\mu\nu}-{p_\mu p_\nu\over M^2}\right).

Unlike the photon propagator below, this is not part of a one-parameter family of gauges: the Proca kinetic operator is invertible for M0M\neq0. The pμpν/M2p_\mu p_\nu/M^2 term records a physical longitudinal polarization rather than a redundant gauge direction.

Massive and massless vector degrees of freedom

A massive vector has four components minus the Proca constraint pϵ=0p\cdot\epsilon=0, leaving three polarizations. A massless vector has the on-shell condition p2=0p^2=0, transversality, and the gauge equivalence ϵμϵμ+βpμ\epsilon^\mu\sim\epsilon^\mu+\beta p^\mu, leaving two helicities.

The longitudinal mode and conserved currents

Section titled “The longitudinal mode and conserved currents”

The apparent singularity in the massive polarization sum is associated with the longitudinal polarization. For momentum pμ=(E,0,0,p)p^\mu=(E,0,0,|\mathbf p|), one convenient longitudinal massive polarization is

ϵLμ(p)=1M(p,0,0,E).\epsilon_L^\mu(p)={1\over M}(|\mathbf p|,0,0,E).

It obeys

pϵL=EpMpEM=0,p\cdot\epsilon_L=E{|\mathbf p|\over M}-|\mathbf p|{E\over M}=0,

and

ϵL2=1.\epsilon_L^2=-1.

As M0M\to0, EpE\sim|\mathbf p|, so

ϵLμ(p)=pμM+O(M).\epsilon_L^\mu(p)={p^\mu\over M}+O(M).

If the vector field couples to a current JμJ_\mu, the longitudinal contribution to an amplitude contains

ϵLμJμ=1MpμJμ+O(M).\epsilon_L^\mu J_\mu ={1\over M}p^\mu J_\mu+O(M).

For a conserved current,

pμJμ=0.p_\mu J^\mu=0.

Therefore the dangerous 1/M1/M part vanishes, and the longitudinal polarization decouples in the massless limit. This is the physical reason conserved currents are tied so tightly to massless spin-one fields.

This statement is often the cleanest way to remember what gauge invariance is doing. The massless photon has no longitudinal physical polarization. Coupling it consistently to matter requires amplitudes to be insensitive to adding a multiple of pμp^\mu to the polarization vector. Current conservation is the corresponding condition on the matter side.

Set the vector mass to zero and keep only the field-strength term:

LMaxwell=14FμνFμν.\mathcal L_{\mathrm{Maxwell}} =-{1\over4}F_{\mu\nu}F^{\mu\nu}.

This Lagrangian is invariant under

AμAμ+μα.A_\mu\mapsto A_\mu+\partial_\mu\alpha.

Indeed,

Fμνμ(Aν+να)ν(Aμ+μα)=Fμν,F_{\mu\nu}\mapsto \partial_\mu(A_\nu+\partial_\nu\alpha) -\partial_\nu(A_\mu+\partial_\mu\alpha) =F_{\mu\nu},

because mixed partial derivatives commute. The field strength, not the potential itself, is gauge invariant.

With the source coupling

Lsource=JμAμ,\mathcal L_{\mathrm{source}}=-J^\mu A_\mu,

the Euler–Lagrange equation is

μFμν=Jν.\partial_\mu F^{\mu\nu}=J^\nu.

The divergence of the left side vanishes identically:

νμFμν=0.\partial_\nu\partial_\mu F^{\mu\nu}=0.

Therefore consistency requires

νJν=0.\partial_\nu J^\nu=0.

The same condition follows from gauge invariance of the coupling to matter. The source action is

Sint=d4xJμAμ.S_{\mathrm{int}}=-\int d^4x\,J^\mu A_\mu.

Under AμAμ+μαA_\mu\mapsto A_\mu+\partial_\mu\alpha,

δSint=d4xJμμα=d4xαμJμ,\delta S_{\mathrm{int}} =-\int d^4x\,J^\mu\partial_\mu\alpha =\int d^4x\,\alpha\,\partial_\mu J^\mu,

up to a boundary term. For arbitrary α(x)\alpha(x), this vanishes only if

μJμ=0.\partial_\mu J^\mu=0.

The logic can be read in either direction. If a theory has a conserved current, it can couple naturally to a massless vector field. If a massless vector field couples consistently to matter, gauge redundancy forces the current seen by the vector field to be conserved.

In vacuum, a plane wave obeys p2=0p^2=0 and may be represented by a polarization satisfying pϵ=0p\cdot\epsilon=0. The latter condition still leaves three components because a null vector is orthogonal to itself. The residual equivalence

ϵμϵμ+βpμ\epsilon_\mu\sim\epsilon_\mu+\beta p_\mu

removes one more component, leaving two helicities. This is different from Proca theory twice over: in Maxwell theory A=0\partial\cdot A=0 is imposed as a convenient Lorenz gauge condition rather than obtained as a Proca constraint, and the remaining longitudinal shift is a redundancy rather than a physical mode.

Why naive Lorentz-covariant quantization fails

Section titled “Why naive Lorentz-covariant quantization fails”

One might try to quantize the four components of AμA_\mu as if they were four scalar fields. This immediately produces trouble. A Lorentz-covariant oscillator algebra has the schematic form

[aμ(p),aν(q)]=(2π)32Epημνδ(3)(pq).[a_\mu(\mathbf p),a_\nu^\dagger(\mathbf q)] =-(2\pi)^3 2E_{\mathbf p}\,\eta_{\mu\nu}\delta^{(3)}(\mathbf p-\mathbf q).

The spatial components then have positive norm because ηij=δij-\eta_{ij}=\delta_{ij}, but the time component has

[a0(p),a0(q)]=(2π)32Epδ(3)(pq).[a_0(\mathbf p),a_0^\dagger(\mathbf q)] =-(2\pi)^3 2E_{\mathbf p}\delta^{(3)}(\mathbf p-\mathbf q).

Thus the one-particle state a00a_0^\dagger|0\rangle has negative norm:

0a0a00<0.\langle0|a_0a_0^\dagger|0\rangle<0.

A Hilbert space with negative-norm physical states is unacceptable. Gauge theory avoids this conclusion by telling us that not all components of AμA_\mu create physical states. The timelike and longitudinal modes are artifacts of the Lorentz-covariant description.

This is not just a minor bookkeeping problem. The Maxwell quadratic action can be written, after integrating by parts, as

S0=12d4xAμ(ημν2μν)Aν.S_0={1\over2}\int d^4x\, A_\mu\left(\eta^{\mu\nu}\partial^2-\partial^\mu\partial^\nu\right)A_\nu.

In momentum space the kinetic operator is proportional to

Kμν(p)=p2ημνpμpν,K^{\mu\nu}(p)=p^2\eta^{\mu\nu}-p^\mu p^\nu,

up to the overall sign convention used for the Fourier transform. It has a zero mode:

Kμν(p)pν=0.K^{\mu\nu}(p)p_\nu=0.

The zero mode is exactly the gauge direction Aμ(p)Aμ(p)+ipμα(p)A_\mu(p)\sim A_\mu(p)+ip_\mu\alpha(p). Therefore the kinetic operator is not invertible until we fix a gauge or otherwise remove the redundant directions.

Gauge orbit, gauge slice, and physical configurations

A gauge orbit consists of many potentials AμA_\mu representing the same physical field strength. Gauge fixing chooses one representative from each orbit. In a path integral, integrating over the whole orbit overcounts the same physical configuration infinitely many times.

To calculate with a photon propagator, we need an inverse kinetic operator. One standard Lorentz-covariant choice adds the gauge-fixing term

Lgf=12ξ(μAμ)2.\mathcal L_{\mathrm{gf}} =-{1\over2\xi}(\partial_\mu A^\mu)^2.

For ξ0\xi\neq0 the quadratic operator is

Kμν(ξ)(p)=p2ημν+(11ξ)pμpν,K^{(\xi)}_{\mu\nu}(p) =-p^2\eta_{\mu\nu} +\left(1-{1\over\xi}\right)p_\mu p_\nu,

with the overall sign fixed by L=F2/4(A)2/(2ξ)\mathcal L=-F^2/4-(\partial\cdot A)^2/(2\xi). Its inverse gives the momentum-space photon propagator

Dμν(p)=ip2+iϵ[ημν(1ξ)pμpνp2].D_{\mu\nu}(p) ={-i\over p^2+i\epsilon} \left[ \eta_{\mu\nu}-(1-\xi){p_\mu p_\nu\over p^2} \right].

The pole prescription in the longitudinal term is understood as the corresponding limit of the gauge-fixed inverse; writing a second p2+iϵp^2+i\epsilon there is an equivalent common shorthand. In Feynman gauge, ξ=1\xi=1, this simplifies to

Dμν(p)=iημνp2+iϵ.D_{\mu\nu}(p)={-i\eta_{\mu\nu}\over p^2+i\epsilon}.

This expression propagates four components, but only gauge-invariant combinations are observable. Between conserved currents, the gauge-dependent term is proportional to

(pJ1)(pJ2)=0,(p\cdot J_1)(p\cdot J_2)=0,

so the exchange amplitude is independent of ξ\xi. In Abelian gauge theory the Faddeev–Popov determinant for a linear covariant gauge is field independent, so its ghosts decouple. In non-Abelian gauge theory the determinant depends on the gauge field, and ghost loops are essential. For the present QFT I discussion, the main point is simpler: the propagator exists only after the redundant gauge directions are handled.

Different gauges emphasize different virtues. Lorenz-type gauges preserve manifest Lorentz covariance. Coulomb gauge, A=0\nabla\cdot\mathbf A=0, displays the two transverse photon polarizations more directly but hides manifest covariance. Temporal gauge, A0=0A_0=0, is sometimes intuitive but leaves residual gauge freedom. Physical answers must be independent of this choice.

The Maxwell kinetic operator has longitudinal zero modes

The Maxwell kinetic operator annihilates longitudinal modes proportional to pμp_\mu. This is why the free Maxwell quadratic form cannot be inverted before gauge fixing. Gauge fixing lifts the degeneracy of the quadratic form without changing gauge-invariant observables.

Local phase symmetry and the covariant derivative

Section titled “Local phase symmetry and the covariant derivative”

Now suppose a matter theory has a global U(1)U(1) symmetry. For a Dirac field,

L0=ψˉ(iγμμm)ψ\mathcal L_0=\bar\psi(i\gamma^\mu\partial_\mu-m)\psi

is invariant under

ψeiqαψ,ψˉψˉeiqα,\psi\mapsto e^{iq\alpha}\psi, \qquad \bar\psi\mapsto \bar\psi e^{-iq\alpha},

when α\alpha is constant. If α\alpha depends on spacetime, then

μψμ(eiqαψ)=eiqα(μψ+iq(μα)ψ).\partial_\mu\psi\mapsto \partial_\mu(e^{iq\alpha}\psi) =e^{iq\alpha}\left(\partial_\mu\psi+iq(\partial_\mu\alpha)\psi\right).

The extra term prevents μψ\partial_\mu\psi from transforming like ψ\psi. Introduce a vector field AμA_\mu and define

Dμψ=(μiqAμ)ψ.D_\mu\psi=(\partial_\mu-iqA_\mu)\psi.

If

AμAμ+μα,A_\mu\mapsto A_\mu+\partial_\mu\alpha,

then

DμψeiqαDμψ.D_\mu\psi\mapsto e^{iq\alpha}D_\mu\psi.

Thus a locally invariant Dirac Lagrangian is

L=ψˉ(iγμDμm)ψ14FμνFμν.\mathcal L=\bar\psi(i\gamma^\mu D_\mu-m)\psi -{1\over4}F_{\mu\nu}F^{\mu\nu}.

Expanding the covariant derivative gives

ψˉiγμDμψ=ψˉiγμμψ+qψˉγμψAμ.\bar\psi i\gamma^\mu D_\mu\psi =\bar\psi i\gamma^\mu\partial_\mu\psi +q\bar\psi\gamma^\mu\psi A_\mu.

Up to the sign convention for qq, the photon couples to the conserved current

Jμ=ψˉγμψ.J^\mu=\bar\psi\gamma^\mu\psi.

The scalar version is similar. For a complex scalar field of charge qq,

L=(Dμϕ)(Dμϕ)m2ϕϕ14FμνFμν.\mathcal L=(D_\mu\phi)^*(D^\mu\phi)-m^2\phi^*\phi -{1\over4}F_{\mu\nu}F^{\mu\nu}.

Here the star matters:

(Dμϕ)=(μ+iqAμ)ϕ.(D_\mu\phi)^* =(\partial_\mu+iqA_\mu)\phi^*.

Thus the conjugate field carries the opposite charge. A common alternative is to absorb the charge into the gauge potential. Define

Aμ=qAμ,Dμ=μiAμ.\mathcal A_\mu=qA_\mu, \qquad \mathcal D_\mu=\partial_\mu-i\mathcal A_\mu.

Then

(Dμϕ)(Dμϕ)=(μ+iAμ)ϕ(μiAμ)ϕ.(\mathcal D_\mu\phi)^*(\mathcal D^\mu\phi) =(\partial_\mu+i\mathcal A_\mu)\phi^* (\partial^\mu-i\mathcal A^\mu)\phi.

If the kinetic term for the rescaled gauge field is written with the coupling outside, the Lagrangian becomes schematically

L=(Dμϕ)(Dμϕ)m2ϕ214q2FμνFμν.\mathcal L=(\mathcal D_\mu\phi)^*(\mathcal D^\mu\phi) -m^2|\phi|^2-{1\over4q^2}\mathcal F_{\mu\nu}\mathcal F^{\mu\nu}.

The placement of the coupling is a convention; the opposite signs on the two conjugate scalar factors are not. They are required because ϕ\phi^* has charge q-q when ϕ\phi has charge qq.

Covariant derivative and field strength as connection and curvature

A local phase rotation makes ordinary derivatives transform inhomogeneously. The gauge field AμA_\mu supplies the connection term in DμD_\mu. The field strength FμνF_{\mu\nu} is the curvature: it measures the phase accumulated around an infinitesimal loop.

The field strength can be characterized without guessing its formula. For the Abelian covariant derivative

Dμ=μiqAμ,D_\mu=\partial_\mu-iqA_\mu,

the commutator is

[Dμ,Dν]=iq(μAννAμ)=iqFμν.[D_\mu,D_\nu] =-iq(\partial_\mu A_\nu-\partial_\nu A_\mu) =-iqF_{\mu\nu}.

Thus FμνF_{\mu\nu} is the obstruction to commuting two covariant derivatives. Geometrically, AμA_\mu is a connection and FμνF_{\mu\nu} is its curvature.

This viewpoint generalizes immediately, but the sign bookkeeping is slightly less forgiving. With Hermitian generators and Dμ=μigAμaTaD_\mu=\partial_\mu-igA_\mu^aT^a, the non-Abelian field strength contains a plus sign in components, Fμνa=μAνaνAμa+gfabcAμbAνcF_{\mu\nu}^a=\partial_\mu A_\nu^a-\partial_\nu A_\mu^a+gf^{abc}A_\mu^bA_\nu^c. In matrix notation this same statement is cleaner. Let U(x)=eigαa(x)TaU(x)=e^{ig\alpha^a(x)T^a} and let ψ\psi transform in a representation of a non-Abelian group GG:

ψ(x)U(x)ψ(x),U(x)G.\psi(x)\mapsto U(x)\psi(x), \qquad U(x)\in G.

Write the gauge field as a matrix

Aμ=AμaTa,[Ta,Tb]=ifabcTc,A_\mu=A_\mu^aT^a, \qquad [T^a,T^b]=if^{abc}T^c,

and define

Dμ=μigAμ.D_\mu=\partial_\mu-igA_\mu.

The gauge field must transform so that

DμψU(Dμψ).D_\mu\psi\mapsto U(D_\mu\psi).

Equivalently,

DμUDμU1.D_\mu\mapsto UD_\mu U^{-1}.

This gives

AμUAμU1ig(μU)U1.A_\mu\mapsto U A_\mu U^{-1}-{i\over g}(\partial_\mu U)U^{-1}.

The field strength is defined by

[Dμ,Dν]=igFμν,[D_\mu,D_\nu]=-igF_{\mu\nu},

so

Fμν=μAννAμig[Aμ,Aν].F_{\mu\nu} =\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].

The commutator term is the new feature of Yang–Mills theory. Since [Ta,Tb]=ifabcTc[T^a,T^b]=if^{abc}T^c, the displayed matrix formula gives

Fμνa=μAνaνAμa+gfabcAμbAνc.F_{\mu\nu}^a =\partial_\mu A_\nu^a-\partial_\nu A_\mu^a+gf^{abc}A_\mu^bA_\nu^c.

References that instead use Dμ=μ+igAμD_\mu=\partial_\mu+igA_\mu have the opposite sign in the commutator term. The defining equation [Dμ,Dν]=igFμν[D_\mu,D_\nu]=-igF_{\mu\nu} fixes the signs used here.

Under a gauge transformation,

FμνUFμνU1.F_{\mu\nu}\mapsto UF_{\mu\nu}U^{-1}.

Therefore tr(FμνFμν)\operatorname{tr}(F_{\mu\nu}F^{\mu\nu}) is gauge invariant, and the Yang–Mills action is

LYM=12tr(FμνFμν).\mathcal L_{\mathrm{YM}} =-{1\over2}\operatorname{tr}(F_{\mu\nu}F^{\mu\nu}).

Here the trace is normalized by tr(TaTb)=12δab\operatorname{tr}(T^aT^b)=\tfrac12\delta^{ab}. For a trace in a representation with trR(TRaTRb)=T(R)δab\operatorname{tr}_R(T_R^aT_R^b)=T(R)\delta^{ab}, the coefficient is 1/[4T(R)]-1/[4T(R)].

For U(1)U(1) the commutator vanishes and this reduces to Maxwell theory. For a non-Abelian group, the gauge bosons carry the gauge charge themselves, and the field strength contains cubic and quartic self-interactions after the action is expanded.

Ward identities from replacing a polarization by momentum

Section titled “Ward identities from replacing a polarization by momentum”

For an external photon with momentum qq, a gauge transformation shifts a plane-wave polarization by

ϵμ(q)ϵμ(q)+βqμ.\epsilon_\mu(q)\longmapsto \epsilon_\mu(q)+\beta q_\mu.

A physical scattering amplitude cannot change under this replacement. If the amplitude with the external photon removed is written as Mμ\mathcal M^\mu, then

M=ϵμMμ.\mathcal M=\epsilon_\mu\mathcal M^\mu.

Gauge invariance requires

qμMμ=0.q_\mu\mathcal M^\mu=0.

This is the simplest form of a Ward identity.

For example, the scalar QED three-point vertex for a scalar of charge QQ is proportional to

+iQ(p1+p2)μ,+iQ(p_1+p_2)_\mu,

where the incoming photon momentum is k=p2p1k=p_2-p_1 with both scalar legs on shell. Replacing the photon polarization by kμk_\mu gives the factor

k(p1+p2)=(p2p1)(p1+p2)=p22p12.k\cdot(p_1+p_2) =(p_2-p_1)\cdot(p_1+p_2) =p_2^2-p_1^2.

If the scalar particles have equal mass and are on shell,

p12=p22=m2,p_1^2=p_2^2=m^2,

so

k(p1+p2)=0.k\cdot(p_1+p_2)=0.

This small calculation contains the whole moral: the longitudinal piece of a photon polarization does not contribute to a physical amplitude. In the next page, this becomes a practical rule for QED vertices and tree amplitudes.

Gauge redundancy is not an ordinary global symmetry

Section titled “Gauge redundancy is not an ordinary global symmetry”

It is tempting to speak of gauge invariance as a symmetry, and this language is standard. But one should remember that it is a special kind of symmetry. A global symmetry maps one physical state to another physical state. Gauge redundancy maps one description of a physical configuration to another description of the same physical configuration.

The difference is visible already in Maxwell theory. The field strength FμνF_{\mu\nu} is unchanged by

AμAμ+μα.A_\mu\mapsto A_\mu+\partial_\mu\alpha.

Thus AμA_\mu and Aμ+μαA_\mu+\partial_\mu\alpha cannot represent distinct measured electromagnetic fields. The physical observables must be gauge invariant, such as FμνF_{\mu\nu}, Wilson loops, scattering amplitudes, or properly dressed charged operators.

This does not make gauge theory empty. Quite the opposite: the redundancy imposes strong constraints. It forbids a photon mass term M2AμAμ/2M^2A_\mu A^\mu/2, requires matter to enter through covariant derivatives, enforces current conservation, constrains counterterms, and leads to Ward identities. In non-Abelian gauge theories, the same principle generates the self-interactions of gluons and underlies the structure of the Standard Model.

A massive vector field is described by the Proca equation. Its divergence gives the dynamical constraint μAμ=0\partial_\mu A^\mu=0, so a four-component Lorentz vector carries three physical spin-one polarizations. As M0M\to0, the longitudinal polarization behaves like pμ/Mp^\mu/M and decouples only when it couples to a conserved current.

A massless vector field is described by Maxwell theory with the gauge redundancy AμAμ+μαA_\mu\sim A_\mu+\partial_\mu\alpha. The field strength FμνF_{\mu\nu} is invariant, and the free Maxwell kinetic operator has zero modes precisely along gauge directions. Gauge fixing is required to define a propagator, but physical quantities must be independent of the gauge choice.

Local phase invariance replaces ordinary derivatives by covariant derivatives. The field strength is the commutator of covariant derivatives. For Abelian gauge theory this gives Fμν=μAννAμF_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu; for non-Abelian gauge theory it adds a commutator term and hence gauge-boson self-interactions. The first practical consequence is the Ward identity: amplitudes vanish when an external photon polarization is replaced by its momentum.

Components are not polarizations. The covariant field AμA_\mu has four components, but a photon has two physical helicities. Transversality and the gauge equivalence must both be used in the massless count.

The Proca constraint is not a gauge condition. For M0M\neq0, A=0\partial\cdot A=0 follows from the equations of motion and the longitudinal polarization is physical. In Maxwell theory a Lorenz condition selects representatives of gauge orbits and does not create a third photon polarization.

Gauge transformations are not ordinary global symmetries. Gauge-related potentials describe the same physical configuration. A global symmetry instead relates physically distinct states with the same energy.

The Maxwell operator has no inverse before gauge fixing. Its zero modes are the gauge directions, not an algebraic accident. A propagator becomes well defined only after restricting or lifting those directions.

A bare vector mass breaks Maxwell gauge invariance. The term M2AμAμ/2M^2A_\mu A^\mu/2 is not invariant by itself. Gauge-invariant massive-vector theories require additional structure, such as a Stueckelberg or Higgs field.

A sign convention is a package. With Dμ=μiqAμD_\mu=\partial_\mu-iqA_\mu and ψeiqαψ\psi\mapsto e^{iq\alpha}\psi, the transformation of AμA_\mu, the matter vertex, and the conjugate derivative are fixed. In particular, (Dμϕ)=(μ+iqAμ)ϕ(D_\mu\phi)^*=(\partial_\mu+iqA_\mu)\phi^*.

Derive the Proca constraint. Starting from

μFμν+M2Aν=0,\partial_\mu F^{\mu\nu}+M^2A^\nu=0,

show that μAμ=0\partial_\mu A^\mu=0 for M0M\neq0, and then show that each component of AμA^\mu obeys the Klein–Gordon equation.

Solution

Take ν\partial_\nu of the Proca equation:

νμFμν+M2νAν=0.\partial_\nu\partial_\mu F^{\mu\nu}+M^2\partial_\nu A^\nu=0.

Since Fμν=FνμF^{\mu\nu}=-F^{\nu\mu} and νμ\partial_\nu\partial_\mu is symmetric in μ,ν\mu,\nu,

νμFμν=0.\partial_\nu\partial_\mu F^{\mu\nu}=0.

Therefore

M2νAν=0.M^2\partial_\nu A^\nu=0.

For M0M\neq0,

νAν=0.\partial_\nu A^\nu=0.

Now expand

μFμν=μ(μAννAμ)=2Aνν(μAμ).\partial_\mu F^{\mu\nu} =\partial_\mu(\partial^\mu A^\nu-\partial^\nu A^\mu) =\partial^2A^\nu-\partial^\nu(\partial_\mu A^\mu).

Using μAμ=0\partial_\mu A^\mu=0, the equation becomes

(2+M2)Aν=0.(\partial^2+M^2)A^\nu=0.

Thus the Proca field is a massive Klein–Gordon field in each component, supplemented by the transversality constraint that removes one component.

Show explicitly that the longitudinal Proca polarization decouples from a conserved current in the massless limit. Use

ϵLμ(p)=1M(p,0,0,E),pμ=(E,0,0,p),\epsilon_L^\mu(p)={1\over M}(|\mathbf p|,0,0,E), \qquad p^\mu=(E,0,0,|\mathbf p|),

with E2p2=M2E^2-\mathbf p^2=M^2.

Solution

First check transversality:

pϵL=EpMpEM=0.p\cdot\epsilon_L =E{|\mathbf p|\over M}-|\mathbf p|{E\over M}=0.

The norm is

ϵL2=p2E2M2=1.\epsilon_L^2={|\mathbf p|^2-E^2\over M^2}=-1.

For small MM,

E=p+M22p+O(M4),E=|\mathbf p|+{M^2\over2|\mathbf p|}+O(M^4),

so

ϵLμ(p)=pμM+O(M).\epsilon_L^\mu(p)={p^\mu\over M}+O(M).

The coupling to a current is

ϵLμJμ=1MpμJμ+O(M).\epsilon_L^\mu J_\mu={1\over M}p^\mu J_\mu+O(M).

If the current is conserved, then in momentum space

pμJμ=0.p_\mu J^\mu=0.

The singular piece vanishes, leaving a term of order MM. Hence the longitudinal polarization decouples as M0M\to0.

Let

Kμν(p)=p2ημνpμpν.K^{\mu\nu}(p)=p^2\eta^{\mu\nu}-p^\mu p^\nu.

Show that Kμνpν=0K^{\mu\nu}p_\nu=0. Then add Lgf=(A)2/(2ξ)\mathcal L_{\mathrm{gf}}=-(\partial\cdot A)^2/(2\xi), invert the resulting operator, and explain why its ξ\xi-dependent part vanishes between conserved currents.

Solution

Compute

Kμνpν=p2ημνpνpμpνpν.K^{\mu\nu}p_\nu =p^2\eta^{\mu\nu}p_\nu-p^\mu p^\nu p_\nu.

Since ημνpν=pμ\eta^{\mu\nu}p_\nu=p^\mu and pνpν=p2p^\nu p_\nu=p^2,

Kμνpν=p2pμpμp2=0.K^{\mu\nu}p_\nu =p^2p^\mu-p^\mu p^2=0.

Thus pνp_\nu is a null eigenvector of the kinetic operator. A matrix or differential operator with a null eigenvector has no inverse on the full vector space. Since a propagator is the inverse of the quadratic operator, the Maxwell propagator is undefined until one removes or lifts the gauge-zero directions by a gauge condition.

Introduce the transverse and longitudinal projectors

PμνT=ημνpμpνp2,PμνL=pμpνp2.P^{\mathrm T}_{\mu\nu}=\eta_{\mu\nu}-{p_\mu p_\nu\over p^2}, \qquad P^{\mathrm L}_{\mu\nu}={p_\mu p_\nu\over p^2}.

The gauge-fixed operator is proportional to

p2(PT+1ξPL),-p^2\left(P^{\mathrm T}+{1\over\xi}P^{\mathrm L}\right),

so its Feynman inverse is

Dμν(p)=ip2+iϵ(PμνT+ξPμνL)=ip2+iϵ[ημν(1ξ)pμpνp2].D_{\mu\nu}(p) ={-i\over p^2+i\epsilon} \left(P^{\mathrm T}_{\mu\nu}+\xi P^{\mathrm L}_{\mu\nu}\right) ={-i\over p^2+i\epsilon} \left[\eta_{\mu\nu}-(1-\xi){p_\mu p_\nu\over p^2}\right].

Contracting the last term with two currents gives a factor (pJ1)(pJ2)(p\cdot J_1)(p\cdot J_2). It vanishes when both currents are conserved, proving that the exchange amplitude is independent of ξ\xi.

For scalar QED, the one-photon vertex for an incoming scalar momentum p1p_1 and outgoing scalar momentum p2p_2 is proportional to

(p1+p2)μ.(p_1+p_2)_\mu.

Let the photon momentum be q=p1p2q=p_1-p_2. Show that replacing the photon polarization by qμq_\mu gives zero for equal-mass on-shell scalar particles.

Solution

The replacement gives the contraction

q(p1+p2).q\cdot(p_1+p_2).

Using q=p1p2q=p_1-p_2,

q(p1+p2)=(p1p2)(p1+p2)=p12p22.q\cdot(p_1+p_2) =(p_1-p_2)\cdot(p_1+p_2) =p_1^2-p_2^2.

For equal-mass on-shell scalar particles,

p12=p22=m2.p_1^2=p_2^2=m^2.

Therefore

q(p1+p2)=0.q\cdot(p_1+p_2)=0.

This is the tree-level Ward identity for the scalar three-point vertex.

With

Dμ=μigAμ,Aμ=AμaTa,D_\mu=\partial_\mu-igA_\mu, \qquad A_\mu=A_\mu^aT^a,

define FμνF_{\mu\nu} by

[Dμ,Dν]=igFμν.[D_\mu,D_\nu]=-igF_{\mu\nu}.

Derive

Fμν=μAννAμig[Aμ,Aν].F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].
Solution

Let DμD_\mu act on a test field ψ\psi. Then

DμDνψ=(μigAμ)(νψigAνψ).D_\mu D_\nu\psi =(\partial_\mu-igA_\mu)(\partial_\nu\psi-igA_\nu\psi).

Expanding,

DμDνψ=μνψig(μAν)ψigAνμψigAμνψg2AμAνψ.D_\mu D_\nu\psi =\partial_\mu\partial_\nu\psi -ig(\partial_\mu A_\nu)\psi -igA_\nu\partial_\mu\psi -igA_\mu\partial_\nu\psi -g^2A_\mu A_\nu\psi.

Subtract the same expression with μ\mu and ν\nu exchanged. The second-derivative terms cancel, and so do the terms where a gauge field multiplies a derivative of ψ\psi. The result is

[Dμ,Dν]ψ=ig(μAννAμ)ψg2(AμAνAνAμ)ψ.[D_\mu,D_\nu]\psi =-ig(\partial_\mu A_\nu-\partial_\nu A_\mu)\psi -g^2(A_\mu A_\nu-A_\nu A_\mu)\psi.

This can be written as

[Dμ,Dν]ψ=ig(μAννAμig[Aμ,Aν])ψ.[D_\mu,D_\nu]\psi =-ig\left(\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu]\right)\psi.

Therefore

Fμν=μAννAμig[Aμ,Aν].F_{\mu\nu}=\partial_\mu A_\nu-\partial_\nu A_\mu-ig[A_\mu,A_\nu].

For an Abelian group the commutator vanishes, recovering the Maxwell field strength.

  • Coleman, Sidney. Lectures of Sidney Coleman on Quantum Field Theory. Edited by Bryan Gin-ge Chen, David Derbes, David Griffiths, Brian Hill, Richard Sohn, and Yuan-Sen Ting. World Scientific, 2019. Chapters 26–27 discuss massive vectors, gauge invariance, and gauge-field quantization.
  • Polyakov, A. M. Gauge Fields and Strings. Harwood Academic Publishers, 1987. Chapter 1 develops the gauge-field viewpoint from long-distance physics.
  • Srednicki, Mark. Quantum Field Theory. Cambridge University Press, 2007. The chapters on spin-one fields and electrodynamics give a compact route from the Maxwell operator to gauge-fixed propagators and Ward identities.
  • Weinberg, Steven. The Quantum Theory of Fields, Volume I: Foundations. Cambridge University Press, 1995. Sections 5.3 and 8.1–8.6 treat vector fields, constraints, gauge invariance, and QED.
  • Zee, A. Quantum Field Theory in a Nutshell. 2nd ed., Princeton University Press, 2010. Chapters III.4 and IV.5 discuss gauge redundancy, Maxwell zero modes, and non-Abelian curvature.