Skip to content

Symbols, Characteristics, and PDE Type

The principal symbol is the part of a differential operator seen by very short-wavelength disturbances. Its zero set identifies characteristic covectors and characteristic hypersurfaces. Elliptic operators have no nonzero real characteristic covectors, hyperbolic operators organize propagation around real characteristic cones, and parabolic evolution uses an anisotropic balance between one time derivative and two spatial derivatives. These distinctions explain, at the structural level, why Laplace, wave, and heat equations require different data and produce different kinds of solutions.

Helpful background. Fourier Series, Fourier Transforms, and Plancherel Theory supplies the derivative-to-momentum correspondence used to interpret symbols.

Let PP be a scalar differential operator of order mm on an open set:

P=αmbα(x)α,α=1α1nαn.P = \sum_{|\alpha|\leq m} b_\alpha(x)\,\partial^\alpha, \qquad \partial^\alpha = \partial_1^{\alpha_1}\cdots\partial_n^{\alpha_n}.

With this site’s Fourier convention,

u~(p)=dnxe+ipxu(x),u(x)=dnp(2π)neipxu~(p),\widetilde u(p) = \int d^n x\,e^{+ip\cdot x}u(x), \qquad u(x) = \int\frac{d^n p}{(2\pi)^n}\, e^{-ip\cdot x}\widetilde u(p),

a derivative acts as jipj\partial_j\mapsto -ip_j. It is therefore convenient to write Dj=ijD_j=i\partial_j, so that DjpjD_j\mapsto p_j. If

P=αmaα(x)Dα,P = \sum_{|\alpha|\leq m}a_\alpha(x)D^\alpha,

then its full symbol and principal symbol in this convention are

p(x,ξ)=αmaα(x)ξα,pm(x,ξ)=α=maα(x)ξα.p(x,\xi) = \sum_{|\alpha|\leq m}a_\alpha(x)\xi^\alpha, \qquad p_m(x,\xi) = \sum_{|\alpha|=m}a_\alpha(x)\xi^\alpha.

Equivalently, starting from the coefficients bαb_\alpha of the α\partial^\alpha expression, substitute jiξj\partial_j\mapsto-i\xi_j and keep the terms of degree mm. Some mathematics references define Dj=ijD_j=-i\partial_j because they use the opposite Fourier sign. That changes the displayed symbol by predictable powers of 1-1, but not the geometric question of whether its homogeneous leading part vanishes.

The principal symbol also follows directly from a high-frequency test. For a smooth real phase ϕ\phi and amplitude aa,

uλ(x)=eiλϕ(x)a(x),λ+,u_\lambda(x)=e^{-i\lambda\phi(x)}a(x), \qquad \lambda\longrightarrow+\infty,

one has

Puλ=eiλϕ[λmpm(x,dϕ)a+O(λm1)].P u_\lambda = e^{-i\lambda\phi} \left[ \lambda^m p_m(x,d\phi)\,a +O(\lambda^{m-1}) \right].

Thus pmp_m controls the leading response to rapidly oscillating waves. Because dϕd\phi is a covector, not a vector, symbols naturally live on the cotangent bundle. Under a change of coordinates, the covector and the principal symbol transform together, so the equation pm(x,ξ)=0p_m(x,\xi)=0 is coordinate independent.

For a square NN-component system, pm(x,ξ)p_m(x,\xi) is an N×NN\times N matrix. The relevant question is then whether this matrix is invertible. Systems that are not square require a rank condition rather than a determinant test.

Characteristics and admissible data surfaces

Section titled “Characteristics and admissible data surfaces”

A nonzero real covector ξTxM\xi\in T_x^*M is characteristic for a scalar operator when

pm(x,ξ)=0.p_m(x,\xi)=0.

For a square system, it is characteristic when

detpm(x,ξ)=0.\det p_m(x,\xi)=0.

Let a hypersurface SS be locally defined by ϕ(x)=0\phi(x)=0 with dϕ0d\phi\neq0 on SS. It is characteristic at xSx\in S precisely when its normal covector is characteristic:

pm(x,dϕx)=0.p_m(x,d\phi_x)=0.

On a noncharacteristic surface, the highest normal derivative can locally be solved for in terms of tangential derivatives and lower-order data. This is the algebraic prerequisite for posing ordinary Cauchy data there. It is not, by itself, a proof of existence, uniqueness, stability, or global well-posedness.

Only the top-order coefficients enter pmp_m. Lower-order terms can change spectra, decay rates, masses, and Green functions, but they do not change the characteristic set of a fixed-order operator.

Elliptic, hyperbolic, and parabolic structures

Section titled “Elliptic, hyperbolic, and parabolic structures”

A scalar operator is elliptic at xx when

pm(x,ξ)0for every real ξ0.p_m(x,\xi)\neq0 \quad\text{for every real }\xi\neq0.

For systems, ellipticity means that pm(x,ξ)p_m(x,\xi) is invertible for every nonzero real ξ\xi. Elliptic equations have no real characteristic hypersurfaces. Their basic problems are consequently boundary-value problems rather than finite-speed initial-value problems. Ellipticity underlies regularity: under suitable hypotheses, singularities in the solution must be tied to singularities in the forcing or boundary data.

For the Euclidean Laplacian,

P=Δ=j=1nj2,p2(ξ)=ξ2,P=-\Delta = -\sum_{j=1}^n\partial_j^2, \qquad p_2(\xi)=|\xi|^2,

so the only real zero is ξ=0\xi=0 and Δ-\Delta is elliptic.

Let pmp_m be a real homogeneous principal polynomial. One standard scalar definition says that it is hyperbolic with respect to a real covector nn if

pm(x,n)0p_m(x,n)\neq0

and, for every real ξ\xi, all roots λ\lambda of

λpm(x,ξ+λn)\lambda\longmapsto p_m(x,\xi+\lambda n)

are real. The covector nn selects a time direction; characteristic roots then describe real propagation speeds relative to it. Distinct roots give strict hyperbolicity away from unavoidable degeneracies. For systems, real characteristic roots alone need not give stable evolution: strong hyperbolicity additionally requires sufficiently uniform control of the characteristic modes.

With the (+)(+---) metric convention, the wave operator is

=t22.\Box=\partial_t^2-\nabla^2.

Its principal symbol in the Fourier convention above is

p2(τ,k)=τ2+k2.p_2(\tau,\mathbf k) = -\tau^2+|\mathbf k|^2.

The characteristic set is the null cone τ2=k2\tau^2=|\mathbf k|^2. The normal dtdt to a constant-time surface gives p2(dt)=1p_2(dt)=-1, so such a surface is noncharacteristic. By contrast, a surface ϕ(t,x)=0\phi(t,\mathbf x)=0 is characteristic when

(tϕ)2+ϕ2=0.-(\partial_t\phi)^2+|\nabla\phi|^2=0.

For a second-order scalar operator in two variables,

P2=Ax2+2Bxy+Cy2,P_2 = A\,\partial_x^2 +2B\,\partial_x\partial_y +C\,\partial_y^2,

the traditional local classification uses

Δtype=B2AC.\Delta_{\mathrm{type}}=B^2-AC.

At a nonsingular point, Δtype<0\Delta_{\mathrm{type}}<0 is elliptic, Δtype>0\Delta_{\mathrm{type}}>0 is hyperbolic, and Δtype=0\Delta_{\mathrm{type}}=0 is parabolic. This discriminant is useful for reducing a scalar second-order equation to a local canonical form.

Evolutionary parabolic equations are more clearly described with weighted orders. For the heat operator

Pheat=tκΔx,κ>0,P_{\mathrm{heat}}=\partial_t-\kappa\Delta_{\mathbf x}, \qquad \kappa>0,

one time derivative balances two spatial derivatives. Assigning weight 22 to the temporal frequency τ\tau and weight 11 to each spatial frequency kjk_j gives the weighted principal symbol

pheat(2)(τ,k)=iτ+κk2.p_{\mathrm{heat}}^{(2)}(\tau,\mathbf k) = -i\tau+\kappa|\mathbf k|^2.

This expresses the scaling tx2t\sim|\mathbf x|^2. If ordinary isotropic second-order counting were used instead, the time derivative would be discarded and the result would miss the evolution structure. Parabolicity also carries a time orientation: spatial Fourier modes satisfy

tu~(t,k)+κk2u~(t,k)=0,u~(t,k)=eκk2tu~(0,k)\partial_t\widetilde u(t,\mathbf k) +\kappa|\mathbf k|^2\widetilde u(t,\mathbf k)=0, \qquad \widetilde u(t,\mathbf k) = e^{-\kappa|\mathbf k|^2t}\widetilde u(0,\mathbf k)

for forward time. Reversing that evolution amplifies high frequencies and is unstable.

OperatorRelevant leading symbolStructureImmediate consequence
Δx+m2-\Delta_{\mathbf x}+m^2k2\lvert\mathbf k\rvert^2ellipticno nonzero real characteristic covector
+m2\Box+m^2τ2+k2-\tau^2+\lvert\mathbf k\rvert^2hyperbolicnull characteristic cone and admissible spacelike Cauchy surfaces
tκΔx\partial_t-\kappa\Delta_{\mathbf x}iτ+κk2-i\tau+\kappa\lvert\mathbf k\rvert^2 with weights (2,1)(2,1)parabolicforward smoothing with tx2t\sim\lvert\mathbf x\rvert^2

The mass m2m^2 does not appear in either second-order principal symbol. It changes the full Fourier-space equation, but not the high-frequency characteristics.

For the Klein–Gordon equation,

(+m2)ϕ=0,(\Box+m^2)\phi=0,

the full plane-wave condition is

p02+p2+m2=0,or equivalentlyp2=m2.-p_0^2+|\mathbf p|^2+m^2=0, \qquad\text{or equivalently}\qquad p^2=m^2.

This massive dispersion relation is not the characteristic equation. The principal symbol omits m2m^2, so the characteristic cone remains p2=0p^2=0. Massive disturbances have subluminal group velocity, while the high-frequency propagation boundary is still null.

The free Dirac operator is first order:

PD=iγμμm.P_D=i\gamma^\mu\partial_\mu-m.

Its principal symbol is

σ1(PD)(p)=γμpμ=p ⁣ ⁣ ⁣/.\sigma_1(P_D)(p) = \gamma^\mu p_\mu = p\!\!\!/.

The Clifford relation gives

(p ⁣ ⁣ ⁣/)2=p21,(p\!\!\!/)^2=p^2\mathbb 1,

so the symbol fails to be invertible on the null cone. The same conclusion follows by composing the two mass signs:

(iγμμm)(iγνν+m)=(+m2)1.(i\gamma^\mu\partial_\mu-m) (i\gamma^\nu\partial_\nu+m) = -(\Box+m^2)\mathbb 1.

After Euclidean continuation, ΔE+m2-\Delta_E+m^2 is elliptic instead. This change of operator type is one reason Euclidean boundary-value methods and Lorentzian causal evolution must not be interchanged without specifying the continuation and its prescription.

The scalar-field application is developed on Klein–Gordon field and mode solutions.

The principal symbol is deliberately incomplete.

  • It does not select retarded, advanced, Feynman, or other Green functions. Those choices require boundary conditions, support conditions, or an explicit i0i0 prescription.
  • It does not replace an energy estimate or a well-posedness theorem. Hyperbolicity is the starting structure for such results, not their conclusion.
  • It does not encode masses or potentials, because these are lower-order terms.
  • It can vary with position. A variable-coefficient equation may change type across the domain, and conclusions valid in one region need not cross that interface unchanged.
  • For gauge systems or constrained equations, the unreduced principal symbol may be singular for structural reasons. Gauge fixing, constraints, or a reduced system must be specified before invertibility is interpreted.

The symbol definitions and ellipticity criterion used here follow Dyatlov 2022, Definition 12.17, §13.3, and Definition 14.1, PDF. The model PDE classifications can be compared with Evans 2010, Chapters 2, 6, and 7, the characteristic-root discussion with Melrose 2016, Hyperbolic Differential Operators, PDF, and the Klein–Gordon and Dirac examples with Schwartz 2014, §§2.3 and 10.3.

For existence and stability in declared function spaces, continue to Weak Solutions, Sobolev Spaces, and Well-Posedness. For inverse kernels, continue to Fundamental Solutions and Green Operators; that page and the symbol page are the two hard inputs to the hyperbolic and elliptic branches.

Using the full dispersion relation as the characteristic equation. Characteristics come from the homogeneous top-order symbol. For +m2\Box+m^2, the mass shell is p2=m2p^2=m^2, but the characteristic cone is p2=0p^2=0.

Classifying the heat operator by ordinary spacetime order alone. The terms t\partial_t and Δx\Delta_{\mathbf x} balance under parabolic scaling, even though their ordinary derivative orders differ. Dropping t\partial_t erases the direction of evolution.

Treating a characteristic covector as a propagation trajectory. A covector is normal to a phase surface. Rays or bicharacteristics arise only after the symbol is used to generate a Hamiltonian flow on phase space.

Ignoring Fourier-sign translations. Under the convention used here, μipμ\partial_\mu\mapsto-ip_\mu and iμpμi\partial_\mu\mapsto p_\mu. A source using D=iD=-i\partial must be translated before formulas are compared term by term.

  1. For P=Ax2+2Bxy+Cy2P=A\partial_x^2+2B\partial_x\partial_y+C\partial_y^2, show that the characteristic-curve slopes are real precisely when B2AC0B^2-AC\geq0, assuming A0A\neq0.

    Solution

    A curve ϕ(x,y)=0\phi(x,y)=0 is characteristic when

    A(xϕ)2+2B(xϕ)(yϕ)+C(yϕ)2=0,A(\partial_x\phi)^2 +2B(\partial_x\phi)(\partial_y\phi) +C(\partial_y\phi)^2=0,

    where an irrelevant overall sign from the Fourier substitution has been removed. Along a level curve y=y(x)y=y(x), xϕ+yϕdy/dx=0\partial_x\phi+\partial_y\phi\,dy/dx=0. Setting s=dy/dxs=dy/dx therefore gives

    As22Bs+C=0,s=B±B2ACA.As^2-2Bs+C=0, \qquad s=\frac{B\pm\sqrt{B^2-AC}}{A}.

    The slopes are real exactly when B2AC0B^2-AC\geq0. Equality gives the repeated characteristic direction of the parabolic case.

  2. Let ϕ(t,x)=tf(x)\phi(t,\mathbf x)=t-f(\mathbf x). Determine when the level surfaces of ϕ\phi are characteristic for the wave operator.

    Solution

    The normal covector is dϕ=(1,f)d\phi=(1,-\nabla f). Substitution into the wave principal symbol gives

    p2(dϕ)=1+f2.p_2(d\phi)=-1+|\nabla f|^2.

    Hence the surface is characteristic exactly when f=1|\nabla f|=1. Constant-time surfaces have f=constantf=\text{constant} and are noncharacteristic.

  3. Solve the spatial Fourier modes of (tκΔ)u=0(\partial_t-\kappa\Delta)u=0 and explain why backward evolution is unstable.

    Solution

    Since Δk2\Delta\mapsto-|\mathbf k|^2, each mode obeys

    tu~+κk2u~=0.\partial_t\widetilde u +\kappa|\mathbf k|^2\widetilde u=0.

    Thus

    u~(t,k)=eκk2(tt0)u~(t0,k).\widetilde u(t,\mathbf k) = e^{-\kappa|\mathbf k|^2(t-t_0)} \widetilde u(t_0,\mathbf k).

    For t>t0t>t_0, high-frequency modes are strongly damped. Recovering earlier data multiplies them by e+κk2(tt0)e^{+\kappa|\mathbf k|^2(t-t_0)}, so arbitrarily small high-frequency errors can become arbitrarily large.