Lecture 7 — The Lagrangian Formalism: Action, Euler–Lagrange Equations, and the Pendulum#
Source: NPTEL Classical Physics, Mod-01 Lec-07 (Lagrangian formalism), Prof. V. Balakrishnan.
Lecture 5–6 classified what a system’s flow can look like once you already have equations of motion. This lecture backtracks to ask a more basic question: where do the equations of motion themselves come from? Newton’s second law was handed down by experience, not derived — the Lagrangian formalism supplies the missing “why,” and does it through the same kind of extremal principle that already shows up everywhere else in physics.
An extremal principle for dynamics#
Static equilibrium is minimized potential energy. Isolated thermal equilibrium is maximized entropy. Equilibrium at fixed temperature and volume minimizes the Helmholtz free energy; at fixed temperature and pressure, the Gibbs free energy. In every one of these static problems, the system sits at an extremum of some scalar. The claim behind the Lagrangian formalism is that motion obeys an extremal principle too — not extremizing a single number, but extremizing an integral along the entire path taken through time.
For a system with generalized coordinates \(q_1,\dots,q_n\), generalized velocities \(\dot q_1,\dots,\dot q_n\), and possibly explicit time dependence, posit a scalar function \(L(q,\dot q,t)\) — the Lagrangian. Given a start point \(q(t_1)\) and an end point \(q(t_2)\), define the action
Of all conceivable paths connecting those two points in \((q,\dot q)\)-space, the system follows the one along which \(S\) is stationary — extremal, not necessarily a minimum, though it usually is one for the problems of interest. That’s the principle of least action: \(\delta S = 0\).
Deriving the Euler–Lagrange equations#
Fix the endpoints \(t_1, t_2\) and consider an arbitrary variation \(q \to q + \delta q\) that vanishes at both ends, \(\delta q(t_1) = \delta q(t_2) = 0\) — the paths all have to start and finish at the same two points, so there’s no freedom to vary there. Demanding \(\delta S = 0\) gives
summed implicitly over each independent degree of freedom. Time itself is never varied — it’s the arena the dynamical variables move through, not a dynamical variable — but the variation and the time derivative do commute, \(\delta \dot q = \frac{d}{dt}\delta q\), since shifting the path and differentiating along it are unrelated operations. Integrating the second term by parts,
The boundary term vanishes identically because \(\delta q = 0\) at both ends by construction. What’s left has to hold for arbitrary \(\delta q(t)\) in between, and for arbitrary \(t_1, t_2\) too — the action principle is true for any two points along the trajectory, so you can slice the interval as finely as you like and apply it locally. That forces the integrand’s bracket to vanish pointwise, giving the Euler–Lagrange equations,
A statement about an entire path — extremize a single number, the action, integrated over the whole trajectory at once — has turned into a local, pointwise differential equation. That’s not a paradox: the system doesn’t need to know the future to extremize \(S\) globally, because \(t_1\) and \(t_2\) were arbitrary all along, so the same stationarity condition holds between any two nearby instants, and a local differential equation is exactly what “stationary between every pair of nearby points” means.
Why only \(q\) and \(\dot q\)?#
The derivation assumed \(L\) depends on positions and velocities alone, never on \(\ddot q\) or higher. That’s not a mathematical necessity — if \(L\) contained \(\ddot q\), the same integration-by-parts trick, applied twice, would produce a \(+\dfrac{d^2}{dt^2} \dfrac{\partial L}{\partial \ddot q}\) term in the equations of motion, and in general one sign-alternating derivative term for every extra order. It’s a physical input: the independent dynamical variables of a mechanical system are exactly the ones you’re free to specify at an instant without reference to the forces acting — position and velocity. Once those are fixed, acceleration is determined (it’s what Newton’s second law says), so it can’t also be an independent slot in \(L\). A few exceptional systems genuinely do need \(\ddot q\), but the overwhelming majority don’t, and this course won’t need them.
Recovering Newton’s second law#
For a conservative system without friction, the Lagrangian that reproduces the known equations of motion is \(L = T - V\), the kinetic energy minus the potential energy — not a derivation, just the choice that happens to work. For a set of particles with Cartesian coordinates \(q_i\),
Since \(V\) carries no velocity dependence, \(\partial L/\partial \dot q_i = m_i \dot q_i\), and the Euler–Lagrange equation collapses to
mass times acceleration equals force — Newton’s second law, recovered rather than assumed. \(L\) itself is required to be a scalar (invariant under coordinate rotations, and later, under Lorentz transformations too), which is exactly why \(T-V\), built from dot products and a scalar potential, is a sensible guess in the first place.
Eliminating constraints: the Atwood machine#
The real payoff shows up with constrained systems. The textbook approach — Atwood’s machine, two masses \(m_1, m_2\) hanging over a frictionless pulley by an inextensible string — normally requires introducing the string tension as an unknown constraint force, solving for it, and only then extracting the acceleration. The Lagrangian method skips the constraint force entirely.
Measure both masses’ positions \(x_1, x_2\) downward from the pulley, with the potential zero there, so \(V = -m_1 g x_1 - m_2 g x_2\) and
The string being inextensible means \(x_1 + x_2 = \ell\) for constant \(\ell\), i.e. \(\dot x_2 = -\dot x_1\). Substitute the constraint directly into \(L\) before varying, eliminating \(x_2\) in favor of the single independent coordinate \(x_1\):
where the constant \(m_2 g \ell\) term is irrelevant — the Euler–Lagrange equations only ever see derivatives of \(L\), so an additive constant changes nothing. (This is the general statement that the Lagrangian is not unique: shifting the zero of potential energy, as in choosing where \(V=0\), never affects the physics, exactly as it doesn’t in Newtonian mechanics. That freedom disappears relativistically, where there’s an absolute zero of energy set by rest mass.) The single Euler–Lagrange equation for \(x_1\) gives
the familiar Atwood-machine acceleration — reached without ever writing down the tension. Rather than adding constraint forces to the problem, the Lagrangian method uses the constraint to remove a coordinate, and the equations of motion for whatever coordinates remain come out automatically consistent with it.
What the Lagrangian formalism buys you beyond Newton#
Eliminating constraint forces is only the most immediate advantage:
Constraints are absorbed by reducing to independent coordinates, rather than needing to be modeled as forces (normal reactions, tensions, and the like).
Special relativity: the formalism generalizes to the relativistic regime, where Newton’s equations no longer hold, essentially by requiring \(L\) to be a Lorentz scalar rather than merely a rotational one.
Fields: it extends to systems with a continuous number of degrees of freedom. Maxwell’s equations for the electromagnetic field don’t resemble Newton’s equations at all, but they do emerge as Euler–Lagrange equations from a suitable field Lagrangian — the same unifying machinery covers particles and fields.
The chief disadvantage is that the formalism isn’t the easiest starting point for quantization — that’s what motivates shifting to the Hamiltonian formalism later, which trades the Euler–Lagrange equations’ second-order-in-time structure (there’s a \(\ddot{}\) hiding inside \(\frac{d}{dt}\partial L/\partial \dot q\)) for genuinely first-order dynamics, at the cost of doubling the number of variables from \(q\)’s alone to \(q\)’s and their conjugate momenta.
Worked example: the simple pendulum#
A massless rigid rod of length \(\ell\) carries a bob of mass \(m\), swinging frictionlessly in a vertical plane with angular displacement \(\theta\) measured from the bottom. The bob’s speed is \(\ell\dot\theta\), and taking the potential zero at the lowest point,
The single Euler–Lagrange equation, \(\frac{d}{dt}\frac{\partial L}{\partial \dot\theta} = \frac{\partial L}{\partial \theta}\), gives \(m\ell^2\ddot\theta = -mg\ell\sin\theta\), i.e.
This is the exact pendulum equation — genuinely nonlinear, since \(\sin\theta\) carries every odd power of \(\theta\). It reduces to simple harmonic motion, \(\ddot\theta \approx -\frac{g}{\ell}\theta\) with \(\omega_0 = \sqrt{g/\ell}\) and period \(2\pi\sqrt{\ell/g}\), only in the small-angle limit \(\sin\theta \approx \theta\). There is no universal cutoff angle (“5.5 degrees,” as some textbooks assert) below which this approximation is simply “valid” — how small \(\theta\) needs to be depends entirely on what accuracy you’re willing to accept, since the size of the error is set by the first neglected term, \(\theta^3/6\), relative to \(\theta\) itself. Only at \(\theta \equiv 0\) are the linear and nonlinear equations exactly equal.
Because the rod (not a string) can swing all the way around, the phase portrait covers both bounded oscillation and full rotation. The equilibria alternate exactly as Lecture 3–4 argued they must: centers (stable) at every even multiple of \(\pi\), sitting at the bottom of the potential, and saddles (unstable) at every odd multiple, at the top. The saddle energy \(E_s = 2mg\ell\) organizes everything:
\(E < E_s\): closed orbits trapped in a single well — ordinary back-and-forth oscillation, harmonic only in the small-amplitude limit.
\(E = E_s\): the separatrix. Released infinitesimally away from a saddle, the pendulum swings out, climbs asymptotically toward the next saddle, and never quite arrives — technically a heteroclinic connection rather than homoclinic, since \(-\pi\) and \(\pi\) are physically the same point but distinct on this unrolled \(\theta\) axis.
\(E > E_s\): the barrier no longer separates anything, and the trajectory becomes an open, unbounded curve — the pendulum has enough energy to loop over the top and rotates continuously rather than oscillating, trading potential for kinetic energy each time it passes back through the bottom.
The period diverges at the separatrix#
Approaching the separatrix from below, the amplitude grows toward \(\pi\) and the pendulum spends longer and longer crawling past the top before turning back — in the limit, infinitely long, since it’s asymptotically approaching an equilibrium point it can never actually reach in finite time. The quarter-period integral,
is a complete elliptic integral of the first kind, \(T = 4\sqrt{\ell/g}\; K\!\big(\sin^2\tfrac{\theta_0}{2}\big)\), which is finite for every \(\theta_0 < \pi\) but diverges logarithmically as \(\theta_0 \to \pi\):
At \(10\%\) of the way to the separatrix the period is still within a percent of the small-angle value; by \(99.9\%\) of the way there it has grown past five times that value, and it keeps climbing without bound. The exact solution \(\theta(t)\) isn’t expressible in elementary functions once the amplitude leaves the small-angle regime — it’s an elliptic function — but exactly at the separatrix (\(E = E_s\)) a closed form reappears, related to the soliton solution of the sine-Gordon equation, a nonlinear wave equation that recurs throughout physics far beyond this one pendulum.
Where this is headed#
The Lagrangian formalism answers “why Newton’s equations” by subsuming them into a single extremal principle, general enough to survive the move to special relativity and to continuous fields — generalizations Newton’s \(F=ma\) has no way to make. What it doesn’t immediately hand you is first-order dynamics: the Euler–Lagrange equations are second order in time, the same complication Lecture 5–6 sidesteps by working directly in \((q,\dot q)\) phase space. Recasting the same physics with momenta instead of velocities as independent variables — trading \(n\) second-order equations for \(2n\) first-order ones — is exactly the move the Hamiltonian formalism makes next.