Multivariable Differential Calculus
We move from functions of one variable to functions of several variables: temperature on a map, the elevation of a terrain, the energy of a system. The derivative splits into partial derivatives gathered in the gradient, the tangent line becomes a tangent plane, concavity becomes the Hessian matrix. With these tools we learn to approximate (Taylor), to find the direction of steepest ascent, and to optimise — both freely and under constraints.
Complete Theory
8A function assigns a number to each point of the plane: think of as the elevation at point of a landscape. The domain is the set of points where makes sense (e.g. requires ).
Level curves. The level curve of value is the set of points where equals exactly :
These are precisely the contour lines of topographic maps (points at the same elevation) or the isotherms of weather maps. Cutting the graph of with horizontal planes and projecting down gives the level curves.
How to read them. Where level curves are dense, the function varies rapidly (steep slope); where they are sparse, it varies slowly (flat). Closed concentric level curves signal a peak or a basin. In the analogue is level surfaces (e.g. equipotential surfaces in electrostatics).
The definition of limit mirrors the one-variable case, but "approaching" now happens in all directions of the plane. The Euclidean norm replaces the absolute value:
The big difference. In you approach only from the right or the left. In you can approach along infinitely many paths: lines, parabolas, spirals. For the limit to exist, it must equal along every path.
Proving non-existence. Just find two paths giving different limits. Classic example: at . Along the -axis (): . Along : . Different limits ⟹ the limit does not exist. (Polar coordinates , are often the fastest test: if the result depends on , the limit does not exist.)
is continuous at if .
The partial derivative measures how changes when moving only in the direction, keeping frozen:
In practice you differentiate with respect to treating as constant. It is the slope of the graph along the slice .
Warning: partials are not enough. Unlike the 1D case, the existence of partial derivatives does not even guarantee continuity! A stronger notion is needed: differentiability. is differentiable at if it is well approximated by a linear function (the differential):
The error tends to zero faster than : near the graph is almost a plane.
Practical criterion (total differential theorem). If the partial derivatives exist and are continuous in a neighbourhood (), then is differentiable, hence continuous. This covers almost all "elementary" functions.
The gradient gathers all partial derivatives into a vector:
The directional derivative in the direction of the unit vector measures the slope of in that direction, and is the projection of the gradient:
Three fundamental properties of the gradient.
- Direction of steepest ascent. is maximal when is aligned with (): the gradient points "uphill" the most steeply. The maximal slope is .
- Perpendicularity to level curves. Along a level curve does not change, so : the gradient is orthogonal to level curves. (That is why water flows perpendicular to contour lines.)
- Steepest descent in the direction : the principle behind gradient descent used to train neural networks.
The tangent plane to the graph of at is the analogue of the tangent line: the plane that best approximates the surface there.
It is exactly the linear approximation given by the differential (§3). For near , is almost its tangent plane.
Second-order Taylor. For a better approximation we add the quadratic term built from the Hessian matrix (the second derivatives):
The linear term tilts, the quadratic term curves. At critical points (where ) only the quadratic term survives: it is the Hessian's quadratic form that decides minimum, maximum or saddle — the topic of the next section.
To optimise we first find the critical points, where the gradient vanishes:
There the tangent plane is horizontal: candidates for maximum, minimum or saddle. To classify them we use the Hessian matrix and the sign of the quadratic form it generates. In 2D, with :
- and : positive definite ⟹ local minimum (bowl up, grows in every direction).
- and : negative definite ⟹ local maximum.
- : indefinite ⟹ saddle point (grows in some directions, decreases in others — like a mountain pass or a horse saddle).
- : test inconclusive, higher orders needed.
Why it works. Near a critical point : the behaviour is that of the quadratic form. Positive definite ⟹ bowl (minimum); indefinite ⟹ saddle. The criterion is the Hessian eigenvalue test in disguise (same-sign positive eigenvalues ⟹ minimum, etc.).
Chain rule. If is evaluated along a curve , the composite has derivative given by the gradient dotted with the velocity:
It is the multivariable version of . It confirms that the directional derivative (§4) is the derivative along a line traversed at unit speed.
Implicit function theorem (Dini). A level curve is generally not the graph of a function. But locally it is, provided the tangent is not vertical. If , and , then there exists with near , and
Derivation: differentiating in via the chain rule, . Example: on the circle , near we have , horizontal at the top ().
Often we want to optimise not everywhere, but subject to a constraint (e.g. maximise an area at fixed perimeter). The method of Lagrange multipliers solves the problem.
The geometric idea. At a constrained optimum, the level curve of is tangent to the constraint . Since the gradient is perpendicular to level curves, this means and are parallel:
The number is the Lagrange multiplier. Solve the three equations (, , ) in the unknowns .
Meaning of . It measures the sensitivity of the optimum to relaxing the constraint: in economics it is the "shadow price", the change of optimal value per unit of constraint. For several constraints, .
Worked Examples
5Exercises with Solutions
6Keep studying
Guided exercises on this topic
Recommended Books
As an Amazon Associate I earn from qualifying purchases.