← Back to list

Optimization Theory and Applications-3

Gradient of Quadratic Forms and the Matrix Product Rule

RADHAMADHAB DALAI · 2026-05-11 04:10 · 0 claps · 6.2 min read
#multivariate #quadratic-forms #multivariate-analysis
Open on Medium ↗

Optimization Theory and Applications-3

Gradient of Quadratic Forms

and the Matrix Product Rule

MATHEMATICAL FOUNDATIONS

Vectors, Norms, and Inner Products

We need a bit of a footing before we can have a meaningful conversation about gradients and quadratic forms. In this section we review the necessary vocabulary of linear algebra and real analysis that will support us along the entire development.

We write vectors as column vectors. Given x∈Rn, we have:

1.2 The Inner Product and Euclidean Norm

1.3 Matrices and Linear Maps

Symmetric matrices are at the heart of the theory of quadratic forms, as we shall see in §5.

1.4 Key Algebraic Identities

We will use these transpose rules repeatedly without proof:

2 PARTIAL DERIVATIVES

Partial Derivatives and Multivariate Differentiation

We shall use again and again The basic operation of calculus, differentiation, generalises naturally to functions of several variables. The main idea is to change one variable at a time , keeping all the others fixed . rules of transposition without proof

2.1 Functions of Several Variables

Let f:Rn→R be a scalar-valued function. We write f(x)=f(x1,x2,…,xn). The domain may be an open subset U⊆Rn.

Geometrically, ∂f/∂xi is the slope of the curve obtained by slicing the graph of f with the hyperplane that fixes all variables except xi.

2.2 Total Differentiability

The linear map L can always be represented by a row vector (the gradient), so that L(h)=∇f(a)⊤h.

2.3 Linear Approximation

Differentiability means f is well-approximated near a by the first-order Taylor expansion:

This is the cornerstone identity that connects local behavior of f to its derivatives.

3 THE GRADIENT

The Gradient: Definition, Geometry, and Properties

The gradient is the “multi-dimensional derivative” of a scalar function. It bundles all directional information into a single vector.

The gradient is the “multi-dimensional derivative” of a scalar function. It bundles all directional information into a single vector.

3.1 Geometric Meaning of the Gradient

3.2 Linearity of the Gradient

3.3 Gradient of a Linear Function

4 DIRECTIONAL DERIVATIVES

Directional Derivatives and the Gradient Connection

The partial derivative measures the rate of change along coordinate axes. But what about an arbitrary direction? The directional derivative answers this question.

4.1 Proof of Theorem 4.1

Therefore:

4.2 Maximum Directional Derivative

By the Cauchy–Schwarz inequality:

Equality holds when u=∇f(x)/‖∇f(x)‖. This confirms that the gradient direction gives the maximum rate of increase.

5 QUADRATIC FORMS

Quadratic Forms: Algebra and Geometry

"A quadratic form is a polynomial in which every term has total degree 
exactly two — the natural generalization of  to many variables."

5.2 WLOG: We May Assume A is Symmetric

Takeaway: Without loss of generality, we always assume A is symmetric when working with quadratic forms.

5.3 Geometric Interpretation

6 CORE DERIVATION

Gradient of a Quadratic Form: Full Derivation

6.2 Derivation Method 2: Perturbation (Definition-Based)

We compute each partial derivative ∂f/∂xk explicitly.

Expand the quadratic form in coordinates:

6.2 Derivation Method 2: Perturbation (Definition-Based)

This approach uses the definition of the gradient as the linear part of the perturbation f(x+h)−f(x).

7 SYMMETRIC MATRICES

Symmetric Matrices and Simplification

The symmetry of a matrix is not just a notational convenience but has deep algebraic and geometric consequences that motivate the entire theory of quadratic forms and optimisation.

7.1 Spectral Theorem

Therefore symmetric matrices always have real eigenvalues and an orthonormal eigenbasis. In the eigenvector coordinate system the quadratic form reduces to a sum of squares:

7.2 Gradient Formula — Symmetric A

For symmetric A=A⊤:

8 MATRIX PRODUCT RULE

The Matrix Product Rule for Gradients

8.1 Product Rule: Scalar × Vector Case

9 CHAIN RULE

Chain Rule in Matrix Calculus

9.1 Jacobian Matrix

9.2 Special Case: Linear Composition

9.3 Chain Rule for Quadratic Compositions

10 SECOND-ORDER STRUCTURE

The Hessian Matrix: Second-Order Derivatives

10.1 Hessian of a Quadratic Form

10.2 Second-Order Taylor Expansion

For a quadratic form f(x)=x⊤Ax, this expansion is exact (the remainder is zero), which makes quadratic forms the simplest non-linear functions.

11 MEAN VALUE THEOREM

Mean Value Theorem for Multivariate Functions

TBC


메타데이터
post_id
07b3e1540b4b
slug
optimization-theory-and-applications-3-07b3e1540b4b
url
https://medium.com/@rmdi115/optimization-theory-and-applications-3-07b3e1540b4b
canonical_url
https://medium.com/@rmdi115/optimization-theory-and-applications-3-07b3e1540b4b
author_url
https://medium.com/@rmdi115
status
ok
fetched_at
2026-06-15 20:49:13