Optimization Theory and Applications-3
Gradient of Quadratic Forms and the Matrix Product Rule
Optimization Theory and Applications-3
Gradient of Quadratic Forms
and the Matrix Product Rule



MATHEMATICAL FOUNDATIONS
Vectors, Norms, and Inner Products
We need a bit of a footing before we can have a meaningful conversation about gradients and quadratic forms. In this section we review the necessary vocabulary of linear algebra and real analysis that will support us along the entire development.

We write vectors as column vectors. Given x∈Rn, we have:

1.2 The Inner Product and Euclidean Norm



1.3 Matrices and Linear Maps


Symmetric matrices are at the heart of the theory of quadratic forms, as we shall see in §5.
1.4 Key Algebraic Identities
We will use these transpose rules repeatedly without proof:

2 PARTIAL DERIVATIVES
Partial Derivatives and Multivariate Differentiation
We shall use again and again The basic operation of calculus, differentiation, generalises naturally to functions of several variables. The main idea is to change one variable at a time , keeping all the others fixed . rules of transposition without proof
2.1 Functions of Several Variables
Let f:Rn→R be a scalar-valued function. We write f(x)=f(x1,x2,…,xn). The domain may be an open subset U⊆Rn.

Geometrically, ∂f/∂xi is the slope of the curve obtained by slicing the graph of f with the hyperplane that fixes all variables except xi.

2.2 Total Differentiability


The linear map L can always be represented by a row vector (the gradient), so that L(h)=∇f(a)⊤h.

2.3 Linear Approximation
Differentiability means f is well-approximated near a by the first-order Taylor expansion:

This is the cornerstone identity that connects local behavior of f to its derivatives.
3 THE GRADIENT
The Gradient: Definition, Geometry, and Properties
The gradient is the “multi-dimensional derivative” of a scalar function. It bundles all directional information into a single vector.

The gradient is the “multi-dimensional derivative” of a scalar function. It bundles all directional information into a single vector.
3.1 Geometric Meaning of the Gradient


3.2 Linearity of the Gradient

3.3 Gradient of a Linear Function

4 DIRECTIONAL DERIVATIVES
Directional Derivatives and the Gradient Connection
The partial derivative measures the rate of change along coordinate axes. But what about an arbitrary direction? The directional derivative answers this question.

4.1 Proof of Theorem 4.1

Therefore:

4.2 Maximum Directional Derivative
By the Cauchy–Schwarz inequality:
Equality holds when u=∇f(x)/‖∇f(x)‖. This confirms that the gradient direction gives the maximum rate of increase.


5 QUADRATIC FORMS
Quadratic Forms: Algebra and Geometry
"A quadratic form is a polynomial in which every term has total degree
exactly two — the natural generalization of to many variables."



5.2 WLOG: We May Assume A is Symmetric

Takeaway: Without loss of generality, we always assume A is symmetric when working with quadratic forms.
5.3 Geometric Interpretation

6 CORE DERIVATION
Gradient of a Quadratic Form: Full Derivation

6.2 Derivation Method 2: Perturbation (Definition-Based)
We compute each partial derivative ∂f/∂xk explicitly.
Expand the quadratic form in coordinates:


6.2 Derivation Method 2: Perturbation (Definition-Based)
This approach uses the definition of the gradient as the linear part of the perturbation f(x+h)−f(x).



7 SYMMETRIC MATRICES
Symmetric Matrices and Simplification
The symmetry of a matrix is not just a notational convenience but has deep algebraic and geometric consequences that motivate the entire theory of quadratic forms and optimisation.
7.1 Spectral Theorem

Therefore symmetric matrices always have real eigenvalues and an orthonormal eigenbasis. In the eigenvector coordinate system the quadratic form reduces to a sum of squares:


7.2 Gradient Formula — Symmetric A
For symmetric A=A⊤:

8 MATRIX PRODUCT RULE
The Matrix Product Rule for Gradients

8.1 Product Rule: Scalar × Vector Case




9 CHAIN RULE
Chain Rule in Matrix Calculus

9.1 Jacobian Matrix

9.2 Special Case: Linear Composition

9.3 Chain Rule for Quadratic Compositions



10 SECOND-ORDER STRUCTURE
The Hessian Matrix: Second-Order Derivatives


10.1 Hessian of a Quadratic Form


10.2 Second-Order Taylor Expansion

For a quadratic form f(x)=x⊤Ax, this expansion is exact (the remainder is zero), which makes quadratic forms the simplest non-linear functions.

11 MEAN VALUE THEOREM
Mean Value Theorem for Multivariate Functions
TBC
메타데이터
- post_id
- 07b3e1540b4b
- slug
- optimization-theory-and-applications-3-07b3e1540b4b
- url
- https://medium.com/@rmdi115/optimization-theory-and-applications-3-07b3e1540b4b
- canonical_url
- https://medium.com/@rmdi115/optimization-theory-and-applications-3-07b3e1540b4b
- author_url
- https://medium.com/@rmdi115
- status
- ok
- fetched_at
- 2026-06-15 20:49:13