es

The material of this section is based on solvability section. Therefore, the reader is advised to study this section first before reading underlying material.

Least Square Approximation

When a linear system A x = b satisfies Corollary 5 of dot section, a perfect solution exists because b lies entirely within the column space 𝒞(A). However, in real-world applications—such as experimental data fitting and overdetermined systems (m > n)—measurement noise or unmodeled dynamics often pull the target vector b outside of 𝒞(A).

Because no true solution exists, our objective shifts from finding a perfect solution to finding an optimal approximation that minimizes the Euclidean norm of the residual vector r = b − A x.

The distance between b and any vector in the column space is minimized when the residual vector r = b − A x̂ is strictly orthogonal to the column space itself (r ⊥ 𝒞(A)).

figure x̂

By Theorem 15 of dot section, the orthogonal complement of the column space is the left null space (𝒞(A)⊥ = ker(A✶)). Therefore, the optimality condition dictates that the residual must live in the left null space:

\[ \mathbf{A}^{\ast} \mathbf{r} = 0 \qquad \Longrightarrow \qquad \mathbf{A}^{\ast} \left( \mathbf{b} - \mathbf{A}\,\hat{\bf x} \right) = 0 . \] Distributing the operator yields the Normal Equations:
\[ \mathbf{A}^{\ast} \mathbf{A} \, \hat{\bf x} = \mathbf{A}^{\ast} \mathbf{b} . \]
If the columns of A are linearly independent, the square matrix is strictly positive-definite and invertible, yielding a unique least-squares solution:
\[ \hat{\bf x} = \left( \mathbf{A}^{\ast} \mathbf{A} \right)^{-1} \mathbf{A}^{\ast} \mathbf{b} . \]
   
Example 1:

Consider three experimental data points (x, y): (1, 2), (2, 5), and (3, 5). We want to fit a straight line y = m x + c to these points. This overdetermined system yields three equations for two unknowns:

\[ \begin{aligned} 1m + c &= 2 \\ 2m + c &= 5 \\ 3m + c &= 5 \end{aligned} \implies \mathbf{A} = \begin{pmatrix} 1 & 1 \\ 2 & 1 \\ 3 & 1 \end{pmatrix}, \quad \mathbf{x} = \begin{pmatrix} m \\ c \end{pmatrix}, \quad \mathbf{b} = \begin{pmatrix} 2 \\ 5 \\ 5 \end{pmatrix} \]

Step 1: Compute the Normal Equation Matrix Aᵀ A and Vector Aᵀ b:

\[ \mathbf{A}^{\mathrm{T}}\mathbf{A} = \begin{pmatrix} 1 & 2 & 3 \\ 1 & 1 & 1 \end{pmatrix} \begin{pmatrix} 1 & 1 \\ 2 & 1 \\ 3 & 1 \end{pmatrix} = \begin{pmatrix} 14 & 6 \\ 6 & 3 \end{pmatrix} \] \[ \mathbf{A}^{\mathrm{T}}\mathbf{b} = \begin{pmatrix} 1 & 2 & 3 \\ 1 & 1 & 1 \end{pmatrix} \begin{pmatrix} 2 \\ 5 \\ 5 \end{pmatrix} = \begin{pmatrix} 27 \\ 12 \end{pmatrix} \]

Step 2: Solve the System AᵀA x̂ = Aᵀ b:

Inverting the 2 × 2 matrix yields:

\[ \begin{pmatrix} 14 & 6 \\ 6 & 3 \end{pmatrix}^{-1} = \frac{1}{(14)(3) - (6)(6)}\begin{pmatrix} 3 & -6 \\ -6 & 14 \end{pmatrix} = \frac{1}{6}\begin{pmatrix} 3 & -6 \\ -6 & 14 \end{pmatrix} \]

Multiplying to find the optimal parameter weights x̂:

\[ \hat{\mathbf{x}} = \frac{1}{6}\begin{pmatrix} 3 & -6 \\ -6 & 14 \end{pmatrix} \begin{pmatrix} 27 \\ 12 \end{pmatrix} = \begin{pmatrix} 1.5 \\ 1.0 \end{pmatrix} \]

Conclusion: The unique least-squares optimal line is y = 1.5x + 1.0. The projected optimal coordinates on the column space are A x̂ = [2.5, 4.0, 5.5]^{\mathrm{T}}$, yielding a minimized residual error vector r = [-0.5, 1.0, -0.5]ᵀ.

   ■
End of Example 1

Wolfram Mathematica Verification Scripts:

Option A: Direct Algebraic Solver via Normal Equations:

ClearAll[A, b, ATA, ATb, xHat, residual]; (* 1. Setup the design matrix and target vector *) A = {{1, 1}, {2, 1}, {3, 1}}; b = {2, 5, 5}; (* 2. Formulate normal equations components *) ATA = Transpose[A] . A; ATb = Transpose[A] . b; (* 3. Solve for parameters [m, c] *) xHat = Inverse[ATA] . ATb; Print["--- Normal Equations Solution ---"]; Print["Optimal Parameter Vector [m, c]: ", xHat // N]; Print["Orthogonality Check (A^T . residual should be zero vector): ", Transpose[A] . (b - A . xHat)];
Option B: Geometrical Visualization Script:
ClearAll[xData, yData, data, A, b, xHat, fitLine, plotPoints, plotLine]; (* 1. Define raw tracking coordinate arrays *) xData = {1, 2, 3}; yData = {2, 5, 5}; data = Transpose[{xData, yData}]; (* 2. Build the standard algebraic design matrix A and target vector b *) A = Transpose[{xData, ConstantArray[1, Length[xData]]}]; b = yData; (* 3. Solve the Normal Equations directly via standard matrix primitives *) xHat = Inverse[Transpose[A] . A] . Transpose[A] . b; (* 4. Extract parameters (Slope m is the first element, Intercept c is the second) *) m = xHat[[1]]; c = xHat[[2]]; fitLine[x_] := m * x + c; (* 5. Render elements independently with string parameters to avoid scope errors *) plotPoints = ListPlot[data, PlotStyle -> {Red, PointSize[0.025]}, PlotLegends -> {"Experimental Data"}]; plotLine = Plot[fitLine[x], {x, 0, 4}, PlotStyle -> {Blue, Thick}, PlotLegends -> {"Least-Squares Line"}]; (* 6. Output composite graphic layout *) Show[plotLine, plotPoints, AxesLabel -> {"x", "y"}, PlotLabel -> "Least-Squares Linear Curve Fit Verification", Frame -> True, GridLines -> Automatic]
Figure 1.1: The least square approximation.