Back to Roadmap
ML Fundamentals 5 min read

Linear Regression

A complete step-by-step intuitive and mathematical guide to Linear Regression with visual diagrams and worked examples.

  1. What is Linear Regression?

Linear Regression is a supervised learning algorithm used to predict a continuous numerical value.

It learns the relationship between an input feature XX and an output target YY by fitting a straight line through the observed training data.

For example:
Years of Experience → Monthly Salary
House Size (m2m^2) → Market Price ($)

If you know how much the input variable changes, the model discovers the exact rate of change so you can predict the outcome for new, unseen data points.

  1. The Core Mental Model & Visual Intuition

Suppose we collect real-world data points comparing house sizes to sale prices. On a two-dimensional grid, they scatter like an elongated cloud:

Feature Space

Data Point (x1, y1)

Data Point (x2, y2)

Data Point (x3, y3)

↓

Fitted Straight Line (y_hat = wx + b)

↓

Minimizes Total Squared Residuals

Because real data contains noise, the points never align perfectly on a line. The algorithm finds the single straight line that passes through the dense center of that cloud.

Residual (Error): The vertical distance between an actual observed point yiy_i and the line's estimate y^i\hat{y}_i.
Core Objective: Shift and tilt the line until the cumulative sum of these squared residuals is minimized.

  1. How Does It Make a Prediction?

Linear Regression uses the standard slope-intercept algebraic formula:

y^=wx+b\hat{y} = wx + b
Where:
  • xx = input feature (e.g., house square footage)
  • y^\hat{y} = predicted value (e.g., estimated price)
  • ww = weight or slope (rate of change per unit of xx)
  • bb = bias or y-intercept (baseline value when x=0x = 0)

Step-by-Step Worked Example:

Suppose our fitted model discovers w=900w = 900 and b=2,100b = 2{,}100:

Salary=(900×Experience)+2,100\text{Salary} = (900 \times \text{Experience}) + 2{,}100

If an applicant has x=4x = 4 years of experience:

y^=(900)(4)+2,100=3,600+2,100=$5,700\hat{y} = (900)(4) + 2{,}100 = 3{,}600 + 2{,}100 = \$5{,}700

The data pipeline runs through the following sequence:

Input: x

→

Multiply by w

→

Add bias b

→

Prediction: y_hat

  1. How Does It Learn the Best Line?

The model starts with initial guesses for ww and bb, then iteratively refines them through an optimization loop:

  1. Initialize weights w and b
↓
  1. Compute Predictions y_hat
↓
  1. Calculate Loss via MSE
↓
  1. Compute Gradients
↓
  1. Update w and b via Learning Rate

Loss Converged?

No → back to Compute Predictions y_hat

Yes → Optimal Parameters Found

The standard optimization metric used is Mean Squared Error (MSE):

MSE=1n∑i=1n(yi−y^i)2MSE = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2
Squaring the errors achieves two critical objectives:
  1. It eliminates negative signs so errors in opposite directions do not cancel out.
  2. It penalizes large deviations much more severely than small ones.

  1. What Do Weight and Bias Mean Visually?

In the equation:

y^=wx+b\hat{y} = wx + b
Weight (w): Slope

Controls tilt, steepness, and direction

Bias (b): Y-Intercept

Shifts line vertically along Y-axis

Weight (ww): Controls the angle/slope of the line. If w=900w = 900, each additional unit of xx increases y^\hat{y} by 900900. A negative ww indicates that y^\hat{y} drops as xx rises.
Bias (bb): The point where the line crosses the vertical axis when x=0x = 0. Without a bias term, the model would be constrained to always pass through (0,0)(0, 0).

  1. Ordinary Least Squares vs. Gradient Descent

There are two primary mathematical approaches for finding the optimal weights:

Optimization Choice

Small to Medium Datasets

Ordinary Least Squares (Analytical)

Solves directly via Normal Equation:
beta = (X^T X)^(-1) X^T y

Large / Streaming Datasets

Gradient Descent (Iterative)

Updates weights step-by-step:
w = w - alpha * dLoss/dw

1. Normal Equation (Closed-form): Computes the exact minimum directly in one matrix operation. Very fast for smaller datasets, but matrix inversion scales with O(p3)\mathcal{O}(p^3) complexity as features grow.
2. Gradient Descent (Iterative): Takes gradual downhill steps along the convex error surface. Scales effectively to millions of samples and dimensions.

  1. Key Assumptions & Boundary Conditions

Linear regression produces reliable estimates when the following core assumptions are satisfied:

AssumptionDescriptionWhat Happens If Violated?
LinearityThe link between XX and YY is straight and additiveModel underfits nonlinear curves
HomoscedasticityResidual variance is uniform across all values of XXConfidence intervals and p-values become invalid
IndependenceErrors are uncorrelated across observationsHigh false-positive rates on time-series data
NormalityResiduals follow a standard normal distributionHypothesis testing results become inaccurate
No MulticollinearityInput features are not highly correlated with each otherParameter weights fluctuate wildly

  1. Regression vs. Classification

The key distinction lies in the nature of the target variable:

DimensionLinear RegressionClassification (e.g., Logistic Regression)
Target TypeContinuous numericalDiscrete categories / classes
Example OutputsHouse price, battery life, salarySpam / Not Spam, Benign / Malignant
Output Range(−∞,+∞)(-\infty, +\infty)Probabilities bounded in [0,1][0, 1]
Function ShapeStraight line (y^=wx+b\hat{y} = wx + b)S-shaped curve (σ(wx+b)\sigma(wx + b))
Loss FunctionMean Squared Error (MSE)Binary Cross-Entropy (Log Loss)

  1. Simple vs. Multiple Linear Regression

Simple Linear Regression (One input variable):

y^=wx+b\hat{y} = wx + b

Multiple Linear Regression (Multiple input variables):

y^=w1x1+w2x2+⋯+wnxn+b\hat{y} = w_1 x_1 + w_2 x_2 + \dots + w_n x_n + b
In vector dot-product notation:
y^=wTx+b\hat{y} = \mathbf{w}^T \mathbf{x} + b

For instance, estimating a house price rarely depends on square footage alone; it also depends on number of bedrooms, distance to schools, and property age. Each feature receives its own distinct weight wiw_i.

  1. Key Takeaways

Linear Regression models the straight-line relationship between inputs and a continuous target by finding parameters that minimize the sum of squared prediction errors.
The fundamental equation:
y^=wx+b\boxed{\hat{y} = wx + b}
The learning loop:
Data⟶Prediction⟶Loss (MSE)⟶Update Parameters⟶Trained Model\text{Data} \longrightarrow \text{Prediction} \longrightarrow \text{Loss (MSE)} \longrightarrow \text{Update Parameters} \longrightarrow \text{Trained Model}

  1. What to Learn Next

Linear Regression

→

Cost Functions

→

Gradient Descent

→

Regularization (Ridge/Lasso)

→

Logistic Regression