- What is Linear Regression?
Linear Regression is a supervised learning algorithm used to predict a continuous numerical value.
It learns the relationship between an input feature and an output target by fitting a straight line through the observed training data.
If you know how much the input variable changes, the model discovers the exact rate of change so you can predict the outcome for new, unseen data points.
- The Core Mental Model & Visual Intuition
Suppose we collect real-world data points comparing house sizes to sale prices. On a two-dimensional grid, they scatter like an elongated cloud:
Feature Space
Data Point (x1, y1)
Data Point (x2, y2)
Data Point (x3, y3)
Fitted Straight Line (y_hat = wx + b)
Minimizes Total Squared Residuals
Because real data contains noise, the points never align perfectly on a line. The algorithm finds the single straight line that passes through the dense center of that cloud.
- How Does It Make a Prediction?
Linear Regression uses the standard slope-intercept algebraic formula:
= input feature (e.g., house square footage)= predicted value (e.g., estimated price)= weight or slope (rate of change per unit of )= bias or y-intercept (baseline value when )
Step-by-Step Worked Example:
Suppose our fitted model discovers and :
If an applicant has years of experience:
The data pipeline runs through the following sequence:
Input: x
Multiply by w
Add bias b
Prediction: y_hat
- How Does It Learn the Best Line?
The model starts with initial guesses for and , then iteratively refines them through an optimization loop:
- Initialize weights w and b
- Compute Predictions y_hat
- Calculate Loss via MSE
- Compute Gradients
- Update w and b via Learning Rate
Loss Converged?
No → back to Compute Predictions y_hat
Yes → Optimal Parameters Found
The standard optimization metric used is Mean Squared Error (MSE):
- It eliminates negative signs so errors in opposite directions do not cancel out.
- It penalizes large deviations much more severely than small ones.
- What Do Weight and Bias Mean Visually?
In the equation:
Controls tilt, steepness, and direction
Shifts line vertically along Y-axis
- Ordinary Least Squares vs. Gradient Descent
There are two primary mathematical approaches for finding the optimal weights:
Optimization Choice
Small to Medium Datasets
Ordinary Least Squares (Analytical)
Solves directly via Normal Equation:
beta = (X^T X)^(-1) X^T y
Large / Streaming Datasets
Gradient Descent (Iterative)
Updates weights step-by-step:
w = w - alpha * dLoss/dw
- Key Assumptions & Boundary Conditions
Linear regression produces reliable estimates when the following core assumptions are satisfied:
| Assumption | Description | What Happens If Violated? |
|---|---|---|
| Linearity | The link between and is straight and additive | Model underfits nonlinear curves |
| Homoscedasticity | Residual variance is uniform across all values of | Confidence intervals and p-values become invalid |
| Independence | Errors are uncorrelated across observations | High false-positive rates on time-series data |
| Normality | Residuals follow a standard normal distribution | Hypothesis testing results become inaccurate |
| No Multicollinearity | Input features are not highly correlated with each other | Parameter weights fluctuate wildly |
- Regression vs. Classification
The key distinction lies in the nature of the target variable:
| Dimension | Linear Regression | Classification (e.g., Logistic Regression) |
|---|---|---|
| Target Type | Continuous numerical | Discrete categories / classes |
| Example Outputs | House price, battery life, salary | Spam / Not Spam, Benign / Malignant |
| Output Range | Probabilities bounded in | |
| Function Shape | Straight line () | S-shaped curve () |
| Loss Function | Mean Squared Error (MSE) | Binary Cross-Entropy (Log Loss) |
- Simple vs. Multiple Linear Regression
Simple Linear Regression (One input variable):
Multiple Linear Regression (Multiple input variables):
For instance, estimating a house price rarely depends on square footage alone; it also depends on number of bedrooms, distance to schools, and property age. Each feature receives its own distinct weight .
- Key Takeaways
Linear Regression models the straight-line relationship between inputs and a continuous target by finding parameters that minimize the sum of squared prediction errors.
- What to Learn Next
Linear Regression
Cost Functions
Gradient Descent
Regularization (Ridge/Lasso)
Logistic Regression