Regression Calculator
Least squares on paired data: the line, the prediction with the extrapolation caveat printed beside it, every residual, the variance partition proved live, and the slope’s own significance test.
Regression Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- Simple (one-predictor) regression only
- No prediction INTERVAL — the point prediction is exact, its uncertainty band needs t quantiles this site refuses to fake
In short: For x = 1..8 and y = 52, 55, 58, 61, 66, 64, 70, 72 the least-squares line is ŷ = 49.5 + 2.833333x: slope b = 119/42 = 2.833333, intercept a = 62.25 − 2.833333×4.5 = 49.5. Predicting at x = 9 lands exactly on 75 — one step past your data (max x = 8), which the page says out loud. The residuals end −0.333333, −0.166667, 0, 0.166667, 2.333333, −2.5, 0.666667, −0.166667 — the 6th point sits 2.5 low. The partition closes: SST 349.5 = SSR 337.166667 + SSE 12.333333, R² = 0.964711, and the slope’s t = 12.807304 with p = 1.392e-5 — identical to the Correlation page’s r test, because they are the same question.
Formula
b = Sxy/Sxx · a = ȳ − b·x̄ · SE(b) = √((SSE/(n−2))/Sxx) · t = b/SE(b)
Least squares minimizes the sum of squared residuals — equivalently, it makes the residuals orthogonal to every column. The partition SST = SSR + SSE is shown closing on your data.
Worked Example
- Paste paired lists — same count, 3 to 500 pairs.
- Optionally set a prediction point x.
- Read the line, the prediction with its caveat, and the residuals.
- Check the partition and the slope test before quoting R².
Defaults: ŷ = 49.5 + 2.833333x; prediction at 9 is 75 exactly (one step past max x 8); residuals close with the 6th point 2.5 low; SST 349.5 = SSR 337.166667 + SSE 12.333333; SE(b) 0.221228, t = 12.807304, p = 1.392e-5.
Strengths & Limits Of This Model
Where this engine is strong
- Residuals and partition printed, not summarized away
- Extrapolation distance stated at the prediction
Where it stops
- No robust or weighted fitting
- No diagnostic plots — the residual list stands in
Practical Use Cases
Dosing and scaling curves
the line under the last run
Budget models
cost as a linear function of units
Coursework
every intermediate printed, nothing hidden
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
Regression Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
What is least squares actually minimizing?
The sum of SQUARED vertical residuals — and the solution has a geometric signature: the residual vector ends up orthogonal to x itself, which is why the fitted line always passes through (x̄, ȳ). Squaring penalizes big misses over small ones; that choice, not the arithmetic, is the assumption.
Why does the prediction card mention extrapolation?
Because the line is a claim fitted INSIDE your x range. At x = 9 with data ending at 8, the arithmetic is exact and the warrant is thinner — nothing in the data says the slope continues past where you measured. The page prints the distance outside the range rather than letting a confident number imply an untested region.
How are the Regression and Correlation pages different?
They share Sxy, and on the same data the slope’s t equals the correlation’s t exactly. The QUESTION differs: correlation grades the co-movement; regression fits the line and answers “what y at this x?” — a prediction machine, not a strength meter.
What does R² mean here?
The share of y’s variance the line carries: SSR over SST, which the partition card shows closing (337.166667 over 349.5 on the default). It equals r² for one predictor — the R Squared page (in build) will own the decomposition as its own question.
When is the slope test the wrong tool?
When the residuals are the story: heavy curvature, an outlier with leverage, or dependence between observations all sit under the t’s assumptions. The residuals card exists so those failures are visible before the p is quoted.
Can I fit more than one predictor?
Not here — this page is the two-variable world. The Multiple Linear Regression page runs the 3×3 normal equations, refuses collinear predictors, and prices each coefficient separately.
How large a residual should worry me?
Judge it against the spread of the others — one residual several times its siblings is a lever arm on the slope, and dropping it will move the line. The residual card names the largest so the question is impossible to skip; investigate the point before blessing the fit.
Why does the fitted line always pass through (x̄, ȳ)?
Because least squares makes the residuals orthogonal to x: the positive and negative residuals cancel exactly above and below the mean point. It is a provable identity, not a tendency — and the hero card states it so the fit can be sanity-checked in one glance.