Correlation Calculator
Pearson’s r from paired lists — direction, strength band, and a significance test on the association — with the live x² demo showing the one shape linear r is built to miss.
Correlation Calculator
Results recalculate instantly on every keystroke. Nothing you type is transmitted.
What this result does not account for
- Linear association only — the x² case is the boundary
- No rank (Spearman) variant — monotone-but-nonlinear data need it
In short: For x = 1..8 against y = 52, 55, 58, 61, 66, 64, 70, 72: Sxy = 119, Sxx = 42, Syy = 349.5, so r = 119/√(42×349.5) = 0.982197 — a very strong positive linear association. Its significance: t = r√(n−2)/√(1−r²) = 12.807304 on df = 6, p = 1.392e-5, so the association is no coincidence even at 1%. And r² = 0.964711 says the linear story explains 96.471149% of the variance. The boundary demo: for x = −3..3 and y = x², dependence is perfect yet r computes to exactly 0 — linear r sees one shape, not all of them.
Formula
r = Sxy / √(Sxx·Syy) · t = r√(n−2)/√(1−r²) · df = n − 2
The significance t is computed from r directly — and it equals the Regression page’s slope t on the same data, an identity you can check across the two pages.
Worked Example
- Paste paired lists — same count, 3 to 500 pairs.
- Read r with its band and direction.
- Check the significance t and p before trusting the band.
- Look at the x² card, then plot your own data before quoting r.
Defaults: r = 0.982197, very strong positive; t = 12.807304, df 6, p = 1.392e-5; r² = 0.964711. The x² demo: r = 0 EXACTLY for y = x² on −3..3.
Strengths & Limits Of This Model
Where this engine is strong
- Significance test attached, not an untested number
- The linear-only boundary demonstrated live on-page
Where it stops
- No Spearman/Kendall alternatives
- No outlier diagnostics beyond the residual story on the Regression page
Practical Use Cases
Measurement agreement
do two instruments move together?
Feature screening
which inputs travel with the outcome?
Coursework
r with its test, not just the number
Methodology & Editorial Standards
Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.
This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.
Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.
Correlation Calculator — 8 Expert FAQs
8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.
Does r = 0 mean no relationship?
It means no LINEAR relationship — and the x² card is the standing counterexample: y = x² on −3..3 is perfectly determined by x, and r computes to exactly 0. Correlation measures one shape. Plot the data; the scatterplot is the test r cannot replace.
Does correlation prove causation?
No, and the band card refuses the word. r measures co-movement. The cause could run either way, or through a third variable moving both — every observational r sits inside that ambiguity, however large it is.
Why test r at all — isn’t 0.98 obviously big?
Because “big” depends on n. With n = 4 even r = 0.95 is easily chance; with n = 400 r = 0.2 is already a coincidence too far. The t test prices the sample size in — this page’s default lands p = 1.392e-5 because n = 8 AND r = 0.982197 together.
What moves r — scale, outliers, or restriction?
Not scale: r is invariant under relabeling either axis (that is its point — standardized co-movement). Outliers and range restriction both drag it hard: one extreme pair can manufacture or destroy a strong r, and a narrow x range caps the r any true relationship can show.
Why must both lists vary?
Sxx or Syy zero means every value on an axis is identical — nothing co-moves with a constant. The page refuses that case rather than dividing by zero, with the flat axis named.
What is the boundary at r = ±1?
A perfect linear fit: the t statistic grows without bound and the p sits below machine precision. The page prints that boundary honestly instead of showing a fake test — perfect dependence needs no significance test, it needs a cause.
Why does correlation need at least three pairs?
Because the significance test spends two degrees of freedom estimating two means: df = n − 2 must be positive for a p to exist at all. Three pairs give df = 1 — technically computable, practically fragile; the band card means little until n is double digits.
Does swapping x and y change r?
No — Pearson r is symmetric: Sxy is the same product sum either way. That symmetry is also the deepest difference from regression: a correlation has no direction of explanation, while a regression line treats one variable as the responder and changes when you swap.