Statistics

Hypothesis Test Calculator

A claim about a mean, put on trial: state H₀, feed the summary stats, and read the verdict — statistic, p, and the decision at your α, with the reasoning printed rather than implied.

Hypothesis Test Calculator

Results recalculate instantly on every keystroke. Nothing you type is transmitted.

Sample
The claim
The verdict
—
The statistic—
The p-value—
In plain language—
The reasoning, stated—

What this result does not account for

  • One-sample t on a mean only — proportions, two samples and paired designs are out of scope
  • Assumes roughly normal data or n large enough for the mean to be
● Zero-Server Execution Updated 11 Aug 2026 Reviewed by Sana Khalid IEEE-754 Double Precision

In short: Testing H₀: μ = 50 against μ ≠ 50 with n = 16, x̄ = 52, s = 5: the statistic is t = (52−50)√16/5 = 1.6 on df = 15 — the same figure the T Score page prints — and the two-sided p is 0.130445. Since 0.130445 ≥ 0.05, the verdict is keep H₀: a world where the true mean is 50 produces sample means this far from 50 about 13% of the time, and that is not strange enough to convict it. Failing to reject is a finding about evidence, not proof the claim is true.

Formula

t = (x̄ − μ₀)·√n / s · · df = n − 1 · p = P(T as extreme as t | H₀)

p is computed from the regularized incomplete beta (continued fractions to machine precision) — the forward tail area the T Score page refuses to fake.

Worked Example

  1. Enter the summary stats and the claimed mean.
  2. Pick the alternative: two-sided, or a direction.
  3. Set α — the evidence bar the p must clear.
  4. Read statistic, p, and the verdict sentence together.

n = 16, x̄ = 52, s = 5, μ₀ = 50, two-sided, α = 5%: t = 1.6, df = 15, p = 0.130445 → keep H₀. One-sided (greater): p = 0.065223 — still short of 0.05.

Strengths & Limits Of This Model

Where this engine is strong

  • Verdict sentence worded against the classic misreadings
  • Two-sided and directed alternatives priced from the same statistic

Where it stops

  • No paired or two-sample designs
  • No power analysis — a kept H₀ is not evidence of equivalence

Risk & accuracy notice. A kept null is the most misquoted result in applied statistics. This page reports the verdict and the p together precisely so the sentence cannot be shortened into a claim the arithmetic never made.

Practical Use Cases

Coursework

the full t trial, verdict included

Process checks

has the mean moved off target?

Audit sampling

does the book value survive the sample?

Methodology & Editorial Standards

Computation runs in IEEE-754 double precision at full internal precision; rounding to two decimal places occurs strictly at the display layer, so no cumulative drift enters the result. All monetary outputs use accounting presentation — grouped thousands, two decimals, negatives in parentheses — so figures can be transcribed directly into a model or working paper. Division-by-zero and out-of-domain inputs return an em-dash rather than a misleading number.

This engine was reconciled against an independent reference implementation and hand-verified for the worked example above before release. Our full five-stage review process is published on the About Us page.

Sana Khalid Principal Front-End Engineer · ApexConverter

Statistical inference, experiment design and numerical stability. Last reviewed: 11 August 2026.

Disclaimer. This calculator is provided for informational and modelling purposes only and does not constitute financial, tax, legal, medical, or engineering advice. Verify all figures with a qualified professional before acting on them.


Hypothesis Test Calculator — 8 Expert FAQs

8 analyst-written answers to the questions practitioners actually ask — optimised for voice and answer-engine retrieval.

What does “keep H₀” actually conclude?

Only that the evidence is not strange enough to reject the claim at your α. It is a verdict of not yet guilty, not innocence: p = 0.13 says data like yours arrive 13% of the time in a world where H₀ holds, which is ordinary, not exculpatory.

Why is the statistic t and not z?

Because s estimates σ. Dividing by an estimated standard deviation spends a degree of freedom and adds wobble, and the t distribution prices that wobble with heavier tails. The T Score page builds this exact statistic; this page carries it through to the probability the T Score page refuses to fake.

What does the p-value NOT mean?

It is not the probability H₀ is true. It is P(data this extreme | H₀) — a property of the data given the claim, not of the claim given the data. Reversing that conditional is the oldest misreading in statistics, and the verdict card is worded to survive it.

Why did my one-sided test succeed where two-sided failed?

Because you halved the p — 0.065223 against 0.130445 — by committing in advance to a direction. That commitment is a payment made BEFORE seeing the data; choosing the direction after is how a 0.065 becomes a 0.13 dressed as significance. The α you quote must match the alternative you declared.

Does rejecting H₀ measure importance?

No — it measures surprise. With enough n a trivial 0.1-unit shift becomes overwhelming evidence. Sizing the EFFECT is a different question with a different tool (the effect-size page, in build); the p only says whether the data argue at all.

Can I test a proportion or two groups here?

Not here. This page is strictly ONE sample against a claimed MEAN — the boundary is structural, not snobbery: proportions and two-group comparisons run different distributions (and the A/B page, in build, owns the two-proportion case). A chi-square page handles categorical counts; ANOVA handles three or more group means at once.

What α should I choose?

Whatever the decision is worth before seeing data — 5% is convention, not law. The honest rule: pick α and the alternative first, then let p fall where it falls. Choosing α after the p is how a 6% surprise gets laundered into a 5% finding.

My sample is tiny — does that change anything here?

The t distribution already prices smallness through df = n − 1: heavier tails, larger p for the same gap. What t cannot price is non-normal data at tiny n — the assumption card is the honest reading there, and more data beats more cleverness.

Related Statistics Engines