Statistical test calculator

Hypothesis Testing Calculator

This hypothesis testing calculator helps students check statistics work and researchers examine experimental evidence. Choose a supported procedure, enter raw observations or counts, and review the test statistic, p-value, assumptions, and a plain-language conclusion in one workspace.

Use the Hypothesis Testing Calculator

The workspace opens on One Sample Z-Test. Every module listed below has its own dedicated guide and the same tool already set to that method.

Interactive calculator workspace

Interactive tool

The calculator loads as you approach this section so the guide remains fast on mobile connections.

Example preview

One-sample mean with known σ

Inputs
n=10, x̄=502, μ₀=500, σ=4; two-sided Z test.
Representative result
z≈1.58 and p≈0.114.

Illustrative only. Load the interactive tool to enter your own values and review assumptions.

Choose a hypothesis-test family

Start with the question and study design, then use a topic guide to select the precise calculator.

P-value calculator from a test statistic

If you already have a test statistic, use its reference distribution to calculate the corresponding tail probability. Choose Standard Normal, Student’s t, Chi-Square, or F, then provide the degrees of freedom required by that distribution.

Example result

p = 0.049995790

Reference
Standard Normal (Z)
Tail definition
Two-sided probability
Lower tail
0.975002105
Upper tail
0.024997895

This is a reference-distribution tail probability conditional on the statistic and model you supplied. It does not verify how the statistic was derived, whether assumptions hold, or whether an effect is practically important.

Result summary

Direct p-value calculation

Conclusion

The supplied statistic corresponds to p = 0.049995790 for the selected two-sided probability. A reject/fail-to-reject decision requires a pre-specified significance level and a correctly derived statistic.

Key values

Statistic
1.96
P-value
0.049995790
Lower tail
0.975002105
Upper tail
0.024997895

Assumptions and limits

  • The statistic follows the selected reference distribution under the null model.
  • The tail direction was chosen before inspecting the result.

Method

The calculator evaluates the selected reference distribution CDF and combines tails according to the requested direction. It does not reconstruct or validate the originating hypothesis test.

Choose a supported module

Each module has its own canonical guide, complete server-rendered explanation, and compact tool already set to the method named in the URL.

One Sample Z-TestRun a one sample z-test when population standard deviation is known, then interpret the z statistic, p-value, tails, assumptions, and limits.One Sample T-TestRun a one sample t-test for a mean with unknown population spread, then interpret the t statistic, degrees of freedom, p-value, and assumptions.Two Sample Z-TestCompare two independent means when both population standard deviations are known, with correct tails, assumptions, examples, and interpretation.Two Sample T-TestCompare two independent means with pooled or Welch t methods, and understand variance choices, degrees of freedom, p-values, and limitations.Paired Sample T-TestTest the mean of matched or before-after differences, with correct pair order, assumptions, t statistic, p-value interpretation, and examples.One-Way ANOVACompare three independent group means with an omnibus F-test and supported post-hoc methods, with assumptions and careful result interpretation.Two-Way ANOVA with ReplicationTest two categorical main effects and their interaction with a balanced two-way ANOVA, complete ANOVA table, assumptions, and separate conclusions.Chi-Square Test of IndependenceAnalyze a 2 by 2 table with Pearson chi-square, expected counts, p-value, and Cramér’s V, while checking independence and sparse-cell limits.Chi-Square Goodness-of-Fit TestCompare observed and expected category frequencies, calculate chi-square, degrees of freedom and p-value, and inspect category contributions.F-Test for Equal VariancesCompare two independent variances with directional or two-sided F-tests, fixed sample order, normality cautions, examples, and limitations.One-Proportion Z-TestTest one sample proportion against a hypothesized value with a z statistic, Wilson confidence interval, effect size, and clear assumptions.Two-Proportion Z-TestCompare two independent sample proportions with a pooled z test, an unpooled confidence interval for the difference, and Cohen’s h effect size.Exact Binomial TestRun an exact binomial test for one proportion with true tail probabilities, a Clopper–Pearson interval, and guidance for small samples.Fisher’s Exact TestTest association in a 2×2 table exactly with Fisher’s method, ideal for small expected counts, with the sample odds ratio and clear interpretation.McNemar TestTest paired binary outcomes for change with McNemar’s procedure, using an exact version for few discordant pairs and chi-square otherwise.Mann-Whitney U TestCompare two independent samples with the Mann-Whitney U test: tie-corrected p-values, a rank-biserial effect size, and guidance on when ranks beat means.Wilcoxon Signed-Rank TestTest paired differences with the Wilcoxon signed-rank procedure: signed ranks, tie-corrected p-values, and a matched-pairs rank-biserial effect size.Sign TestRun the exact sign test on paired data using only difference directions, with binomial tail p-values and honest guidance about its low power.Kruskal-Wallis TestCompare several independent groups with the tie-corrected Kruskal-Wallis H test and follow up with Dunn-Bonferroni pairwise comparisons.Friedman TestTest three or more related treatments with the Friedman procedure: within-block ranks, a tie-corrected Q statistic, and Kendall’s W agreement.Pearson Correlation TestTest whether a linear correlation differs from zero, with the exact t reference, a Fisher-z confidence interval, and the r² share of variance.Spearman Correlation TestTest monotonic association with Spearman’s rank correlation: midranks, a t-approximation p-value, and a Fisher-z interval robust to outliers.Kendall Tau TestTest ordinal association with Kendall’s tau-b: concordant versus discordant pairs, tie corrections, and an asymptotic normal p-value.Levene TestCheck whether groups share a common variance with Levene’s mean-centered W statistic, the classical gate before pooled-variance analyses.Brown-Forsythe TestTest variance equality with the median-centered Brown-Forsythe statistic, the robust default for skewed or outlier-prone group data.Welch ANOVACompare group means without the equal-variance assumption using Welch’s F and Games-Howell pairwise follow-up with Satterthwaite degrees of freedom.Repeated Measures ANOVACompare within-subject condition means with subject variability removed and the Greenhouse-Geisser sphericity correction applied automatically.TOST Equivalence TestDemonstrate practical equivalence of two means with two one-sided Welch t tests and the interval-inside-bounds reading of the result.Power & Sample SizeSolve for the sample size that reaches a target power, or the power a fixed n delivers, across t-test and proportion designs with honest approximations.Confidence Interval CalculatorBuild z, t, Wilson, Clopper-Pearson, and chi-square intervals directly from summary statistics, with margins of error and method guidance.Critical Value LookupLook up exact critical values and rejection regions for the four classical reference distributions at any alpha and tail configuration.

Watch the explanation

2:36 min

The player loads only after you press play. You can also watch on YouTube.

Read the complete transcript

How to Choose a Statistical Hypothesis Test. Choose a hypothesis test before looking for a favorable p-value. State the population parameter, null value, alternative direction, and practically important effect first. Classify the outcome. A numerical mean, binary proportion, categorical count, rank association, correlation, or variance question requires a different statistical target. Then map the design: one sample, two independent groups, paired observations, three or more groups, repeated measures, or a factorial experiment. For one numerical mean, use Z only when population standard deviation is known independently. If spread is estimated from this sample, use a T procedure. For two numerical groups, independence versus pairing matters. Independent groups often use Welch's T; matched measurements are analyzed through within-pair differences. For three or more means, consider ANOVA, Welch ANOVA, or a repeated-measures method. The design and variance conditions decide among them. Binary and categorical outcomes follow other routes. Use proportion or exact binomial tests, Mc Nemar for paired binary data, and chi square or Fisher procedures for count tables.

Rank tests can address ordinal outcomes or fragile parametric assumptions. Pearson, Spearman, and Kendall target different forms of association; they are not interchangeable buttons. Check independence, distributional conditions, influential observations, expected counts, and the planned tail before calculating. A large sample does not repair a broken design. For a worked example, take ten observations with mean five hundred two, test a null mean of five hundred, and use independently known sigma four. The standard error is four divided by square root ten, or one point two six four nine. The Z statistic is one point five eight one one. The two-sided p-value is point one one three eight. At a predeclared point zero five threshold, this result does not cross the evidence cutoff. That is not the probability the null is true, and not proof of no important effect. If four were a sample standard deviation, use T with nine degrees of freedom. Report the effect, uncertainty, statistic, degrees of freedom, p-value, assumptions, and multiplicity plan. Choose and calculate supported tests free at Distri Scope dot com.

What Is a Hypothesis Test?

A hypothesis test compares sample evidence with a null hypothesis, which is a specific claim about a population parameter or data pattern. The alternative hypothesis states the difference, direction, or association the study is designed to detect. A test statistic expresses how far the observed result falls from the null model in reference-distribution units, and the p-value summarizes how unusual that result would be if the null model and the test assumptions were correct.

DistriScope combines a hypothesis test calculator, a t test calculator, and a p value calculator workflow for supported mean, ANOVA, categorical, and variance procedures. Before selecting one, identify the outcome type, number of groups, pairing or independence, and the parameter in the research question. A small p-value does not measure effect size, practical importance, data quality, or the probability that either hypothesis is true.

Before You Run a Hypothesis Test

Write down the statistical question, the unit of observation, and the quantity you want to estimate or explain before opening Statistical Hypothesis Test Calculator. Confirm where the values came from, what units they use, and whether repeated observations are independent. Preserve the original inputs and record every parameter, transformation, and option used in the workspace. This creates a reproducible trail and makes it easier to compare the result with another package.

Treat the graph and numerical output as evidence within a model, not as a substitute for the study design. If a conclusion changes when a plausible parameter or assumption changes, report that sensitivity. Clear documentation is part of statistical accuracy because it allows another person to understand what was calculated, test the same conditions, and identify where an interpretation may need revision. See the DistriScope methodology for formulas, numerical methods, and independent-verification guidance.

Hypothesis Testing Calculator Inputs Explained

Raw observations or counts

Enter numerical sample values for mean tests and ANOVA, paired values for a paired t test, contingency counts for independence, or observed and expected counts for goodness of fit. Keep units and category order consistent.

Sample mean and sample size

For mean tests, the calculator derives the sample mean (x̄) and sample size (n) from the raw observations. These values are outputs, not separate fields, which reduces transcription errors.

Null or hypothesized value

The population mean (μ₀) or hypothesized mean difference defines the null claim for the relevant test. For goodness of fit, the expected counts play the corresponding reference role.

Standard deviation

A z test requires a population standard deviation (σ) known independently of the current sample. A t test derives the sample standard deviation (s) from the observations because σ is unknown.

Significance level (α)

Alpha is the decision threshold chosen before calculation, commonly 0.05. It controls the planned false-positive rate under the null model; it is not the probability that the conclusion is wrong.

One-tailed or two-tailed alternative

Choose the direction from the written alternative hypothesis before viewing the data. Directional options apply only where the procedure supports them; ANOVA and chi-square statistics use their defined right-tail reference probabilities.

Method-specific options

Some procedures require additional choices, such as Welch versus pooled variance, post-hoc comparisons, factor levels, expected frequencies, or numerator and denominator sample order. Use the module guide for the exact input definition.

How to Use This Hypothesis Test Calculator

  1. 1Write the null hypothesis and alternative hypothesis, then choose α and any directional tail before inspecting the result.
  2. 2Select the test that matches the outcome type, number of groups, pairing, independence, and whether population variability is known.
  3. 3Enter the raw observations, paired values, or counts requested by that module. Add only the procedure-specific parameters shown in the form.
  4. 4Run the calculation and review the derived sample statistics, test statistic, degrees of freedom when applicable, critical value, and p-value.
  5. 5Compare p with α, check assumptions, and report the numerical result with its practical context. Preserve the inputs so another person can reproduce it.

Step-by-Step Example: One-Sample t Test

Test H₀: μ = 9 against H₁: μ ≠ 9 at α = 0.05 using the five observations 8, 9, 10, 11, and 12. Their mean is x̄ = 10. The squared deviations from 10 sum to 10, so the sample standard deviation is s = √[10 / (5 − 1)] = 1.5811. The standard error is s / √n = 1.5811 / √5 = 0.7071. Therefore, t = (x̄ − μ₀) / (s / √n) = (10 − 9) / 0.7071 = 1.4142, with df = n − 1 = 4. A standard Student’s t table uses the 0.975 quantile for this two-sided test and gives the critical value ±2.7764. The exact two-sided probability is p = 2P(T₄ ≥ 1.4142) = 0.2302.

Interpretation

Because p = 0.2302 is greater than α = 0.05, fail to reject H₀. The sample does not provide sufficient evidence that the population mean differs from 9. This result does not prove that μ equals 9; it reflects limited evidence from five observations under the independence and normality assumptions.

Understanding the Results: Test Statistic, Degrees of Freedom, and p-Value

  • Read the sign and magnitude of the test statistic in the direction defined by the sample order and alternative. For two-sided tests, both tails contribute to the p-value.
  • Report degrees of freedom for t, F, chi-square, and ANOVA results because they determine the reference curve used for the probability calculation.
  • Compare p with the preselected α. If p ≤ α, reject H₀ for the stated alternative; if p > α, fail to reject H₀ rather than declaring it true.
  • Interpret the observed effect, uncertainty, assumptions, and study design alongside the threshold decision. Multiple tests require a planned correction or a clearly limited family of comparisons.

Null and Alternative Hypotheses

Null and alternative hypotheses

The null hypothesis (H₀) supplies the reference claim used to calculate the statistic. The alternative hypothesis (H₁ or Hₐ) states the planned difference, direction, or association. Write both before analyzing the data.

Test statistic and degrees of freedom

A z, t, F, or chi-square statistic compares the observed pattern with variation expected under H₀. Degrees of freedom identify the correct reference distribution for tests that estimate information from the sample.

P-value

The p-value is the probability, assuming H₀ and the other model conditions, of obtaining a result at least as incompatible with H₀ as the observed one. It is not the probability that H₀ is true.

Statistical versus practical significance

Crossing α is a decision under a specified error-rate rule. It does not show that an effect is large or useful. Report an effect estimate and confidence interval when the selected tool or independently validated software provides them.

Hypothesis Testing Calculator FAQs

Does p > 0.05 prove the null hypothesis?

No. It means the observed result did not cross that selected threshold under the null model and assumptions. Report that you failed to reject H₀; do not claim that H₀ is true or that the effect is exactly zero.

When should I use a z test instead of a t test?

Use a z test for a mean only when the population standard deviation is known independently of the current sample. Use a t test when the population standard deviation is unknown and the sample standard deviation estimates it. A large sample alone does not make σ known.

Can I choose a one-tailed test after seeing the data?

No. Choose the direction from the research question before examining the result. Switching tails after seeing the sign of the statistic changes the planned false-positive rate and makes the reported decision unreliable.

Why might another calculator return a slightly different p-value?

Check whether both tools use the same test, tail, variance assumption, sample order, degrees of freedom, and rounding. Small display differences can come from numerical precision; larger differences usually indicate different inputs or method settings.

Limitations and verification

  • The calculator supports selected procedures and cannot diagnose every violation or design complication.
  • Rounded or summary-only inputs can prevent checks that are possible with raw observations.
  • A non-significant result is not proof of no effect, and a significant result is not proof of causation.
  • Confirm high-stakes results with validated software, the original data, and qualified statistical review.

Educational information. Last reviewed 2026-08-09. Calculations should be independently verified for consequential decisions.