Interactive module guide

Two Sample Z Test Calculator — Free & Interactive

This two sample z test calculator compares two independent population means when both population standard deviations are known.

Open calculator

Free · No sign-up · Calculations stay in your browser

A two sample z-test compares the means of two independent numerical populations when both population standard deviations are known from information external to the samples.

It tests a pre-specified mean difference, implemented here as zero.

This condition is uncommon: large samples do not turn sample standard deviations into known population values, and an independent two sample t-test is normally used when either σ must be estimated.

Use the two sample z test calculator for two independent numerical samples only when both population standard deviations are known outside the current data. Confirm that observations do not overlap across groups, measurement units match, and the planned alternative direction was chosen before inspecting the sample difference.

When to use this two sample z test calculator

Use it when

  • Use the procedure for two independently sampled or independently assigned groups, a numerical outcome, and genuinely known σ₁ and σ₂ that apply to the current populations. Each row should belong to one group only.
  • Choose a two-sided alternative for any difference, a right-sided alternative when μ₁−μ₂>0 is the planned claim, or a left-sided alternative when μ₁−μ₂<0 is planned. The order of samples determines the direction.
  • Normal populations support the exact reference. With sufficiently informative independent samples, a normal approximation for the difference in means may be reasonable under finite-variance conditions, but process skew, dependence, and outliers still require review.

Choose another method when

  • Do not use this test for the same subjects measured twice; pairwise dependence should be modeled with a paired test. Do not use it for proportions merely because a normal approximation is familiar.
  • Avoid supplying sample standard deviations in the population-standard-deviation fields. That discards estimation uncertainty and can yield p-values that are too small.

Interactive tool

The calculator loads as you approach this section so the guide remains fast on mobile connections.

Example preview

Two calibrated production lines

Inputs
x̄₁=102, x̄₂=99, n₁=n₂=5, σ₁=3, σ₂=4; two-sided Z test.
Representative result
z≈1.342 and p≈0.18.

Illustrative only. Load the interactive tool to enter your own values and review assumptions.

Watch the explanation

3:24 min

The player loads only after you press play. You can also watch on YouTube.

Read the complete transcript

The Two-Sample z-Test. Are two calibrated production lines different enough to rule out chance?

A two sample z-test compares the means of two independent numerical populations when both population standard deviations are known from information external to the samples. It tests a pre-specified mean difference, implemented here as zero. This condition is uncommon, large samples do not turn sample standard deviations into known population values, and an independent two sample t-test is normally used when either sigma must be estimated.

The null implemented is H naught, mu one minus mu two equals 0. The numerator x bar one minus x bar two measures the observed difference, and the denominator combines independent known-variance contributions. Covariance would be needed if samples were paired or otherwise related. Under H naught and the normal model, z follows the standard normal distribution. The p-value is a tail probability for statistics at least as extreme under H naught, not the probability that the groups have equal means or that randomization failed. Use the procedure for two independently sampled or independently assigned groups, a numerical outcome, and genuinely known sigma one and sigma two that apply to the current populations. Each row should belong to one group only. Interpret the sign relative to Sample 1 minus Sample 2 and report both sample means so the direction is transparent. The z magnitude shows how large the difference is relative to its known standard error. A rejection supports evidence for the stated directional or two-sided alternative at alpha; it does not measure practical value or establish causation without a randomized design. Failure to reject leaves the difference unresolved rather than demonstrating equality.

Two independent calibrated lines produce the same component. Historical validation supplies sigma one equals 3 grams and sigma two equals 4 grams. Samples are Line 1, 101, 103, 100, 102, 104 and Line 2, 98, 100, 99, 101, 97. A two-sided comparison at alpha equals 0.05 was planned. Enter the two samples and known standard deviations. The sample means are 102 and 99 grams, so the observed difference is 3 grams. The known standard error is the square root of the sum of two terms: 3 squared divided by 5, plus 4 squared divided by 5, giving square root of 5, about 2.236. Therefore z approximately 1.342 and the two-sided p-value is around 0.18. Because p exceeds 0.05, the test fails to reject equal population means. The three-gram observed gap is not sufficiently extreme relative to the known variability and small sample sizes.

The experiment provides insufficient evidence, at the planned threshold, that the two line means differ. It does not show the lines are equivalent and does not assess whether a three-gram difference is operationally acceptable. The external sigma values and independence between sampled units must remain applicable.

Try it free at Distri Scope dot com.

How to read the result

Interpret the sign relative to Sample 1 minus Sample 2 and report both sample means so the direction is transparent. The z magnitude shows how large the difference is relative to its known standard error.

A rejection supports evidence for the stated directional or two-sided alternative at α; it does not measure practical value or establish causation without a randomized design. Failure to reject leaves the difference unresolved rather than demonstrating equality.

How to use the two sample z test calculator

Enter raw values for Sample 1 and Sample 2, keeping units and inclusion rules consistent. The calculator derives each mean and sample size; unequal sample sizes are permitted.

Enter σ₁ and σ₂ as known individual-level population standard deviations. The standard error is sqrt(σ₁²/n₁+σ₂²/n₂). Swapping samples changes the sign of z but not a correctly computed two-sided p-value.

Set alpha and tail before calculation. If lower outcomes are better in the application, “greater” is not automatically “better”; write the directional parameter claim in symbols to prevent a sign error.

Formula, hypotheses, and assumptions

z=(xˉ1xˉ2)Δ0σ12n1+σ22n2,Δ0=0 herez = \frac{(\bar{x}_1 - \bar{x}_2) - \Delta_0}{\sqrt{\frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}}}, \quad \Delta_0=0 \text{ here}

Conditions to review

  • Population standard deviations are known
  • Samples are randomly selected
  • Populations are normally distributed or n₁, n₂ ≥ 30

Calculator parameters

  • Significance Level (α): default 0.05.
  • Population Std Dev 1 (σ₁): default 1.
  • Population Std Dev 2 (σ₂): default 1.
  • Test Type: default Two-tailed.

What the method is doing

The null implemented is H₀:μ₁−μ₂=0. The numerator x̄₁−x̄₂ measures the observed difference, and the denominator combines independent known-variance contributions. Covariance would be needed if samples were paired or otherwise related.

Under H₀ and the normal model, z follows the standard normal distribution. The p-value is a tail probability for statistics at least as extreme under H₀, not the probability that the groups have equal means or that randomization failed.

Worked example: two calibrated production lines

Two independent calibrated lines produce the same component. Historical validation supplies σ₁=3 grams and σ₂=4 grams. Samples are Line 1: 101, 103, 100, 102, 104 and Line 2: 98, 100, 99, 101, 97. A two-sided comparison at α=0.05 was planned.

  1. 1Enter the two samples and known standard deviations. The sample means are 102 and 99 grams, so the observed difference is 3 grams.
  2. 2The known standard error is sqrt(3²/5+4²/5)=sqrt(5)≈2.236. Therefore z≈1.342 and the two-sided p-value is around 0.18.
  3. 3Because p exceeds 0.05, the test fails to reject equal population means. The three-gram observed gap is not sufficiently extreme relative to the known variability and small sample sizes.

Interpretation

The experiment provides insufficient evidence, at the planned threshold, that the two line means differ. It does not show the lines are equivalent and does not assess whether a three-gram difference is operationally acceptable. The external σ values and independence between sampled units must remain applicable.

Common mistakes

  • Calling two groups independent because they are stored in separate columns is not enough. Matching, shared batches, repeated operators, or common time shocks create dependence.
  • Switching sample order while keeping a directional alternative unchanged reverses the hypothesis. Write μ₁−μ₂ explicitly.
  • Known σ values from an obsolete process or different population are not known for the current comparison. Their provenance belongs in the report.

Limits and independent validation

The output provides a known-σ confidence interval and standardized effect, but not an equivalence test or adjustment for covariates and clustering. It assumes the supplied population spreads are fixed; prospective power is a separate planning calculation.

The procedure compares means only. Different variances, shapes, or tails may matter even when mean evidence is weak.

Before using the result

  • Confirm independent assignment or sampling, comparable measurement, and the external source of σ₁ and σ₂. Plot each group and review time or batch order.
  • Recompute both means, the combined known standard error, z, and the appropriate normal tail independently. State sample order, test direction, α, and the limited conclusion.
  • Perform a provenance audit for each population standard deviation: name the dataset or calibration that supplied it, its measurement conditions, and the date or process version it represents. Repeat the calculation with plausible alternative spread values to show how strongly the conclusion depends on treating those quantities as fixed and transportable.

See every option in the Statistical Hypothesis Test Calculator or review the DistriScope methodology. Educational information; last reviewed 2026-08-09. Verify consequential calculations independently.