Interactive module guide

Sign Test Calculator — Exact Paired Analysis

This sign test calculator evaluates paired differences using only their directions, with exact binomial p-values and no distributional assumptions.

Open calculator

Free · No sign-up · Calculations stay in your browser

The sign test is the most assumption-light paired test: it keeps only the direction of each paired difference and asks whether positives and negatives are equally likely, referring the count of positives to an exact Binomial(n, 1/2) distribution.

It requires no normality, no symmetry, and not even meaningful difference magnitudes—only trustworthy directions—and it pays for that robustness with lower power than the Wilcoxon signed-rank or paired t test.

This calculator drops tied pairs, reports the positive/negative split, and computes exact binomial tail p-values with no approximation at any sample size.

Use the sign test calculator when only the direction of each paired difference is trustworthy; it needs the fewest assumptions and pays with lower power.

When to use this sign test calculator

Use it when

  • Use it when only the direction of each paired difference is reliable: preference judgments, better/worse clinical assessments, or instruments whose readings are ordinal at best.
  • Use it as a robustness floor under a Wilcoxon or paired t result. Because its null hypothesis—the median difference is zero—needs almost no assumptions, a significant sign test is very hard to argue away.
  • Use it for quick analyses where an exact p-value matters more than power: the binomial reference is exact at every n, so there is no approximation to defend in a report.

Choose another method when

  • Avoid it when difference magnitudes are meaningful and roughly symmetric; the Wilcoxon signed-rank test uses that extra information and is systematically more powerful.
  • Avoid it for independent samples—without pairing there are no differences to sign; the Mann-Whitney U test covers that design.
  • Avoid it when most pairs are ties: dropped ties shrink the effective sample, and a conclusion from the few remaining pairs may not represent the process.

Interactive tool

The calculator loads as you approach this section so the guide remains fast on mobile connections.

Example preview

Direction-only paired readings

Inputs
Nine pairs: 7 positive and 2 negative differences.
Representative result
Exact two-sided p=2·P(X≤2)≈0.180 under Binomial(9, 1/2); no significant direction.

Illustrative only. Load the interactive tool to enter your own values and review assumptions.

Watch the explanation

4:53 min

The player loads only after you press play. You can also watch on YouTube.

Read the complete transcript

The Sign Test. Paired changes can carry trustworthy directions even when their magnitudes do not. The sign test keeps only plus or minus and asks whether those directions balance. Keep every pair intact and subtract sample two from sample one. Positive means sample one is larger; negative means sample two is larger. An exactly zero difference has no direction. This page drops every tie before inference, while preserving the original pair count for reporting. Count positives and negatives after ties leave. Their sum is informative n, the effective sample size used by the sign test. Under the directional null, positive and negative are equally likely among informative pairs. The positive count follows a binomial distribution with probability one half. This calculator evaluates those finite binomial tail probabilities directly at every n. It uses no normal approximation and no continuity correction. For a two-sided test, the page finds the smaller inclusive tail, doubles it, and caps the result at one. That is its exact convention.

Pairing is the design. Row i in sample one and row i in sample two must describe the same person or deliberately matched unit. Different pairs must remain independent. Households, clinics, or one person contributing several pairs require an analysis that represents that dependence. Only direction needs a meaningful order. The test does not require normality, equal spacing, or a symmetric distribution of paired differences. That robustness has a price: a tiny improvement and a huge improvement both contribute one positive sign. Their magnitudes disappear. The operational null is equal positive and negative probabilities among non-ties. Calling it a zero-median test needs extra continuity and point-mass care. Tail direction follows sample one minus sample two. Right-tailed favors positives, left-tailed favors negatives, and two-sided allows either imbalance. Exact calculation does not make small samples powerful. Discrete p-values move in coarse steps, so the attained error rate may sit below nominal alpha.

Now compare nine matched measurements under two conditions. Keep sample one minus sample two, a two-sided alternative, and alpha point zero five fixed. The nine paired differences contain seven positives, two negatives, and no ties. All nine pairs are informative for the sign count. Under the null, imagine nine independent fair signs. Every positive-count pattern lies on a binomial reference with five hundred twelve equally likely sign sequences. The smaller side is two. Counts from zero through two contain forty-six of those five hundred twelve sequences, giving tail probability zero point zero eight nine eight four four. Double that smaller inclusive tail. The exact two-sided p-value is zero point one seven nine six eight eight, above alpha point zero five. Fail to reject equal directional probabilities. This does not prove no change, perfect balance, or a zero median in the population. Seven ninths, or zero point seven seven eight, were positive among non-ties. Wilcoxon rejected on these same pairs because magnitudes added information; that is not a contradiction.

With five informative pairs, even five positives and no negatives give two-sided p zero point zero six two five. Alpha point zero five is unreachable. Many tied pairs shrink informative n and can make the nonzero subset selective. Always report original pairs, ties dropped, and directions retained. Unequal sample lengths fail validation. If every paired difference is zero, no direction remains, so the calculator stops instead of inventing evidence. Swap the sample order and seven positives become two. The positive share and one-sided tails reverse, while the two-sided p-value stays unchanged. Lock pair order, subtraction direction, tail, and alpha before results. Choosing a favorable tail or test afterward invalidates the evidence. Report pair design, sign orientation, positive, negative, and tied counts, informative n, exact convention, p, alpha, and guarded conclusion. This page provides no change magnitude or confidence interval. Reduce valid paired changes to honest directions, count exact binomial evidence, and run the Sign Test calculator free at Distri Scope dot com.

How to read the result

The statistic display shows the raw split, such as 7 positive vs 2 negative; everything else is a probability statement about that split under a fair coin.

Exactness cuts both ways: the p-value is exactly right, but with few pairs even extreme splits cannot reach small p-values—with five informative pairs the best possible two-sided p is 0.0625.

If the sign test and the Wilcoxon disagree, the difference magnitudes are doing the work: a few large differences on one side can move the Wilcoxon while leaving the sign count balanced.

How to use the sign test calculator

Enter the two paired samples in matching order with equal lengths; the calculator forms each difference (sample 1 minus sample 2) and keeps only its sign.

Tied pairs—zero differences—carry no directional information and are dropped, with the count reported; the test then runs on the informative pairs only.

A right-tailed alternative claims positive differences are more likely, meaning sample 1 tends to exceed sample 2 within pairs.

Formula, hypotheses, and assumptions

S=#{di>0}Binomial(n,12) under H0S = \#\{d_i > 0\} \sim \text{Binomial}(n, \tfrac{1}{2}) \text{ under } H_0

Conditions to review

  • Paired observations with a defined direction of difference
  • Independent pairs
  • Only an ordinal scale is required
  • Ties (zero differences) carry no information and are dropped

Calculator parameters

  • Significance Level (α): default 0.05.
  • Test Type: default Two-tailed.

What the method is doing

Under the null hypothesis each informative pair is a fair coin flip, so the positive count follows Binomial(n, 1/2) exactly; the calculator sums exact tail masses rather than using any normal approximation.

The two-sided p-value is twice the smaller tail (capped at one), the same convention scipy’s binomtest applies to a symmetric null.

The proportion of positive differences is reported as a simple effect summary; with all magnitudes discarded, that proportion is the entire observable effect.

Worked example: direction-only paired readings

Nine subjects are measured under two conditions, and only the direction of each within-pair change is considered trustworthy. Seven subjects score higher under condition A and two score higher under condition B, with no ties. Run a two-sided sign test at α = 0.05.

  1. 1Count the informative pairs: n = 7 + 2 = 9, with no zero differences to drop.
  2. 2Under the null, the positive count follows Binomial(9, 1/2); the observed smaller tail is P(X ≤ 2) = (1 + 9 + 36)/512 ≈ 0.0898.
  3. 3Double the smaller tail for the two-sided p-value: p = 2 × 0.0898 ≈ 0.180.
  4. 4Compare with α = 0.05: 0.180 > 0.05, so the null hypothesis of equally likely directions is not rejected; the positive share 7/9 ≈ 0.78 is the effect summary.

Interpretation

A 7-to-2 split looks lopsided, but under a fair coin such an imbalance among nine pairs happens about 18% of the time, so the evidence is insufficient at α = 0.05. The same data analyzed with the Wilcoxon signed-rank test—which also uses the magnitudes—did reject; the contrast is a clean illustration of the power the sign test trades for its minimal assumptions.

Common mistakes

  • Do not report the original pair count when ties were dropped; the test ran on fewer pairs, and the write-up should say so.
  • Do not read a non-significant sign test as evidence of no effect—its power is the lowest of the paired tests, and with small n even strong effects often fail to reach significance.
  • Do not switch between sign and Wilcoxon results after seeing both p-values; pick the test from the measurement quality before looking.

Limits and independent validation

Discarding magnitudes is the point, but it is still a loss: the test cannot distinguish many small improvements from a few large ones, and no confidence interval for the median difference is reported here.

With very few informative pairs the achievable p-values are coarse, and the smallest possible p may exceed conventional thresholds no matter how one-sided the data.

The test says nothing about the size of a change; pairing it with a descriptive summary of the differences is essential for an interpretable report.

Before using the result

  • Verify the pairing order across the two fields; a misaligned row silently flips signs.
  • Recount the positive, negative, and tied differences by hand—three integers—and check them against the display; this is the entire data reduction, so checking it validates everything.
  • Cross-check with the Wilcoxon signed-rank test: agreement strengthens the conclusion, and disagreement tells you the magnitudes, not the directions, carry the signal.

See every option in the Statistical Hypothesis Test Calculator or review the DistriScope methodology. Educational information; last reviewed 2026-08-09. Verify consequential calculations independently.