Interactive module guide

Kruskal-Wallis Test Calculator — Rank ANOVA

This Kruskal Wallis test calculator compares up to twelve independent groups through ranks, with Dunn-Bonferroni post-hoc pairwise comparisons.

Open calculator

Free · No sign-up · Calculations stay in your browser

The Kruskal-Wallis test extends the rank-based comparison to several independent groups: pool every observation, rank them, and ask whether the group mean ranks spread out more than chance allows.

Its tie-corrected H statistic is referred to a chi-square distribution with k − 1 degrees of freedom, making it the rank-based alternative to one-way ANOVA for skewed, ordinal, or outlier-prone data.

This calculator supports two to twelve groups, reports η²(H) as an effect size, and—after a significant omnibus result—runs Dunn-Bonferroni pairwise comparisons on the mean ranks so you can see which groups differ.

Use the Kruskal Wallis test calculator for several independent groups with skewed or ordinal values, then read the Dunn-Bonferroni comparisons only after a significant omnibus result.

When to use this Kruskal Wallis test calculator

Use it when

  • Use it to compare three or more independent groups whose values are ordinal ratings, skewed measurements, or contain outliers that would dominate group means in a one-way ANOVA.
  • Use it when group variances are unstable in ways that worry the ANOVA machinery; ranks tame heavy tails, and the sharp null—identical distributions—does not lean on normality anywhere.
  • Use it with the Dunn-Bonferroni follow-up when the research question is which groups differ, not only whether any do; the post-hoc comparisons use the pooled rank variance and adjust for multiplicity.

Choose another method when

  • Avoid it for repeated measures—when the same subjects appear in every condition, the Friedman test respects the blocking that this test would ignore.
  • Avoid reading a rejection as a difference in means or medians without the equal-shape assumption; unequal spreads alone can produce significance, which is a finding, but a different one.
  • Avoid it with only two groups when the Mann-Whitney U test answers directly; H with k = 2 is equivalent but the two-sample framing reports a more interpretable effect size.

Interactive tool

The calculator loads as you approach this section so the guide remains fast on mobile connections.

Example preview

Three small skewed groups

Inputs
Groups of 5, 4, and 5 values compared through pooled ranks.
Representative result
Tie-corrected H≈0.771 with df=2 and p≈0.680; no evidence of a location difference.

Illustrative only. Load the interactive tool to enter your own values and review assumptions.

Watch the explanation

5:12 min

The player loads only after you press play. You can also watch on YouTube.

Read the complete transcript

The Kruskal-Wallis Test. Several independent groups may differ even when raw means are fragile. The Kruskal Wallis test replaces values with pooled ranks and asks whether group rank positions separate. Independence is the design. Every person or unit belongs to one group only, and observations must remain independent both within and between groups. Pool every observation while keeping its group label. Rank the combined values from smallest to largest on one shared rail, never separately inside each group. Equal pooled values receive their average rank. These midranks preserve ties without inventing an arbitrary ordering, and each tied block also contributes to the correction. Return the pooled ranks to their original groups. Each group now has a rank sum and a mean rank that describe where its observations sit in the common ordering. The H formula increases when group rank sums spread apart. It scales each squared rank sum by its group size, then subtracts the common rank baseline. Ties reduce the available rank variation. Divide the uncorrected H by the standard tie factor; if every value is identical, that factor is zero and inference stops.

This page compares H with the upper tail of a chi square distribution using the number of groups minus one degrees of freedom. Larger H is stronger evidence. That chi square reference is an approximation. Any group below five observations triggers a warning, but the calculator never switches automatically to an exact or permutation method. The sharp null says every group shares one distribution. Rejection means at least one rank distribution differs, but it does not identify which group caused the omnibus result. A median or location interpretation needs reasonably similar group shapes. Unequal spread or shape can also move pooled ranks, so check distributions before naming a median shift. Meaningful ordering is enough for the rank calculation. Normality and equal raw spacing are not required, but rank robustness does not rescue dependent or wrongly grouped data. The result table shows mean ranks, not raw means or medians. A high mean rank means observations tend to appear later in the pooled ordering, nothing more specific. The page reports eta squared H as a rank effect, floored at zero. It is not raw outcome variance explained, and this page provides no confidence interval.

Now compare three independent groups with five, four, and five observations. Keep alpha point zero five fixed and select Dunn Bonferroni as a possible follow up. Pool all fourteen values. There are no ties, so ranks one through fourteen stay ordinary and the tie correction equals one. The group rank sums are thirty-six, thirty-six, and thirty-three. Dividing by group size gives mean ranks seven point two, nine point zero, and six point six. Insert those rank sums and sizes into the formula. H equals zero point seven seven one four, far smaller than a strongly separated rank pattern. With two degrees of freedom, the upper chi square p-value is zero point six seven nine nine six five. It exceeds alpha point zero five. Fail to reject the common distribution null. Dunn rows do not run after this non-significant omnibus, and the floored rank effect is zero. This is not proof that the three populations are identical. It means these fourteen observations do not separate the group rank positions beyond routine chance variation.

For contrast, three perfectly separated groups have mean ranks three, eight, and thirteen. H equals twelve point five and p equals zero point zero zero one nine three. A significant omnibus opens the selected Dunn follow up. Each pair compares two mean ranks with a tie adjusted pooled rank standard error and a two-sided normal z probability. Three groups create three comparisons. Bonferroni multiplies every raw p-value by three and caps at one; here only group one versus group three remains significant. An omnibus rejection does not require every adjusted pair to reject. Report the global result and every adjusted row honestly instead of forcing a story that all groups differ. Tied values receive midranks and alter H through the tie factor. If all values are identical, or a group has fewer than two observations, validation blocks the result. Report group sizes, independence, pooled midranks, tie correction, H, degrees of freedom, chi square approximation, p, alpha, rank effect, warning, and adjusted follow ups. Use Friedman for repeated measures. Turn several independent groups into one honest rank comparison, guard the omnibus before pairwise claims, and run the Kruskal Wallis calculator free at Distri Scope dot com.

How to read the result

The group table reports each group’s mean rank; those mean ranks—not means or medians—are what the test actually compares, and reading them first grounds every other number.

A significant H with all pairwise Dunn comparisons non-significant is possible: the omnibus pools evidence that no single pair concentrates. Report both levels honestly rather than forcing a pairwise story.

η²(H) near zero with a significant p indicates a real but small ordering effect made detectable by sample size; the effect size keeps the practical relevance question separate from the significance question.

How to use the Kruskal Wallis test calculator

Enter each group in its own field, using the add and remove controls for two to twelve groups; every group needs at least two observations, and sizes may differ.

Ties across the pooled data receive midranks, and the H statistic is divided by the standard tie-correction factor; heavily tied ordinal data are expected and handled.

The post-hoc selector controls whether Dunn-Bonferroni comparisons run after a significant omnibus result; they never run after a non-significant one.

Formula, hypotheses, and assumptions

H=12N(N+1)j=1kRj2nj3(N+1)H = \frac{12}{N(N+1)} \sum_{j=1}^{k} \frac{R_j^2}{n_j} - 3(N+1)

Conditions to review

  • Independent observations within and between groups
  • At least an ordinal measurement scale
  • Under the sharp null, all groups share one distribution
  • For a location reading, similar distribution shapes across groups

Calculator parameters

  • Significance Level (α): default 0.05.
  • Post-hoc Test: default Dunn-Bonferroni.

What the method is doing

The calculator computes H from the group rank sums, applies the tie correction 1 − Σ(t³−t)/(N³−N), and refers the result to χ²(k − 1)—the same convention as scipy’s kruskal, so omnibus results are directly reproducible.

Dunn comparisons use z = (R̄ᵢ − R̄ⱼ)/√(σ²(1/nᵢ + 1/nⱼ)) with the tie-adjusted pooled rank variance, and multiply each p-value by the number of comparisons (Bonferroni), capped at one.

η²(H) = (H − k + 1)/(N − k), floored at zero, estimates the share of rank variability attributable to groups; small chi-square approximation samples (any group under 5) trigger a warning.

Worked example: three small skewed groups

Three independent groups of measurements are compared: group 1 has 2.9, 3.0, 2.5, 2.6, 3.2; group 2 has 3.8, 2.7, 4.0, 2.4; group 3 has 2.8, 3.4, 3.7, 2.2, 2.0. With fourteen observations in total and no distributional information, run a Kruskal-Wallis test at α = 0.05 with Dunn-Bonferroni follow-up selected.

  1. 1Pool the 14 values and rank them; there are no ties, so the correction factor is one.
  2. 2Sum the ranks within each group and compute H = 12/(14×15) · ΣRⱼ²/nⱼ − 3×15 ≈ 0.771 with k − 1 = 2 degrees of freedom.
  3. 3The p-value is P(χ²₂ > 0.771) ≈ 0.680, far above α = 0.05, so the omnibus null is not rejected.
  4. 4Because the omnibus is not significant, no Dunn-Bonferroni comparisons are produced; η²(H) is floored at zero, matching the absence of evidence for any group effect.

Interpretation

The three groups’ mean ranks are as close as chance would routinely produce (p ≈ 0.680), so the data provide no evidence that any group tends to larger values. The example also shows correct post-hoc discipline: pairwise comparisons are not fished for after a non-significant omnibus result, which is exactly when they would be most misleading.

Common mistakes

  • Do not run or report pairwise comparisons when the omnibus test is not significant; the calculator enforces this order deliberately.
  • Do not describe a rejection as “the medians differ” without checking that the group distributions share a shape; dominance of ranks is the defensible reading.
  • Do not forget that Bonferroni is conservative: with many groups, real pairwise differences can vanish under the adjustment, and that is a power cost, not an error.

Limits and independent validation

The chi-square reference is an approximation; with any group under five observations the calculator warns rather than switching to exact permutation p-values, which it does not compute.

Dunn-Bonferroni is the only post-hoc offered; step-down procedures (Holm) or rank-based intervals are not yet available.

The test detects ordering tendencies, not specific parameter differences, and reports no confidence intervals for group contrasts.

Before using the result

  • Confirm independence within and across groups from the design—no subject may appear twice anywhere.
  • Recompute one group’s mean rank by hand from the pooled ranking and compare with the group table; a mismatch means a data-entry slip.
  • When distributions look roughly normal with similar spreads, cross-check with one-way ANOVA: agreement strengthens the finding, and disagreement points at outliers or skew driving one of the tests.

See every option in the Statistical Hypothesis Test Calculator or review the DistriScope methodology. Educational information; last reviewed 2026-08-09. Verify consequential calculations independently.