Two-Way ANOVA Calculator
Enter one line per level of factor A and one | separated cell per level of factor B to get the full two-way ANOVA table for both main effects and the A × B interaction, with partial η², cell means and standard deviations, a variance check and an interaction plot. Works with several replicates per cell or with one.
Every cell needs the same number of values (a balanced design). Unequal cell counts need general linear-model software, because the sums of squares then depend on the type used. One factor only? Use the one-way ANOVA calculator.
One line per level of factor A. Separate the cells (levels of factor B) with |, and the replicates within a cell with commas or spaces.
Related Calculators
ANOVA Calculator
Run one-way ANOVA calculations and inspect variance across groups.
Tukey HSD Calculator
Pairwise comparisons after ANOVA with Tukey HSD or Tukey-Kramer from raw data or an ANOVA table: differences, q statistics, adjusted p-values, and simultaneous confidence intervals.
Levene's Test Calculator
Test equal variances across groups with Levene (mean), Brown-Forsythe (median) or a 10% trimmed mean: W, p-value, F critical value, group SDs and the ANOVA on absolute deviations.
How to enter your data
Type one line for each level of factor A. Inside a line, separate the cells with a vertical bar (|): one cell for each level of factor B. The numbers inside a cell, separated by commas or spaces, are the replicates: the observations made under that exact combination of the two factors. Every cell needs the same number of values. A 2 × 3 design with four replicates per cell looks like this:
12, 14, 11, 13 | 18, 17, 20, 19 | 25, 27, 24, 26
10, 9, 12, 11 | 15, 16, 14, 17 | 18, 20, 19, 21
Lines that start with # are ignored. If your data are in Excel, this is the layout of the input range of Anova: Two-Factor With Replication: each line here is one block of consecutive rows there (the number of rows in a block is Excel's "Rows per sample"), and each | separated cell is one column.
What a two-way ANOVA tests
The ANOVA splits the variation of the data into four parts: what factor A explains, what factor B explains, what their combination adds beyond the two (the interaction), and the error, which is the variation of the replicates around their own cell mean. Each effect is compared with the error by an F ratio.
| Effect | Null hypothesis | Question it answers |
|---|---|---|
| Factor A | The mean response is the same at every level of A (averaged over the levels of B) | Does A matter on its own? |
| Factor B | The mean response is the same at every level of B (averaged over the levels of A) | Does B matter on its own? |
| A × B | The effect of A is the same at every level of B | Does the effect of one factor depend on the other? |
With a levels of A, b levels of B and r replicates in every cell (n = abr observations):
SS_A = b·r·Σ(A level mean − grand mean)², df = a − 1
SS_B = a·r·Σ(B level mean − grand mean)², df = b − 1
SS_AB = r·Σ(cell mean − A level mean − B level mean + grand mean)², df = (a − 1)(b − 1)
SS_error = Σ(value − its cell mean)², df = ab(r − 1)
MS = SS / df, F = MS_effect / MS_error, p = P(F(df_effect, df_error) ≥ F)
Partial η² is SS_effect / (SS_effect + SS_error). The results also show the critical F for your α: the value F must exceed for p to fall below α.
Worked example: tooth growth in guinea pigs
The ToothGrowth data set that ships with R (Crampton 1947; Bliss 1952) records the tooth length of 60 guinea pigs that received vitamin C as orange juice (OJ) or as ascorbic acid (VC) at three doses, ten animals per combination. Enter OJ as the first line (A1), VC as the second (A2) and the doses 0.5, 1 and 2 mg/day as the three cells (B1, B2, B3), or press Load ToothGrowth example.
| Mean length (SD) | 0.5 mg (B1) | 1 mg (B2) | 2 mg (B3) | Mean |
|---|---|---|---|---|
| OJ (A1) | 13.23 (4.46) | 22.7 (3.91) | 26.06 (2.66) | 20.6633 |
| VC (A2) | 7.98 (2.75) | 16.77 (2.52) | 26.14 (4.8) | 16.9633 |
| Mean | 10.605 | 19.735 | 26.1 | 18.8133 |
The calculator returns this table, which matches R's summary(aov(len ~ supp * factor(dose), data = ToothGrowth)):
| Source | SS | df | MS | F | p | Partial η² |
|---|---|---|---|---|---|---|
| Factor A (supplement) | 205.35 | 1 | 205.35 | 15.572 | 0.000231 | 0.2238 |
| Factor B (dose) | 2426.4343 | 2 | 1213.2172 | 92 | 4.0463e-18 | 0.7731 |
| A × B | 108.319 | 2 | 54.1595 | 4.107 | 0.02186 | 0.132 |
| Error | 712.106 | 54 | 13.1871 | — | — | — |
| Total | 3452.2093 | 59 | — | — | — | — |
- Dose has by far the largest effect (partial η² = 0.77): tooth length rises with every step in dose, and the three doses differ pairwise (Tukey HSD, p < 0.001).
- Orange juice gives longer teeth than ascorbic acid on average (20.66 against 16.96, p = 0.000231), but the interaction is significant too (p = 0.0219), so that average hides the real pattern.
- At 0.5 mg orange juice is 5.25 units ahead and at 1 mg 5.93 units ahead (t(54) = 3.23 and 3.65, p = 0.0021 and 0.0006, before any correction for three comparisons), but at 2 mg the two supplements are equal (difference −0.08, p = 0.96). In the interaction plot the two lines converge at the highest dose.
- Checks: the Brown-Forsythe test across the six cells gives W = 1.709, p = 0.148 (df 5, 54), and a Shapiro-Wilk test of the residuals gives W = 0.985, p = 0.669, so neither assumption is in doubt.
Reported in APA style: there were significant main effects of supplement, F(1, 54) = 15.57, p < .001, partial η² = .22, and of dose, F(2, 54) = 92.00, p < .001, partial η² = .77, and a significant supplement × dose interaction, F(2, 54) = 4.11, p = .022, partial η² = .13.
Check the small example by hand
Press Load example: two levels of A, two of B and three replicates per cell. The cell means are 1.2, 2.2, 3.2 and 4.2, the A level means 1.7 and 3.7, the B level means 2.2 and 3.2, and the grand mean 2.7.
SS_A = 2·3·[(1.7 − 2.7)² + (3.7 − 2.7)²] = 12 and SS_B = 2·3·[(2.2 − 2.7)² + (3.2 − 2.7)²] = 3. Every A level is the one above plus 2, so the interaction is exactly 0. Each cell has deviations of −0.1, 0 and 0.1 from its mean, so SS_error = 4·0.02 = 0.08 with 8 df, MS_error = 0.01 and F_A = 12 / 0.01 = 1200.
Reading the interaction plot
The plot draws one line per level of factor A through the cell means at each level of factor B. Parallel lines mean that the effect of A is the same at every level of B (no interaction). Lines that converge or fan out mean that the size of the effect changes, and lines that cross mean that its direction changes. The F test for A × B decides whether the departure from parallel is larger than the noise between replicates.
Look at the interaction first. When it is not significant, read the main effects from the level means in the margins of the cell means table. When it is significant, a main effect is an average over levels where the effect differs, and it can even vanish: enter 5, 6, 7 | 9, 10, 11 on one line and 9, 10, 11 | 5, 6, 7 on the next, and both main effects come out exactly 0 (all four margins equal 8) while the interaction has F = 48 (p = 0.000121). Then describe the cell means, or test the effect of one factor separately at each level of the other, and correct for the number of comparisons with Bonferroni or Tukey HSD.
Assumptions, and what to do when they fail
- Independent observations. Each value comes from a different subject or unit. Measuring the same subject in several cells breaks this (see the repeated measures question below).
- Roughly normal residuals. The residuals are the values minus their cell means. With about ten or more replicates per cell moderate skew does little harm. Check them with the Shapiro-Wilk test and read normality tests explained.
- Equal variances in every cell. The results list the cell standard deviations and run a Brown-Forsythe test across all cells (it needs at least 3 replicates per cell). With the same number of replicates in every cell the F tests tolerate moderately unequal variances; if the spread differs a lot, try a log or square-root transformation, or look closer with the Levene test calculator.
- A numeric outcome and a balanced layout. Ordinal outcomes and unequal cell counts call for other methods.
When the assumptions clearly fail, a transformation is the simplest fix. For a rank-based or robust two-way analysis use an R package such as ARTool or WRS2; the nonparametric calculators here handle one factor at a time, for example Kruskal-Wallis.
One observation per cell (no replication)
If every cell holds a single value, put one number in each cell (for example 45.2 | 47.9 | 52.1). There is no replicate to estimate the interaction separately, so the calculator fits the additive model: the interaction residual becomes the error term, with (a − 1)(b − 1) degrees of freedom, and only the two main effects are tested. This is Excel's Anova: Two-Factor Without Replication and the standard analysis of a randomized complete block design, where one factor is the treatment and the other a blocking factor. It is valid only when the two factors do not interact.
Example (made-up yields of four varieties, A, in three fields, B):
45.2 | 47.9 | 52.1
43.8 | 49.5 | 51.2
47.1 | 50.3 | 55.6
44.5 | 46.8 | 50.9
gives F(3, 6) = 8.0964 (p = 0.015686) for A, F(2, 6) = 58.7668 (p = 0.000115) for B, and an error sum of squares of 5.445 with 6 degrees of freedom.
Balanced and unbalanced designs
With the same number of replicates in every cell the effects are orthogonal: SS_A + SS_B + SS_AB + SS_error equals the total sum of squares exactly, and Type I, II and III sums of squares give the same table. With unequal cell counts they differ. Type I is sequential and depends on the order of the factors (R's anova() and summary(aov()) use it), Type II tests each main effect after the other main effect and ignores the interaction, and Type III (the default in SPSS) tests each effect after all the others. The right choice depends on the question you are asking, so this calculator refuses an unbalanced grid instead of choosing a type for you. Use general linear-model software for unbalanced data.
Two-way ANOVA in Excel, R, Python and SPSS
| Software | How to run it | Output rows |
|---|---|---|
| Excel | Data > Data Analysis > Anova: Two-Factor With Replication. Select the range including the labels and set Rows per sample to the number of replicates. | Sample (factor A), Columns (factor B), Interaction, Within (error), Total |
| Excel, one value per cell | Data > Data Analysis > Anova: Two-Factor Without Replication | Rows, Columns, Error, Total |
| R | summary(aov(len ~ supp * factor(dose), data = ToothGrowth)) | supp, factor(dose), supp:factor(dose), Residuals |
| Python | sm.stats.anova_lm(ols("len ~ C(supp) * C(dose)", data=df).fit(), typ=2) | C(supp), C(dose), C(supp):C(dose), Residual |
| SPSS | Analyze > General Linear Model > Univariate. Outcome in Dependent Variable, both factors in Fixed Factor(s). | Tests of Between-Subjects Effects (Type III by default) |
In Python, first run import statsmodels.api as sm and from statsmodels.formula.api import ols. Make numeric codes categorical (factor(dose) in R, C(dose) in Python); otherwise the dose is fitted as a slope and the degrees of freedom differ. The sums of squares, F ratios and p-values agree with this calculator for balanced data.
After the ANOVA: follow-up tests and effect sizes
A significant main effect with three or more levels only says that some levels differ. To find which, run Tukey HSD on the level means of that factor, using the error mean square and error degrees of freedom from the ANOVA table (the calculator has a "Means + MSE and df" mode; for the dose in the example that is the means 10.605, 19.735 and 26.1 with 20 observations each, MSE 13.1871 and 54 df). When the interaction is significant, compare the levels of one factor within each level of the other instead.
Report an effect size with every test: partial η² here, or use the effect size calculator and read effect size explained. To plan how many replicates you need before collecting data, see statistical power and sample size explained. Critical values are in the F table, and which statistical test to use helps when two factors are not what you have.
Frequently Asked Questions
What is a two-way ANOVA?
A two-way ANOVA tests how two categorical factors affect one numeric outcome. It gives an F test for the main effect of each factor and an F test for their interaction, which asks whether the effect of one factor depends on the level of the other. The tooth growth example above asks whether length depends on the supplement, the dose, and the combination of the two.
What is the difference between one-way and two-way ANOVA?
One-way ANOVA compares the means of the levels of a single factor. Two-way ANOVA adds a second factor, so it can test the interaction, and because the second factor explains part of the variation it also lowers the error term and can make the test of the first factor more sensitive. Use the one-way ANOVA calculator when you have a single grouping variable.
What does a significant interaction mean?
The effect of one factor is different at different levels of the other, which shows as non-parallel lines in the interaction plot. Main effects are averages over those different effects, so report the interaction first and describe the cell means, or test the effect of one factor separately at each level of the other.
Should I interpret the main effects when the interaction is significant?
With care. When the lines cross, a main effect can be zero or misleading even though the factor matters a great deal: the crossing example above has F = 0 for both main effects and F = 48 for the interaction. When the lines never cross and only differ in steepness, the direction of a main effect may hold in every cell, but its size varies, so say so.
How many replicates does each cell need?
At least 2, otherwise the interaction cannot be separated from error and the calculator fits the additive model instead. More replicates give more power and a steadier variance estimate: the Brown-Forsythe check in the results needs at least 3 per cell. To plan the design, a power analysis such as the F test for fixed-effects ANOVA (main effects and interactions) in G*Power sizes it from an effect size and α.
Why does the calculator reject my grid?
Every line must have the same number of cells separated by |, every cell needs at least one number, and every cell must hold the same number of values. Unequal cell counts (an unbalanced design) are not accepted, because the sums of squares then depend on the type (II or III) you choose; use general linear-model software for those.
Can I run a two-way ANOVA with one observation per cell?
Yes. With a single value per cell there is no way to estimate the interaction separately from error, so the calculator fits the additive model and uses the interaction residual as the error term, with (a − 1)(b − 1) degrees of freedom. This matches Excel's Anova: Two-Factor Without Replication and is valid only if the factors do not interact.
How do I match the output to Excel's Anova: Two-Factor With Replication?
Excel's Sample row is factor A (the blocks of rows), Columns is factor B, Interaction is A × B and Within is the error. Enter the same numbers with one line per block of rows and one | separated cell per column, and the sums of squares, degrees of freedom, F ratios and p-values agree. Excel's F crit column appears here too.
What is partial eta squared, and what counts as a large effect?
Partial η² is SS_effect / (SS_effect + SS_error): the share of the variation left after removing the other effects that this effect explains. The usual rules of thumb are 0.01 for a small, 0.06 for a medium and 0.14 for a large effect. In the tooth growth example dose has partial η² = 0.77, supplement 0.22 and the interaction 0.13.
How do I report a two-way ANOVA?
Give F with its two degrees of freedom, the p-value and an effect size for each effect, for example: there were significant main effects of supplement, F(1, 54) = 15.57, p < .001, partial η² = .22, and of dose, F(2, 54) = 92.00, p < .001, partial η² = .77, and a significant supplement × dose interaction, F(2, 54) = 4.11, p = .022, partial η² = .13. Add the cell means and standard deviations.
Which post-hoc test follows a two-way ANOVA?
When the interaction is not significant, compare the levels of a factor with Tukey HSD on that factor's means, using the error mean square and error degrees of freedom from the ANOVA table. When the interaction is significant, compare the levels of one factor within each level of the other and correct for the number of comparisons (Bonferroni or Tukey).
Can I use a two-way ANOVA for repeated measures?
Not in general. In a repeated-measures design the same subject appears in several cells, which changes the error structure. The one-observation-per-cell layout treats the subjects as blocks and assumes there is no subject by condition interaction. Use a repeated-measures ANOVA or a mixed model when subjects are measured more than once and that assumption is doubtful.
What if the assumptions are not met?
With equal replicates the F tests tolerate moderate non-normality and moderately unequal variances, especially with ten or more replicates per cell. For strong skew a log or square-root transformation often helps; otherwise use a rank-based or robust two-way method in R, for example the ARTool or WRS2 packages.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.