Statistics Concepts

Cronbach's Alpha Explained: Formula, Worked Example and Interpretation

Cronbach's alpha is a number, at most 1, that summarizes how consistently the items of a questionnaire or test vary together across respondents. It is computed from the number of items k and the ratio of the summed item variances to the variance of the total scores: α = k/(k − 1) × (1 − Σs²ᵢ / s²ₜ). For five respondents answering three items, the item variances 1.3, 1 and 1.2 add up to 3.5 and the total scores have variance 8.3, so α = 3/2 × (1 − 3.5/8.3) = 0.8675.

What Cronbach's alpha measures

A scale is a set of items, such as questions or statements, that are meant to measure one thing, for example job satisfaction, and are scored by adding the answers. Internal consistency asks whether the items hang together: do respondents who score high on one item tend to score high on the others? Cronbach (1951) proposed alpha as a summary of that. It is large when the items vary together, small when each item mostly varies on its own, and negative when items pull against each other.

Three properties are worth keeping in mind. Alpha depends on the number of items as well as on how related they are. It is a property of the scores of a particular sample, so it should be computed each time a scale is used rather than quoted from an earlier study (Tavakol and Dennick, 2011). And it is a coefficient, not a test: there is no p-value, only an interval that shows how precisely it is known.

The formula and its equivalent forms

α = k/(k − 1) × (1 − Σ s²ᵢ / s²ₜ)

α = k c̄ / (v̄ + (k − 1) c̄)

α_std = k r̄ / (1 + (k − 1) r̄)

Here k is the number of items, s²ᵢ the variance of item i and s²ₜ the variance of the total score. In the second form v̄ is the average item variance and c̄ the average covariance between two items; in the third, r̄ is the average correlation between two items. The first two are the same number, because the variance of a sum is the sum of the item variances plus all the covariances: s²ₜ = k v̄ + k(k − 1) c̄. Substituting this into the first form gives the second. Converting every item to z-scores makes v̄ = 1 and c̄ = r̄, which turns the second form into the third, the standardized alpha. It shows that alpha is fixed by two things only: the number of items and how strongly they correlate on average.

Worked example: three items, five respondents

Five people answered three statements from 1 (strongly disagree) to 5 (strongly agree).

RespondentItem 1Item 2Item 3Total
145413
233410
355414
42327
544513

Item 1: mean 3.6, Σ(x − x̄)² = 5.2, s² = 5.2 / 4 = 1.3

Item 2: mean 4, Σ(x − x̄)² = 4, s² = 4 / 4 = 1

Item 3: mean 3.8, Σ(x − x̄)² = 4.8, s² = 4.8 / 4 = 1.2

Σ s²ᵢ = 1.3 + 1 + 1.2 = 3.5

Totals 13, 10, 14, 7, 13: mean 11.4, Σ(x − x̄)² = 33.2, s²ₜ = 33.2 / 4 = 8.3

α = 3/2 × (1 − 3.5 / 8.3) = 1.5 × 0.57831 = 0.8675

The Cronbach's alpha calculator returns the same alpha and adds what the formula does not show. The standardized alpha is 0.8669, from a mean inter-item correlation of 0.6847 (the pairs run from 0.4564 to 0.8771). The 95% interval is 0.3304 to 0.9852: with 4 and 8 degrees of freedom F(0.975) = 5.0526 and F(0.025) = 0.1114, so the limits are 1 − 0.1325 × 5.0526 and 1 − 0.1325 × 0.1114. The item statistics are:

ItemCorrected item–total correlationAlpha if item deleted
10.93160.625
20.72340.8372
30.61630.9302

Two lessons follow. On the George and Mallery scale an alpha of 0.8675 is "good", yet with five respondents the interval reaches down to 0.33, so the data do not pin alpha down. And deleting item 3 would raise alpha to 0.9302, but with five respondents that difference is easily noise: check the wording of the item and collect more responses before changing the scale.

Why the number of items matters

Because alpha depends on k and r̄, a longer scale has a higher alpha even when its items are no more closely related. The table gives the standardized alpha for two values of the average inter-item correlation.

Items (k)α when r̄ = 0.2α when r̄ = 0.4
40.50.7273
80.66670.8421
160.80.9143
320.88890.9552

A 32-item scale whose items correlate only 0.2 on average reaches 0.89, higher than a 4-item scale whose items correlate 0.4 (0.73). Read alpha together with the mean inter-item correlation, which the calculator reports, and do not compare alphas of scales of very different lengths as if they measured the same thing.

How to interpret alpha

The rule of thumb of George and Mallery (2003) calls an alpha above 0.9 excellent, above 0.8 good, above 0.7 acceptable, above 0.6 questionable, above 0.5 poor and 0.5 or below unacceptable. A minimum of 0.7 is a widely used convention, usually credited to Nunnally (1978). These are conventions, not tests, and a higher value is not always better: Tavakol and Dennick (2011) note that reports of acceptable values range from 0.70 to 0.95, that a maximum of 0.90 has been recommended, and that a high alpha may mean that some items are redundant, asking the same question in different words.

What alpha cannot tell you

  • That the scale measures one thing. Alpha is not a measure of unidimensionality. Sijtsma (2009) explains why a high alpha is not evidence that the items measure the same thing, and Tavakol and Dennick (2011) add that a test with more than one concept can still reach a high alpha because the larger number of items inflates it.
  • The reliability of the scale. Alpha is a lower bound to reliability. It equals reliability only when the items are essentially tau-equivalent (Novick and Lewis, 1967), and in many cases it is a gross underestimate (Sijtsma, 2009). The documentation of R's psych package likewise says that alpha underestimates the reliability of a test and overestimates the first factor saturation, and points to omega as a fuller analysis.
  • How well it is known. The value comes from one sample; the confidence interval, which the calculator gives, is the measure of that uncertainty.
  • What to do with the items. A high or low alpha says nothing about which item to keep. The corrected item–total correlations and the alpha-if-item-deleted values help, but only together with what the items say.

What to do about a low alpha

  • Reverse-score negatively worded items first. An item worded in the opposite direction correlates negatively with the rest and can push alpha down, even below 0. In the calculator's own example, ten respondents and five statements give alpha 0.59 until statement 5 is reverse-scored and 0.89 afterwards.
  • Look at the corrected item–total correlations. Tavakol and Dennick (2011) suggest that items whose correlation with the total score approaches zero are the candidates for revision or deletion.
  • Check whether the items measure more than one concept. If they do, calculate alpha for each concept separately.
  • Add items if the scale is short. More related items that test the same concept raise alpha (see the table above).
  • Delete items last, one at a time, recalculating each time and reading the wording of the item, rather than removing everything whose alpha-if-deleted is higher.

What to report

  • The number of respondents and of items, and the range of the answer scale.
  • Alpha with its confidence interval, and the standardized alpha if the items have very different variances.
  • Which items were reverse-scored, and which items (if any) were deleted after seeing the data.
  • Alpha for each subscale rather than only for the whole questionnaire.

Alpha is built from variances and covariances, the quantities behind the covariance calculator and correlation calculator; see covariance vs correlation for how the two relate, and sample size explained for planning how many respondents to collect.

Try the Cronbach's Alpha Calculator

Alpha with a confidence interval, standardized alpha, item-total correlations, alpha if item deleted and reverse scoring.

Try the Correlation Calculator

The correlation between two items, with a confidence interval and a test.

Try the Variance Calculator

The sample variance of one item or of the total scores, with the working.

Frequently Asked Questions

Can Cronbach's alpha be greater than 1?

No. The variance of a sum of k items can never exceed k times the sum of their variances, which keeps α = k/(k − 1) × (1 − Σs²ᵢ / s²ₜ) at 1 or below, and it reaches 1 only when the items are perfect copies of one another after centering. There is no lower limit: alpha is negative when the items differ more within a respondent than respondents differ from each other.

Is a Cronbach's alpha of 0.7 good enough?

0.7 is the usual minimum in the literature, but it is a convention. It depends on how many items the scale has, and it is computed from one sample, so it comes with uncertainty. In the worked example an alpha of 0.87 from five respondents has a 95% interval of 0.33 to 0.99, which includes values well below 0.7. Report the interval next to the value.

Can I calculate Cronbach's alpha for a single question?

No. Alpha compares the variance of the items with the variance of their total, so it needs at least two items and at least two respondents. A single question has no internal consistency to measure; its reliability has to be studied by repeating the measurement.

Does Cronbach's alpha work for yes/no questions?

Yes. For items scored 0 and 1, alpha equals the Kuder–Richardson formula 20 (KR-20), the coefficient that alpha generalizes, provided the variance of the total score in KR-20 uses the same divisor (n) as the item variances p·q. Enter the 0/1 answers in the calculator like any other scores.

Should I report alpha for the whole questionnaire or for each subscale?

For each concept that the questionnaire measures. If a questionnaire has several concepts, one alpha for all items is inflated by the larger number of items and does not describe any of the concepts. Compute alpha for the items of each subscale separately.

How many respondents do I need?

There is no minimum at which alpha becomes reliable by itself; the confidence interval shows how well it is known. With 5 respondents and 3 items the 95% interval for an alpha of 0.87 is 0.33 to 0.99, and it narrows as the number of respondents grows. Look at the interval before deciding that a scale is good enough or that an item should be dropped.