Grouped Data Calculator
Calculate the mean, median, mode, variance, standard deviation and quartiles of grouped data from a table of classes and frequencies. Enter one class per line as the lower limit, the upper limit and the frequency. The calculator finds the class boundaries and class marks, shows the working table with f·x, f·x² and the cumulative frequency, and writes out every step.
Have the raw values instead of a table? The frequency distribution calculator groups them into classes, and the mean, median and mode calculator works on the values themselves. Every number is read as an exact decimal, so the fractions behind the results are exact.
One class per line: lower limit, upper limit, frequency. For example "20, 29, 12" or "20-29, 12".
Related Calculators
Frequency Distribution Calculator
Make a frequency table from numbers, classes or categories with relative and cumulative frequency, percent, the working and a bar chart or histogram.
Mean, Median, Mode Calculator
Find mean, median, mode, range, sum, and count with a full step-by-step solution.
Class Width Calculator
Divide a data range into equal histogram classes with boundaries, frequencies, and Sturges' rule.
What is grouped data?
Grouped data are observations summarised in a frequency table. Each row is a class, an interval such as 20-29, and its frequency is the number of observations that fall in it. Survey ages in brackets, exam scores by band, published income and census tables and most textbook exercises come in this form.
The individual values are not kept, so every statistic of grouped data is an estimate. The mean and the variance treat each value as if it were at the class mark, the midpoint of its class, and the median, quartiles and mode assume that the values are spread evenly across a class. The narrower the classes, the closer the estimates are to the numbers you would get from the raw data.
How to enter the classes
| You type | It means |
|---|---|
| 20, 29, 12 | The class from 20 to 29 holds 12 values |
| 20-29, 12 | The same class, written with a dash |
| 20 to 29: 12 | The same class, written with the word to |
| -10 to -1, 4 | A class of negative numbers with 4 values |
| 0.5-1.4 3 | Decimals work, and a space or a tab also separates the parts, so a table pasted from a spreadsheet works |
| 3, 3, 5 | A single value: the value 3 occurs 5 times |
- List the classes from the lowest to the highest. They may touch (10-20 followed by 20-30) but not overlap.
- Frequencies are whole numbers of 0 or more. A class with frequency 0 is allowed and keeps its place in the table.
- You can enter up to 200 classes, and blank lines are ignored.
- Every line is checked, and a line that cannot be read is reported with its line number, so you can correct it.
Open-ended classes such as "60 and over" have no upper limit and so no class mark. Close them with a limit you can defend, and say so in your answer, or leave them out and calculate for the closed classes only.
Grouped data formulas
Class mark: x = (lower limit + upper limit) / 2
n = Σf
Mean = Σ f·x / n
Median = L + ((n/2 − F) / f) × w
Q1 = L + ((n/4 − F) / f) × w, Q3 = L + ((3n/4 − F) / f) × w
Mode = L + ((fm − f1) / ((fm − f1) + (fm − f2))) × w
Population variance σ² = (Σ f·x² − (Σ f·x)² / n) / n
Sample variance s² = (Σ f·x² − (Σ f·x)² / n) / (n − 1)
Standard deviation = √variance
For the median and the quartiles, L is the lower boundary of the class where the cumulative frequency first reaches n/2, n/4 or 3n/4, F is the cumulative frequency of the classes before it, f is its frequency and w is its width. For the mode, fm is the highest frequency and f1 and f2 are the frequencies of the classes just before and just after it (0 if there is none). The interquartile range is Q3 − Q1.
Class limits and class boundaries
The limits are the numbers you type. The boundaries are the cut points between neighbouring classes, and the calculator puts each one halfway between the upper limit of a class and the lower limit of the next. That gives the usual boundaries for both common layouts without asking which one you have:
| Layout | Classes as written | Boundaries used |
|---|---|---|
| Classes that touch | 0-10, 10-20, 20-30 | 0, 10, 20, 30 (the limits themselves) |
| Whole numbers with a gap of 1 | 20-29, 30-39, 40-49 | 19.5, 29.5, 39.5, 49.5 |
| Decimals with a gap of 0.1 | 0.5-1.4, 1.5-2.4, 2.5-3.4 | 0.45, 1.45, 2.45, 3.45 |
The first and last boundaries move out by half of the gap next to them, and a single class uses its limits. The class mark is the same whether you take it from the limits or the boundaries. The median, quartiles and mode are measured from the lower boundary of their class, which is why the median of the ages below starts from 39.5 and not from 40.
Worked example: exam marks in classes that touch
Fifty students' marks are grouped as below. Choose Load example to enter this table.
| Class | Boundaries | Class mark x | f | f·x | f·x² | Cumulative f |
|---|---|---|---|---|---|---|
| 0 to 10 | 0 to 10 | 5 | 5 | 25 | 125 | 5 |
| 10 to 20 | 10 to 20 | 15 | 8 | 120 | 1800 | 13 |
| 20 to 30 | 20 to 30 | 25 | 15 | 375 | 9375 | 28 |
| 30 to 40 | 30 to 40 | 35 | 16 | 560 | 19600 | 44 |
| 40 to 50 | 40 to 50 | 45 | 6 | 270 | 12150 | 50 |
| Total | 50 | 1350 | 43050 |
- Mean = Σ f·x / n = 1350 / 50 = 27.
- Median: n/2 = 25 is first reached in the class 20 to 30, with 13 values below it and f = 15, so the median is 20 + ((25 − 13) / 15) × 10 = 28.
- Quartiles: n/4 = 12.5 falls in the class 10 to 20 (F = 5, f = 8), so Q1 = 10 + ((12.5 − 5) / 8) × 10 = 19.375. 3n/4 = 37.5 falls in 30 to 40 (F = 28, f = 16), so Q3 = 30 + ((37.5 − 28) / 16) × 10 = 35.9375. The interquartile range is 16.5625.
- Mode: the modal class is 30 to 40 with fm = 16, f1 = 15 and f2 = 6, so the mode is 30 + (1 / (1 + 10)) × 10 = 30.9091.
- Variance: the sum of squared deviations is 43050 − 1350² / 50 = 6600. The population variance is 6600 / 50 = 132, so σ = 11.4891, and the sample variance is 6600 / 49 = 134.6939, so s = 11.6058.
Worked example: ages in whole years
Forty customers are grouped by age in whole years. The classes 20-29 and 30-39 leave a gap of 1, so the boundaries fall at 19.5, 29.5, 39.5 and so on.
| Age | Boundaries | Class mark x | f | f·x | Cumulative f |
|---|---|---|---|---|---|
| 20 to 29 | 19.5 to 29.5 | 24.5 | 4 | 98 | 4 |
| 30 to 39 | 29.5 to 39.5 | 34.5 | 9 | 310.5 | 13 |
| 40 to 49 | 39.5 to 49.5 | 44.5 | 15 | 667.5 | 28 |
| 50 to 59 | 49.5 to 59.5 | 54.5 | 8 | 436 | 36 |
| 60 to 69 | 59.5 to 69.5 | 64.5 | 4 | 258 | 40 |
| Total | 40 | 1770 |
- Mean = 1770 / 40 = 44.25.
- Median: n/2 = 20 is first reached in the class 40 to 49 (L = 39.5, F = 13, f = 15), so the median is 39.5 + ((20 − 13) / 15) × 10 = 44.1667.
- Q1: n/4 = 10 falls in 30 to 39 (L = 29.5, F = 4, f = 9), so Q1 = 29.5 + ((10 − 4) / 9) × 10 = 36.1667. Q3: 3n/4 = 30 falls in 50 to 59 (L = 49.5, F = 28, f = 8), so Q3 = 49.5 + ((30 − 28) / 8) × 10 = 52.
- Mode: 39.5 + ((15 − 9) / ((15 − 9) + (15 − 8))) × 10 = 39.5 + (6 / 13) × 10 = 44.1154.
- Σ f·x² = 83220, so the sum of squared deviations is 83220 − 1770² / 40 = 4897.5. The population variance is 4897.5 / 40 = 122.4375 (σ = 11.0651) and the sample variance is 4897.5 / 39 = 125.5769 (s = 11.2061).
Why the results are estimates
Grouping throws the individual values away. Two data sets with the same table have the same grouped mean and variance even when their raw values differ, so the numbers here can differ a little from the ones you would get from the raw data. If you still have the raw values, the mean, median and mode calculator, the standard deviation calculator and the quartile calculator give the exact answers.
The grouped variance also tends to be a little too large. Sheppard's correction subtracts w²/12 (w is the class width) from it for a smooth distribution whose tails taper off at both ends. This calculator reports the plain textbook formula and does not apply the correction.
Sample or population standard deviation
Use the sample values, with the divisor n − 1, when the table is a sample from a larger population. Use the population values, with the divisor n, when the table covers every member of the group you are describing. The calculator shows both. The article on n versus n − 1 explains why they differ, and standard deviation covers the measure itself.
Assumptions and pitfalls
- The mode needs classes of equal width. With different widths the interpolation formula does not apply, so the calculator names the modal class but does not interpolate the mode, and it draws no histogram because the bars would have to show frequency density.
- Two classes can tie for the highest frequency. Then there is no single modal class and no interpolated mode.
- Enter counts, not relative frequencies. Percentages that add up to 100 give the right mean, median, mode, quartiles and population variance, but the sample variance would treat n = 100 as the sample size. Use the actual counts if you can.
- A table of single values is exact only for the mean and the variances. Entering 3, 3, 5 makes the class mark 3, so the mean and variances are those of the values. The median, quartiles and mode are still interpolated between boundaries half a unit either side, so they can differ from the median of the raw values. For a weighted average with other weights use the weighted mean calculator.
- Wide classes make poor estimates. A wide top class of incomes, for example, is not well represented by its midpoint. Narrower classes, or the raw data, give better answers.
- Choose the class width before you group. The class width calculator suggests one, and the frequency distribution calculator builds the table and its histogram from raw values.
Grouped data in other software
| Tool | Command |
|---|---|
| Excel / Google Sheets | Class marks in B2:B6 and frequencies in C2:C6. Mean: =SUMPRODUCT(B2:B6,C2:C6)/SUM(C2:C6). Sample variance: =(SUMPRODUCT(C2:C6,B2:B6^2)-SUMPRODUCT(C2:C6,B2:B6)^2/SUM(C2:C6))/(SUM(C2:C6)-1) |
| Python | import numpy as np; mean = np.average(x, weights=f); values = np.repeat(x, f); values.var(ddof=1) and values.std(ddof=1) for the sample variance and standard deviation |
| R | weighted.mean(x, f); values <- rep(x, f); var(values) and sd(values) for the sample variance and standard deviation |
In each case x holds the class marks and f the frequencies. Repeating each class mark f times reproduces the grouped mean and variance, not the raw ones.
Frequently Asked Questions
What is grouped data?
Grouped data are observations summarised in a frequency table: each row is a class, an interval such as 20-29, with the number of observations in it, its frequency. The individual values are not kept, so statistics of grouped data are estimates that treat every value in a class as if it were at the class mark, the midpoint of the class.
How do I find the mean of grouped data?
Find the class mark x of each class as (lower limit + upper limit) / 2, multiply it by the frequency f of the class, add up the products and divide by the total frequency: mean = sum(f x) / n. For the classes 0-10, 10-20, 20-30, 30-40 and 40-50 with frequencies 5, 8, 15, 16 and 6, sum(f x) = 1350 and n = 50, so the mean is 27.
How do I find the median of grouped data?
Find n/2, find the first class whose cumulative frequency reaches it (the median class) and interpolate: median = L + ((n/2 - F) / f) x w, where L is the lower boundary of the median class, F the cumulative frequency before it, f its frequency and w its width. In the marks example n/2 = 25, the median class is 20-30 with F = 13 and f = 15, so the median is 20 + (12/15) x 10 = 28.
How do I find the mode of grouped data?
Take the class with the highest frequency, the modal class, and interpolate with the frequencies of its neighbours: mode = L + ((fm - f1) / ((fm - f1) + (fm - f2))) x w. In the marks example the modal class is 30-40 with fm = 16, f1 = 15 and f2 = 6, so the mode is 30 + (1/11) x 10 = 30.9091. The formula assumes classes of equal width, and there is no single mode when two classes tie for the highest frequency.
How do I calculate the standard deviation of grouped data?
Work out sum(f x) and sum(f x^2), then the sum of squared deviations sum(f x^2) - (sum(f x))^2 / n, divide it by n for the population variance or by n - 1 for the sample variance, and take the square root. In the marks example sum(f x^2) = 43050 and sum(f x) = 1350, so the sum of squared deviations is 43050 - 1350^2/50 = 6600, the population standard deviation is sqrt(6600/50) = 11.4891 and the sample standard deviation is sqrt(6600/49) = 11.6058.
What is a class mark?
The class mark is the midpoint of a class, (lower limit + upper limit) / 2, and it stands for every value in the class in the mean and variance calculations. The class mark of 20-29 is 24.5 and that of 0-10 is 5.
What is the difference between class limits and class boundaries?
Class limits are the smallest and largest values a class can hold as written, such as 20 and 29. Class boundaries are the cut points between classes, halfway between the upper limit of one class and the lower limit of the next: 19.5, 29.5 and 39.5 for the classes 20-29, 30-39 and 40-49. The class mark is the same either way, and the median, quartiles and mode start from the lower boundary of their class.
How do I find the quartiles of grouped data?
Use the median method with another position: Q1 uses n/4 and Q3 uses 3n/4. Find the class where the cumulative frequency first reaches the position and interpolate, Q = L + ((position - F) / f) x w. In the marks example Q1 = 19.375 and Q3 = 35.9375, so the interquartile range is 16.5625.
Why are the results for grouped data only estimates?
Because the individual values are lost when they are grouped. The mean and variance assume every value sits at the class mark, and the median, quartiles and mode assume the values are spread evenly inside a class. Narrower classes give estimates closer to those of the raw data, and if you still have the raw data, use it.
Can I use open-ended classes such as 60 and over?
Not directly. An open-ended class has no upper limit and so no class mark. Close it with an upper limit you can justify and state that assumption in your answer, or leave it out and calculate for the closed classes only.
Should I use the sample or the population standard deviation for grouped data?
Use the sample standard deviation, with the divisor n - 1, when the table is a sample from a larger population, and the population standard deviation, with the divisor n, when the table covers every member of the population. The calculator shows both.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.