Frequency Distribution Calculator

Make a frequency distribution table from your data. Enter numbers to get one row for each value or to group them into classes of equal width, or enter words to count categories. You get the frequency, relative frequency, percent, cumulative frequency, cumulative relative frequency and cumulative percent of every row, the working, and a bar chart or histogram.

To choose the classes, use the class width calculator. For a larger chart with bins set by width or count, see the histogram maker, and for the centre of the same data the mean, median and mode calculator. When you already have the classes and their frequencies, the grouped data calculator finds the mean, median, mode and standard deviation from the table. Numbers are read as exact decimals, so a value on a class boundary is never moved by rounding.

Enter numbers separated by commas, spaces or new lines

What a frequency distribution shows

A frequency distribution lists every value, class or category of a data set together with how often it occurs. Counting the observations in each row turns a long list into a table that can be read at a glance, and it is the first step toward a histogram, a bar chart or a pie chart.

The frequency f of a row is the number of observations in it, and the frequencies add up to n, the total number of observations. The other columns put the counts in proportion: the relative frequency f / n, the percent, and the cumulative columns, which add up the rows from the top.

Formulas

Relative frequency = f / n

Percent = 100 × f / n

Cumulative frequency = f of this row + f of every row above it

Cumulative relative frequency = cumulative frequency / n

Class midpoint = (lower boundary + upper boundary) / 2

Class width = upper boundary − lower boundary

The relative frequencies of all rows add up to 1 and the percentages to 100, and the last cumulative frequency equals n. The table shows rounded numbers, so a rounded column can add up to slightly more or less; the calculator says so whenever it happens.

How to read the results

OutputWhat it tells you
Number of Observations (n)How many numbers or items were read from your data.
Number of Different Values, Classes or CategoriesHow many rows the table has.
Highest FrequencyThe largest frequency in the table.
Most Frequent Value, Modal Class or Most Frequent CategoryThe row with the highest frequency: the mode of the data, or the modal class for grouped data. A tie lists every row that shares the highest frequency.
Frequency (f)How many observations the row contains.
Relative Frequency (f / n)The share of all observations that are in the row, as a fraction of 1.
Percent (%)The same share as a percentage.
Cumulative FrequencyHow many observations are in this row and in every row above it.
Cumulative Relative FrequencyThe share of all observations up to and including this row.
Cumulative Percent (%)The same share as a percentage.
MidpointFor a class, the value halfway between its lower and upper boundary.

Worked example: number of siblings

Twenty students were asked how many siblings they have: 2, 1, 3, 0, 2, 1, 2, 4, 1, 2, 3, 2, 0, 1, 2, 3, 1, 2, 1, 0. Choose Numbers: one row for each value and Load example to fill in these answers.

ValueFrequencyRelative frequencyPercentCumulative frequencyCumulative relative frequency
030.151530.15
160.33090.45
270.3535160.8
330.1515190.95
410.055201

The most frequent answer is 2 siblings: 7 of the 20 students, a relative frequency of 7 / 20 = 0.35. The cumulative frequency of the value 1 is 9, so 9 students (45%) have at most one sibling. The frequencies add up to 20, the relative frequencies to 1 and the last cumulative frequency is 20.

Worked example: exam scores in classes

Twenty exam scores range from 58 to 100: 72, 85, 91, 68, 77, 95, 88, 64, 70, 100, 81, 79, 66, 74, 83, 90, 58, 62, 87, 76. Choose Numbers: grouped into classes, then Load example, which sets the first class to start at 50 and the class width to 10.

ClassMidpointFrequencyRelative frequencyCumulative frequencyCumulative relative frequency
[50, 60)5510.0510.05
[60, 70)6540.250.25
[70, 80)7560.3110.55
[80, 90)8550.25160.8
[90, 100)9530.15190.95
[100, 110)10510.05201

The scores 70, 90 and 100 sit on class boundaries. Each class includes its lower boundary, so they are counted in 70 to 80, 90 to 100 and 100 to 110. The modal class is 70 to 80 with 6 scores (30%). The cumulative frequency of that class is 11, so 11 of the 20 scores (55%) are below 80. To estimate the mean, median and standard deviation from a table like this one, without the raw scores, use the grouped data calculator.

How to choose the classes

  • Equal widths, no gaps, no overlaps. Every value must belong to exactly one class, and the bars of a histogram are only comparable when the classes are equally wide.
  • Start at or below the smallest value. The first class must reach the smallest value. A round number such as 50 or 0 makes the boundaries easy to read.
  • Pick a number of classes, then a width. Textbooks usually suggest between 5 and 20 classes. Sturges' rule gives k = ⌈log₂ n⌉ + 1 classes (the default of R's hist function) and the square-root choice gives ⌈√n⌉. For 20 values that is 6 and 5 classes. Divide the range by k and round up to a convenient width.
  • Look at the shape. Too few classes hide it, too many make it ragged. Try another width when the picture changes a lot.

The class width calculator turns a range and a number of classes into a width and lists the boundaries, and the range calculator finds the range of your data.

Which class does a value on a boundary belong to?

Every class here includes its lower boundary and excludes its upper one, written [lower, upper). A value that equals a boundary goes to the class that starts there, so 70 with classes of width 10 is counted in 70 to 80. The last class is never closed: if the largest value sits on a boundary, one more class is added for it. Other tools follow other rules, and the counts differ when values sit on the boundaries:

ToolRule for a value on a boundaryWhere 70 goes with limits 60, 70, 80
This calculator[lower, upper): the lower boundary is included70 to 80
Excel and Google Sheets FREQUENCYThe bin limits are upper limits, so a value equal to a limit is counted in that binThe bin with upper limit 70
pandas.cut(lower, upper] by default, [lower, upper) with right=False60 to 70 by default
R hist and cut(lower, upper] by default (right = TRUE), [lower, upper) with right = FALSE60 to 70 by default
numpy.histogram[lower, upper) except the last bin, which also includes its upper edge70 to 80

Whichever rule you use, state it with the table, and give the class boundaries rather than only the words "60-70" when the data are continuous.

Relative and cumulative frequency

Relative frequency lets you compare data sets of different sizes: 7 of 20 students and 35 of 100 students are the same share, 0.35. Cumulative frequency answers questions of the form "how many values are below this point?". For a class table the cumulative frequency of a class counts the values below its upper boundary, and the cumulative relative frequency is the share of values below that boundary.

Plotting the cumulative relative frequency against the upper class boundaries gives an ogive, from which you can read approximate percentiles. For exact percentiles of the raw values use the percentile calculator, and for the quartiles and the extremes the five number summary.

Frequency distribution of categories

For words or labels, enter one item per observation. Each different spelling is one category, so "Yes" and "yes" are two categories: use one spelling throughout. The rows can be listed in the order the categories first appear or from the highest frequency down, which puts the most common category first and makes the cumulative column show how much the top categories cover.

To test whether the observed counts fit an expected pattern, take the frequencies to the chi-square goodness of fit calculator.

Exact decimals instead of rounding surprises

Most decimals have no exact binary form, so a spreadsheet or program that computes the class of a value as FLOOR((x − start) / width) can move a value on a boundary into the wrong class. In double precision 0.3 / 0.1 is 2.9999999999999996, which would put 0.3 in the class before the one that starts at 0.3. This calculator reads every number as an exact decimal and compares and divides them exactly, so 0.3 is counted in [0.3, 0.4).

Values written differently are the same value: 1, 1.0 and 1e0 form one row, and numbers with up to 60 digits are accepted. A number that is not a number, a value below the first class and a table that would need too many classes are reported with a message instead of being guessed at.

Assumptions and pitfalls

  • Enter every observation. The calculator counts for you and does not read a column of counts: to enter a value that occurs 5 times, write it 5 times.
  • Grouping hides detail. A table of classes no longer shows the individual values, so a mean or median calculated from it is an approximation.
  • Rounded columns. Rounded percentages can add up to 99.99 or 100.01. The Total row shows the exact totals.
  • Compare with relative frequency. Two groups of different sizes cannot be compared by their counts.
  • Limits. A table is built from at most 10,000 observations and shows up to 500 different values or categories, or up to 200 classes. The chart is drawn for up to 60 rows.

Frequency tables in other software

ToolCommand
Excel / Google Sheets=FREQUENCY(A2:A21, C2:C6) with the upper class limits in C2:C6; =COUNTIFS(A2:A21, ">=60", A2:A21, "<70") counts one class [60, 70); =COUNTIF(A2:A21, "Apple") counts one category
Python (pandas)s.value_counts().sort_index(); s.value_counts(normalize=True); pd.cut(s, bins, right=False).value_counts().sort_index(); .cumsum() gives the cumulative column
Rtable(x); prop.table(table(x)); cumsum(table(x)); table(cut(x, breaks = seq(50, 110, 10), right = FALSE))
SPSSAnalyze > Descriptive Statistics > Frequencies shows Frequency, Percent, Valid Percent and Cumulative Percent

Frequently Asked Questions

What is a frequency distribution?

A frequency distribution is a table that lists each value, class or category of a data set with the number of times it occurs, its frequency. The frequencies add up to the total number of observations. Adding the relative frequency (the share of the total) and the cumulative frequency (the running total) makes the table easier to interpret and compare.

How do I make a frequency distribution table?

List each different value, or divide the range into classes of equal width that do not overlap. Count how many observations fall in each row: that is the frequency f. Divide each f by the total n for the relative frequency, multiply by 100 for the percent, and add up the rows from the top for the cumulative columns. This calculator does all of it from a pasted list of data.

What is relative frequency and how do I calculate it?

Relative frequency is the share of all observations that fall in a row: relative frequency = f / n. If 7 of 20 students have 2 siblings, the relative frequency of 2 siblings is 7 / 20 = 0.35, or 35%. The relative frequencies of all rows add up to 1.

What are cumulative frequency and cumulative relative frequency?

Cumulative frequency is the running total of the frequencies from the first row down to the current one: it counts the observations up to and including that row. The cumulative relative frequency divides it by n, so it is the share of observations up to that row, and the last one is always 1 (100%). In the exam scores example the cumulative frequency of the class 70 to 80 is 11, so 11 of the 20 scores (55%) are below 80.

How many classes should a frequency distribution have?

Textbooks usually suggest between 5 and 20 classes. Sturges' rule gives k = ceiling(log2 n) + 1 classes and the square-root choice gives ceiling(sqrt n): for 20 values that is 6 and 5. Set the class width to about the range divided by k, rounded up to a convenient number. Too few classes hide the shape of the data and too many make it ragged.

Which class does a value on a boundary belong to?

In this calculator every class includes its lower boundary and excludes its upper one, [lower, upper), so a score of exactly 70 with classes of width 10 is counted in 70 to 80, not in 60 to 70. Other tools differ: Excel's FREQUENCY function and pandas.cut count a value equal to a limit in the class below it by default. State the rule you used when you report a table.

Can I make a frequency table for words or categories?

Yes. Choose Categories and enter one item per observation, separated by commas or new lines. Each different spelling is its own category, so 'Yes' and 'yes' are two categories. The rows can be listed in the order the categories first appear or from the highest frequency down.

How do I make a frequency distribution table in Excel?

For single values, list the different values and use =COUNTIF(range, value) beside each one. For classes, type the upper limit of each class in a column and enter =FREQUENCY(data, bins) as an array formula; the limits are upper limits, so a value equal to a limit is counted in that class. You can also count [lower, upper) classes with COUNTIFS and the conditions >= and <, or use a pivot table with grouping.

Why do my percentages not add up to 100?

Each relative frequency and percent is rounded for display, so the displayed numbers can add up to slightly more or less than 1 or 100: three rows of 33.33 percent add up to 99.99. The exact values add up to 1 and 100, as the Total row shows, and the calculator adds a note whenever the rounded numbers do not.

How much data can the calculator handle?

A table is built from at most 10,000 observations. It can show up to 500 different values or categories, or up to 200 classes; for more distinct numbers, group them into classes. The bar chart or histogram is drawn for tables of up to 60 rows.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.