Line of Best Fit Calculator

Find the line of best fit for a set of points. Enter the X and Y values and get the least-squares line y = mx + b with its slope, intercept, correlation and R², the full working table for small data sets, a scatter plot with the line drawn through the points, every residual, and the predicted Y for any X.

Need p-values, confidence intervals or the ANOVA table? Use the linear regression calculator. If the points bend instead of following a line, fit a curve with the quadratic regression calculator or the exponential regression calculator. To check the fit with a residual plot, use the residual calculator. Background: linear regression explained.

Enter numbers separated by commas, spaces or new lines

One Y for every X, in the same order

What the line of best fit is

The line of best fit, also called the least-squares regression line or trend line, is the straight line y = mx + b that passes as close as possible to a scatter of points. "As close as possible" has an exact meaning: of all straight lines, it is the one with the smallest sum of squared vertical distances between the points and the line. Those distances are the residuals, and the line always passes through the point (x̄, ȳ) made of the two means.

A line drawn by eye is a rough guide. The least-squares line is the one every statistics package and graphing calculator reports, and it is reproducible: the same data always give the same equation.

How to find the line of best fit by hand

1. Means: x̄ = Σx / n, ȳ = Σy / n

2. Deviations: x − x̄ and y − ȳ for every pair

3. Sxx = Σ(x − x̄)² and Sxy = Σ(x − x̄)(y − ȳ)

4. Slope: m = Sxy / Sxx

5. Intercept: b = ȳ − m·x̄

6. Line: y = m·x + b

For 12 pairs or fewer the calculator prints this whole table, so you can check each column against your own work. The slope m is the change in Y for each one-unit increase in X, and b is the value of the line at X = 0.

Worked example: hours studied and exam score

Five students study for X = 1, 2, 3, 4, 5 hours and score Y = 60, 65, 70, 80, 85. Load example fills in these numbers and predicts the score for 3.5 hours.

Pairxyx − x̄y − ȳ(x − x̄)²(x − x̄)(y − ȳ)
1160−2−12424
2265−1−717
33700−200
44801818
5585213426
Sum15360001065
  1. Means: x̄ = 15 ÷ 5 = 3 and ȳ = 360 ÷ 5 = 72.
  2. Slope: m = Sxy ÷ Sxx = 65 ÷ 10 = 6.5.
  3. Intercept: b = 72 − 6.5 × 3 = 52.5, so the line is y = 6.5x + 52.5.
  4. Fit: the fitted scores are 59, 65.5, 72, 78.5 and 85, so the residuals are 1, −0.5, −2, 1.5 and 0 (they always sum to 0). SSE = 7.5 and SST = 430, giving R² = 1 − 7.5 ÷ 430 = 0.9826 and r = 0.9912.
  5. Prediction: for 3.5 hours, ŷ = 6.5 × 3.5 + 52.5 = 75.25.

Interpretation: each extra hour of study goes with about 6.5 more points, and 98% of the variation in scores is explained by hours studied in this small sample. A student who studied 0 hours is predicted to score 52.5, but 0 hours lies outside the data, so that intercept is an extrapolation.

The line through two points

With exactly two points the least-squares line is the line through both of them, and both residuals are 0. Enter X = 1, 4 and Y = 3, 9: the slope is (9 − 3) ÷ (4 − 1) = 2 and the intercept is 3 − 2 × 1 = 1, so y = 2x + 1 with r = 1. With two points there is no scatter to measure, so use at least three when you want a sense of how well a line describes the data.

Using the line: interpolation, extrapolation and fit

  • Interpolation (predicting inside the range of your X values) is usually reliable. Extrapolation (predicting outside it) assumes the pattern continues, which the data cannot confirm; the calculator flags such predictions.
  • r is between −1 and 1 and gives the direction and strength of the straight-line pattern; R² = r² is the share of the variation in Y that the line explains.
  • Residuals that show a curve, a funnel or one very large value mean a straight line is not the right summary. Look at the plot and the residual table, or draw a residual plot with the residual calculator, before quoting the equation.
  • Which variable is X? The line minimizes vertical distances, so swapping X and Y gives a different line. Put the variable you want to predict in Y.

The linear regression calculator adds the standard errors, p-values and prediction intervals for the same line, and the correlation calculator focuses on the strength of the relationship.

Line of best fit in other tools

ToolHow to get the line
Excel / Google Sheets=SLOPE(y, x) and =INTERCEPT(y, x); or add a linear trendline to a scatter chart and tick Display Equation on chart
TI-84STAT > CALC > 4:LinReg(ax+b) gives a (slope) and b (intercept); turn DiagnosticOn to see r and r²
DesmosEnter the points in a table, then type y1 ~ mx1 + b for the regression
Rcoef(lm(y ~ x))
Pythonnumpy.polyfit(x, y, 1) returns [slope, intercept]

Frequently Asked Questions

What is the line of best fit?

It is the straight line y = mx + b that lies closest to the points of a scatter plot in the least-squares sense: it has the smallest possible sum of squared vertical distances from the points. It always passes through the point (x̄, ȳ) of the two means, and its residuals sum to zero.

How do I find the equation of the line of best fit?

Compute the means x̄ and ȳ, then Sxx = Σ(x − x̄)² and Sxy = Σ(x − x̄)(y − ȳ). The slope is m = Sxy / Sxx and the intercept is b = ȳ − m·x̄. The calculator does this for you and, for up to 12 pairs, prints the table so that you can follow each column.

Is the line of best fit the same as the regression line or the trend line?

For a straight line, yes: the least-squares regression line, the line of best fit and a linear trendline in a spreadsheet chart are the same line. A spreadsheet can also fit curves such as exponential and polynomial trendlines, which are not straight lines.

Does the line of best fit have to pass through any of the data points?

No. It usually misses most of the points; it only has to balance them, so the positive and negative residuals cancel and the line passes through the mean point (x̄, ȳ). It passes through every point only when the points already lie on one straight line, in which case r is 1 or −1.

How do I draw the line of best fit on a scatter plot?

Plot the points, calculate the equation, then compute y for two X values near the ends of your data (for example the smallest and largest X) and draw a straight line through those two points. As a check, the line must also pass through (x̄, ȳ).

What does a negative slope mean?

A negative slope means Y tends to decrease as X increases. If the slope is −2.5, each one-unit increase in X goes with a 2.5-unit decrease in Y on average, and the correlation r is negative as well.

What is a good R² for a line of best fit?

It depends on the field: physical measurements often reach 0.99, while human behavior or economic data may be interesting at 0.2 to 0.4. A high R² does not prove the line is the right model, and a low one does not mean the relationship is unimportant. Check the residuals and the scatter plot, and use the linear regression calculator to test whether the slope differs from zero.

What if my data follow a curve?

A straight line will fit a curve badly, and the residuals will show a systematic pattern such as positive at both ends and negative in the middle. In that case fit a curve instead of a line, or transform one variable (for example with a logarithm) and fit a line to the transformed data.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.