Mean Squared Error Calculator
Enter the observed values and the predictions made for them to get the mean squared error (MSE), the root mean squared error (RMSE), the mean absolute error (MAE), the mean absolute percentage error (MAPE), R² and the mean error. The calculator shows the working, a plot of observed against predicted values and a table of every error.
To fit a model first, use the linear regression calculator, the quadratic regression calculator or the exponential regression calculator; for R² alone see the R-squared calculator; to plot the errors and find the pairs that pull a fitted line, use the residual calculator. Related guide: how to interpret R².
Enter numbers separated by commas, spaces or new lines
One prediction for every observed value, in the same order
Related Calculators
R-Squared Calculator
Get R², adjusted R², r, and the SSR/SSE/SST breakdown from X–Y data or observed vs predicted values.
Linear Regression Calculator
Fit a least-squares line and get the equation, R², coefficient tests, ANOVA table, confidence and prediction intervals, residuals and a plot.
Standard Deviation Calculator
Calculate standard deviation, variance, and spread with clear statistical outputs.
What mean squared error measures
Mean squared error is the average of the squared differences between the observed values and the values a model predicted for them. Squaring makes every error positive and makes a large miss count far more than a small one: an error of 4 adds 16, four errors of 1 add only 4. The MSE is 0 only when every prediction is exact, and it grows as the predictions get worse.
Because the errors are squared, the MSE is in the squared units of Y (dollars squared, degrees squared). Its square root, the RMSE, is back in the units of Y and reads as a typical miss that leans towards the biggest errors. The MAE averages the absolute errors, so every unit of error counts the same. The MAPE expresses the errors as a percentage of the observed values, and the mean error keeps their signs to show whether the predictions are too high or too low on average.
Formulas
MSE = (1 / n) · Σ (yᵢ − ŷᵢ)²
RMSE = √MSE
MAE = (1 / n) · Σ |yᵢ − ŷᵢ|
MAPE = (100% / n) · Σ |yᵢ − ŷᵢ| / |yᵢ|
Mean error = (1 / n) · Σ (yᵢ − ŷᵢ)
R² = 1 − SSE / SST, SSE = Σ (yᵢ − ŷᵢ)², SST = Σ (yᵢ − ȳ)²
Here yᵢ are the observed values and ŷᵢ the predictions. This MSE divides the sum of squared errors by n, the convention of forecasting and machine learning (it is what scikit-learn's mean_squared_error returns). The mean squared error in a regression output, also called the residual mean square, divides by the degrees of freedom instead, n − 2 for a straight line, to estimate the error variance; its square root is the residual standard error that the linear regression calculator reports. That value is somewhat larger than the one computed here for the same residuals.
How to read the results
| Output | What it tells you |
|---|---|
| Mean Squared Error (MSE) | The average squared error, in squared units of Y. Lower is better, 0 is perfect. |
| Root Mean Squared Error (RMSE) | The square root of the MSE, in the units of Y. A typical miss, weighted towards the largest errors. |
| Mean Absolute Error (MAE) | The average size of the errors in the units of Y. The RMSE is never smaller than the MAE; a wide gap means a few large errors. |
| Mean Absolute Percentage Error (MAPE) | The average error as a percentage of the observed values. Undefined when an observed value is 0. |
| R² (coefficient of determination) | 1 − SSE / SST: how much better the predictions are than always predicting the mean of the observed values. It is negative when they are worse. |
| Sum of Squared Errors (SSE) | The total squared error, n times the MSE. |
| Total Sum of Squares (SST) | The variation of the observed values around their mean: the SSE of the baseline that always predicts the mean. |
| Mean Error (bias) | The average of observed minus predicted. Negative means the predictions are too high on average, positive means too low; 0 means no bias, not no error. |
Worked example: four predictions
Four values are predicted: observed = 3, −0.5, 2, 7 and predicted = 2.5, 0, 2, 8. Load example fills in these numbers, the same ones as in the scikit-learn documentation.
- Errors (observed − predicted): 3 − 2.5 = 0.5, −0.5 − 0 = −0.5, 2 − 2 = 0 and 7 − 8 = −1.
- MSE: the squares are 0.25, 0.25, 0 and 1, so SSE = 1.5 and MSE = 1.5 ÷ 4 = 0.375. RMSE = √0.375 = 0.6124.
- MAE: the absolute errors are 0.5, 0.5, 0 and 1, which add up to 2, so MAE = 2 ÷ 4 = 0.5.
- MAPE: the percentage errors are 0.5 ÷ 3 = 16.67%, 0.5 ÷ |−0.5| = 100%, 0 and 1 ÷ 7 = 14.29%, which average to 32.7381%. The observed value −0.5 is close to 0, so its percentage error dominates.
- Mean error: (0.5 − 0.5 + 0 − 1) ÷ 4 = −0.25: the predictions are 0.25 too high on average.
- R²: the observed mean is 2.875, SST = 29.1875, so R² = 1 − 1.5 ÷ 29.1875 = 0.9486.
scikit-learn gives the same figures: mean_squared_error returns 0.375, root_mean_squared_error 0.612… and r2_score 0.948….
MSE, RMSE, MAE or MAPE?
- MSE and RMSE suit problems where big misses are much worse than small ones, and they are the quantity that least squares minimises. They are sensitive to outliers: one wild prediction can dominate the total.
- MAE treats every unit of error equally and is easier to explain, so it is the better headline number when outliers are noise rather than signal. Compare it with the RMSE: a large gap points to a few big errors.
- MAPE is scale-free, so it compares series measured in different units, but it is undefined at an observed 0 and blows up when observed values are close to 0. It also punishes over-prediction more than under-prediction, because for non-negative predictions an under-prediction can never exceed 100%.
- R² puts the error in perspective by comparing it with the spread of the observed values. An RMSE of 10 is excellent when the observed values have a standard deviation of 200 and poor when it is 12; see the standard deviation calculator.
Assumptions and pitfalls
- The pairs must line up. The first prediction belongs to the first observed value, and so on. The calculator refuses lists of different lengths instead of trimming one of them.
- Score on data the model has not seen. Errors on the data a model was fitted to are optimistic, and a flexible model can drive them to zero without predicting anything.
- Do not compare values across datasets. The MSE depends on the scale of Y. To compare a model on different data, use R² or relate the RMSE to the standard deviation of the observed values.
- MAPE has no answer for an observed 0. This calculator reports it as undefined rather than replacing the zero with a tiny number, which would give an astronomically large percentage.
- A small MSE can hide a pattern. Two models can have the same MSE while one leaves errors that look random and the other a curve or a funnel. Plot the errors, for instance with the residual calculator, before trusting the number.
- R² can be negative. It is a comparison with the mean, not a squared correlation, so predictions worse than the mean of the observed values give a negative R².
Error metrics in other software
| Tool | Command |
|---|---|
| Excel / Google Sheets: MSE and RMSE | =SUMXMY2(observed, predicted) / COUNT(observed) and =SQRT(SUMXMY2(observed, predicted) / COUNT(observed)) |
| Excel / Google Sheets: MAE and MAPE | =SUMPRODUCT(ABS(observed - predicted)) / COUNT(observed) and =SUMPRODUCT(ABS((observed - predicted) / observed)) / COUNT(observed) |
| Python (scikit-learn) | mean_squared_error(y_true, y_pred), root_mean_squared_error(y_true, y_pred) since version 1.4, mean_absolute_error(y_true, y_pred), r2_score(y_true, y_pred) |
| Python MAPE | mean_absolute_percentage_error(y_true, y_pred) returns a fraction, so multiply by 100; for an observed 0 it divides by a tiny epsilon instead of reporting undefined |
| NumPy | numpy.mean((y - y_hat) ** 2) for MSE, numpy.sqrt(numpy.mean((y - y_hat) ** 2)) for RMSE |
| R | mean((y - y_hat)^2) for MSE, sqrt(mean((y - y_hat)^2)) for RMSE, mean(abs(y - y_hat)) for MAE |
All of these divide by n, so they agree with this calculator to rounding error. MAPE is the exception that needs care: check whether the tool returns a fraction or a percentage.
Frequently Asked Questions
What is mean squared error?
Mean squared error (MSE) is the average of the squared differences between observed values and the predictions made for them: MSE = (1 / n) · Σ (observed − predicted)². It is 0 for perfect predictions and grows with the size of the errors, with large errors counting much more than small ones because of the squaring.
How do I calculate MSE by hand?
Subtract each prediction from its observed value, square every difference, add the squares to get the sum of squared errors, and divide by the number of pairs n. For observed 3, -0.5, 2, 7 and predicted 2.5, 0, 2, 8 the squares are 0.25, 0.25, 0 and 1, the sum is 1.5 and the MSE is 1.5 / 4 = 0.375.
What is a good MSE?
There is no universal threshold, because the MSE depends on the scale of the data. Judge it against a baseline: take the RMSE and compare it with the standard deviation of the observed values, or look at R², which is 1 for perfect predictions, 0 for predicting the mean and negative for anything worse. Also compare it across models on the same data.
What is the difference between MSE, RMSE and MAE?
MSE averages the squared errors, so it is in squared units and is dominated by large errors. RMSE is its square root, in the units of the data. MAE averages the absolute errors, so each unit of error counts equally. RMSE is never smaller than MAE, and the gap between them grows when a few errors are much larger than the rest.
Why does my regression output show a different MSE?
Regression software often reports the residual mean square, the sum of squared residuals divided by the degrees of freedom (n − 2 for a straight line), because that estimates the error variance without bias. This calculator divides by n, as scikit-learn does, which describes the average squared error of the predictions. The two agree as n grows.
Why is MAPE undefined for my data?
MAPE divides each error by the observed value, so an observed value of 0 makes the percentage error impossible to compute. This calculator reports MAPE as undefined instead of dividing by a tiny number. Use MAE or RMSE for data that include zeros, or a scale-free measure such as R².
Can the mean squared error be negative?
No. Every squared error is zero or positive, so the MSE is at least 0. R² is different: it can be negative when the predictions are worse than always predicting the mean of the observed values. The mean error can also be negative, which just means the predictions are too high on average.
Should I use MSE or MAE to compare models?
Use MSE or RMSE when large errors are especially costly or when the model is trained by least squares. Use MAE when every unit of error costs the same or when outliers are noise. Reporting both is common, and a large gap between RMSE and MAE is worth investigating in the residual table.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.