Coding The Brains · Published 16 September 2026

Does your predictive model beat a simple baseline?

Evaluate a predictive model on held-out data, compare it with a simple baseline, and inspect errors in the groups where decisions matter. A better overall score does not by itself justify deployment. For numeric predictions, start by reporting error in the same units as the target, alongside sample size and the consequences of a wrong estimate.

This guide is for a project owner reviewing a regression-model proposal or delivery report. It follows our pre-model data-quality audit: clean records are necessary, but they do not establish that a prediction is useful.

1. Write down the decision before choosing the score

A price estimate used to help an analyst is different from one that automatically approves a purchase. State the target, currency, prediction time, permitted inputs and who acts on the output. Decide what happens when the input is outside the evaluated population. Define acceptance thresholds with the business before inspecting the final test results.

For a pricing task, an advertised price is not a completed sale value. An accurate prediction of the wrong target can still be a poor business tool. Our used-car project case note separates documented implementation work from performance figures that have not been published.

2. Keep the baseline independent of test labels

For a simple reference, predict the training-set median for every test row. Fit that number using training records only. Compare the candidate and baseline on exactly the same held-out rows. A segment-specific or seasonal rule may be a stronger reference when it matches the actual workflow; a weak baseline is not a convincing substitute for the process already in use.

Keep related entities together where needed, and use a time-respecting split when the task predicts future outcomes. Fit preprocessing on training data, not the combined training and test set. Do not keep changing the model after seeing the final test scores and still describe that set as untouched.

3. Read MAE and RMSE as errors, not percentages of accuracy

Mean absolute error (MAE) is the average absolute difference between predicted and actual values. Root mean squared error (RMSE) squares those differences before averaging and taking the square root, so large errors have more influence. Both use the target's units. An MAE of $2,000 is not “98% accuracy,” and neither metric says whether the system is safe to operate.

For implementation details, see scikit-learn's regression metrics, constant prediction baselines and data-leakage guidance.

4. Run an example where the average hides a worse segment

Download the standalone baseline check. Read it first, then run it with Node.js. It has no packages, network requests or file writes.

node regression-baseline-check.mjs

Every record and candidate prediction in this example is synthetic. The script does not train a model. Its candidate outputs are deliberately hand-authored to show how an overall improvement can hide a worse result for one input category. These numbers are not a client result, production benchmark or expected level of performance.

The five illustrative training prices are $10,000, $14,000, $15,000, $16,000 and $20,000. Their median is $15,000. The four separate test targets are $14,000, $16,000, $24,000 and $26,000; the candidate outputs are $17,000, $19,000, $23,000 and $25,000. The compact/utility category is an input known before observing the target.

Synthetic mean absolute error in USD; lower is better
Test sliceRowsMedian baseline MAECandidate MAE
All rows4$5,500$2,000
Compact2$1,000$3,000
Utility2$10,000$1,000

The candidate's overall error is lower, but compact-category error is three times the baseline. Four invented test rows cannot establish statistical confidence. The result illustrates a question for a real review: which segments got worse, how much evidence supports each result, and who bears the cost?

5. Ask for an evaluation handoff you can inspect

A sensible outcome may be to keep the baseline, restrict the model to a supported segment, improve data collection, or run a monitored pilot. Do not turn an attractive average into a promise of revenue or cost savings without measuring that outcome separately.