The best model is not an opinion. It is a league table.

Which model should make a given prediction is an empirical question with a daily answer. Most of the industry answers it with architectural ideology instead, committed before your data was ever seen.

2 min read

Ask a vendor why their model is right for your problem and you will receive an architecture story. Transformer this, foundation that, proprietary the other. Architecture stories are ideology. Which model best predicts a given target, on given data, at a given horizon, is an empirical fact, and an unstable one. It changes as the data changes, sometimes within a quarter.

The forecasting field has known this for decades, most famously through the M-competitions, where simple methods repeatedly embarrassed sophisticated ones on real series and no single method dominated across question types. A vendor whose product is one model, however capable, has answered an empirical question with a commitment made before your data was seen.

What evidence-driven selection requires

If model choice is an empirical question, the machinery for answering it follows directly, and none of it is exotic.

Selection has to be a standing competition, not a coronation. Many model families compete on each question, because the winner for steady weekly demand is routinely the wrong choice for intermittent spares or regulatory case timing. One question, one competition, one current champion.

The scoring has to happen on unseen data, in rolling backtests, with a placebo check behind it so a lucky fit cannot take a seat it did not earn.

The table has to be re-scored continuously as outcomes arrive. A champion that drifts loses its seat to the contender that has not, without a meeting, without a migration project, without anyone defending last year's architecture choice.

And every competition needs one entrant that does not care about elegance. The naive baseline, last year plus trend, sits in every table, and a champion must beat it or there is no champion. Research keeps finding that a large share of sophisticated forecasts fail exactly that test. When nothing beats naive, the honest output is a sentence. This target is not predictable yet with this data, and here is what would change that.

That sentence prevents the quiet catastrophe of enterprise forecasting, which is not bad models but confident automation of guesswork.

What this replaces

Evidence-driven selection retires two familiar failure modes. The first is the data science queue, where each question waits months for a hand-built model that then ossifies because nobody has time to revisit it. The second is the platform monoculture, where every question is answered by the vendor's one architecture, at whatever quality that architecture happens to achieve on it.

Against both, the competitive mechanism is boring, continuous and auditable. Every model's history is on the table, every substitution has a scoring reason, and the answer to "why this model?" is never a belief. It is a row.

Inside Prophesee's Foresight engine this runs as a standing league table of 62 models from 13 families, re-scored daily, naive baseline enforced as the floor. The table is the evidence for the argument, not the other way round. To see who wins on your data, start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.