The Retrodiction Fallacy: Backfitting, Power Laws, and the Manufacture of Predictive Authority

2026-06-17 · 10,906 words · Singular Grit Substack · View on Substack

Why Models That Explain Everything After the Fact Predict Nothing Before It

This essay examines the distinction between prediction and retrodiction, and argues that models derived chiefly by fitting curves to history possess little scientific value, however elegant they appear. Drawing on the philosophy of science, econometrics, statistics, and the long catalogue of financial manias, it shows that many of the “laws” celebrated in markets are artefacts of selective observation and parameter optimisation rather than discoveries about the world. It traces how power-law claims are manufactured, why they persuade, and how they fail every serious requirement of scientific explanation; and it closes with a practical test for telling a genuine predictive model from a flattering story told after the fact.

Keywords: Power laws, backfitting, overfitting, forecasting, philosophy of science, Bitcoin


I. The Seduction of the Perfect Fit

Man is the animal that cannot leave a pattern alone. Give him three points and he will draw a line; give him a line and he will call it a law; give him a law and he will sell you the future. The faculty that lets us find the lion in the grass and the harvest in the season is the same faculty that finds faces in clouds, conspiracies in coincidence, and prophecy in the random walk of a price. It is a magnificent instrument, and like most magnificent instruments it is most dangerous precisely where it is most confident.

Look at what we have done with the night sky. The stars are scattered by accident across distances so vast that no two in a constellation have anything to do with one another, and yet every culture that ever looked up drew the same kind of pictures: hunters, bears, scales, kings. We did not discover the Plough; we imposed it. The constellation is not in the heavens. It is in the eye. And the eye, having drawn the figure, insists that the figure was always there.

Move from the sky to the screen and the habit is unchanged. The chartist sees a head and shoulders where a statistician sees noise. The numerologist finds significance in a birth date. The forecaster rules a confident line through a decade of data and pronounces it destiny. In each case the procedure is identical: a pattern is detected, the detection is mistaken for a discovery, and the discovery is mistaken for a mechanism.

Here is the trouble. The mind evolved to detect patterns, and it was not penalised for detecting too many. The ancestor who mistook the wind for a predator wasted a sprint; the ancestor who mistook a predator for the wind was removed from the gene pool. We are the descendants of the nervous, and nervousness has a signature: it cries wolf. We see pattern where there is only chance, because in the long arithmetic of survival a false alarm is cheaper than a missed threat. The cost of that bargain is that we are constitutionally credulous about regularities.

Now place before such a creature a graph. Let a smooth line pass cleanly through the scattered points of the past, each observation sitting obediently near the curve. The effect is hypnotic. The line seems not merely to describe the points but to command them, as if the data had been trying all along to obey a rule it had not yet been told. The observer draws the only conclusion his wiring permits: if it explained the past so beautifully, it must govern the future.

That conclusion is false. It is the central falsehood this essay exists to dismantle. A curve that fits the past has demonstrated exactly one thing — that it fits the past — and the past is the one stretch of time about which prediction is impossible, because it has already happened. The fit is a record of obedience already rendered, not a promise of obedience to come.

The whole question of whether a model is science or theatre turns on a single distinction, and almost everyone gets it backwards. The distinction is not how well the model fits the data. It is whether the model predicted the data or merely accommodated it. A theory that announced the number before the number was known has staked something and won. A theory that was shaped, after the fact, until it matched the number already on the table has staked nothing and proved nothing. The fitted line is the most flattering of mirrors. It shows the past exactly as the past was, and asks to be congratulated for the resemblance.

II. Prediction and Retrodiction

A prediction is a statement about an observation that has not yet been made. A retrodiction is a statement about an observation that has already been made. The words are nearly twins and the things they name are nearly opposites, and the entire health of a science can be read in which of the two it lives by.

The difference is not pedantry; it is the difference between a wager and a memoir. When you predict, the world has not yet voted, and it may vote against you. The prediction can be wrong, and because it can be wrong, its coming true means something. When you retrodict, the world has already voted, the result is on the record, and you are free to compose a theory that the record cannot embarrass. A retrodiction cannot lose the bet, because the bet was placed after the race was run. This is why the past is so cheap and the future so dear. Anyone can be a prophet of yesterday.

Consider the case that did more than any other to fix this standard in the modern mind. In 1915 Einstein completed the general theory of relativity, and the theory said something specific and frightening: that light passing close to the Sun would be bent by a definite amount, nearly twice the deflection that Newton’s physics allowed. No one had measured it. There was no result to fit. The number existed only in the theory, daring the universe to contradict it. In 1919 an eclipse permitted the measurement, the starlight was seen to bend, and the deflection matched. That is prediction in its purest and most dangerous form: a claim made before the fact, falsifiable in principle, confirmed in practice. Einstein had wagered his theory against the sky, and the sky had paid out.

But the same episode contains a subtlety that almost every retelling flattens, and the subtlety is the real lesson. General relativity also accounted for an old anomaly in the orbit of Mercury, whose perihelion advanced by a tiny excess that Newtonian mechanics could not explain and that astronomers had known about, and puzzled over, since Le Verrier described it in the eighteen-fifties. The excess was not predicted in the temporal sense. It was already on the books. Strictly, relativity retrodicted it. And yet physicists rightly counted Mercury as powerful confirmation, every bit as serious as the eclipse. Why? Because the theory accounted for Mercury with no adjustable parameters. Einstein did not reach into the equations and tune a dial until the orbit came out right. The figure fell out of a structure that had been built for entirely independent reasons, with nothing left free to fudge. The fit was not purchased. It was forced.

There is the whole matter in a single contrast. Temporal order — before or after — is not the deepest thing. The deepest thing is whether the agreement was bought with free parameters. A retrodiction from a rigid theory that could not have been bent to fit is worth more than a prediction from a theory so limp it could have fitted anything. What damns the curve-fitter is not that he speaks after the event. It is that he speaks with a fistful of dials, and turns them until the past obeys, and then presents the obedience he manufactured as a discovery he made. Mercury was a retrodiction and it was science, because nothing was adjusted. The typical market “law” is a retrodiction and it is not science, because everything was.

Gregor Mendel’s peas make the same point from the other side. He proposed a mechanism — discrete inherited factors, combining by simple rules — and from the mechanism flowed ratios that one could go and count: three to one, nine to three to three to one. The numbers were consequences of the theory, not inputs to it. He did not observe the ratios and then invent factors to explain them; he posited the factors and the ratios came due. That is what it looks like when explanation runs in the honest direction, from cause to consequence, and then submits the consequence to the world for checking.

Set beside these the characteristic productions of the forecasting trades. The economic cycle theory that appears, with impeccable timing, after the crash, and shows that the crash was written in the stars all along. The chartists’ patterns, named and catalogued, every one of them discovered in hindsight and none of them able to tell you on Tuesday what Wednesday holds. The power-law overlay drawn across the price history of a speculative asset, fitted to the very observations it claims to explain, advertised as a glimpse of an iron necessity governing the years to come. Each of these is a retrodiction in costume. Each was shaped to fit a past already known, with parameters free to be chosen for exactly that purpose. Each explains everything and forecloses nothing.

Three voices stand behind this distinction, and it is worth naming them, because they disagreed about much and converged on this. Karl Popper made the asymmetry the centre of his philosophy: a theory earns its standing not by the confirmations it collects but by the refutations it survives, and a theory that forbids nothing — that is compatible with every possible observation — tells us nothing about the world. Milton Friedman, from a different tradition and with a different temperament, insisted that the only test of a theory worth the name is the accuracy of its predictions. There is an irony here we shall return to: Friedman’s famous indifference to whether a theory’s assumptions are realistic, pressed to its limit, licenses the very overfitting he would have despised, for if all that matters is the fit, the cynic will simply manufacture the fit. A man who cares only that the answer comes out right, and not at all about why, has built the back door through which the charlatan walks. And Richard Feynman, in his lecture on what he called cargo-cult science, drew the line in plain language: the first principle is that you must not fool yourself, and you are the easiest person to fool. The fitted curve is the most efficient instrument of self-deception ever drawn, because it wears the dispassionate face of mathematics while flattering the one hope the modeller cannot give up.

The proposition that organises everything that follows is therefore very simple, and it is merciless. Explaining what has already happened is easy; any sufficiently flexible scheme can do it, and the more flexible the scheme the more easily it is done. Predicting what has not yet happened is hard, and stays hard, and cannot be faked by anyone forced to commit before the result is in. Science is the enterprise that submits to the second standard. Everything that ducks it and lives on the first is, whatever its mathematics, a kind of storytelling — and we shall see that the mathematics makes the story more dangerous, not less.

III. The Mathematics of Backfitting

Why is the past so easy to fit? The answer is not psychological but mathematical, and once it is grasped the spell of the well-fitted curve is broken for good. The answer is parameters.

A parameter is a dial: a number in a model that you are free to choose. Give a model one dial and it can shift up and down. Give it two and it can shift and tilt. Give it enough and there is almost no shape it cannot be bent into. A straight line has two parameters and can pass exactly through any two points you like. A parabola has three and can pass exactly through any three. A polynomial of degree n has n plus one, and can be driven exactly through n plus one arbitrary points — through the closing prices, if you wish, of n plus one consecutive days, no matter how those prices jump and stagger. The fit will be perfect. It will also be worthless, because a curve threaded exactly through every past point has learned the noise along with the signal, and will lurch off into nonsense the moment it is asked about a point it has not already seen.

The most famous sentence in the literature of overfitting makes the case with more wit than a textbook can, and its provenance is itself a small parable. Freeman Dyson, late in his life, recounted a meeting with Enrico Fermi in 1953 at which Fermi demolished in a few minutes a research programme Dyson and his students had pursued for years. Fermi asked how many free parameters Dyson had used. Dyson said four. Fermi replied: “I remember my friend Johnny von Neumann used to say, with four parameters I can fit an elephant, and with five I can make him wiggle his trunk.” With that, Dyson wrote, the conversation was over. Notice the chain of custody. The remark is von Neumann’s, reported by Fermi, recalled by Dyson, and printed in a journal half a century later. It is an aphorism about the manufacture of false authority that has itself acquired authority by being passed from one great name to the next — and, lest anyone think it mere wit, mathematicians have since shown that four complex parameters really do suffice to draw a passable elephant. The joke is a theorem. With enough freedom you can fit the beast; with a little more you can animate it; and you will have learned precisely nothing about elephants.

The discipline that studies this formally calls it the bias–variance trade-off, and the dry name conceals an iron law. A model that is too rigid has high bias: it cannot bend to the real structure in the data, and it is wrong in a stable, honest way. A model that is too flexible has high variance: it bends to everything, including the accidents, and so it is wrong in an unstable, flattering way — superb on the data it was fitted to, hopeless on the next sample, because the next sample carries different accidents. Every additional parameter buys a closer fit to the past at the cost of a looser grip on the future. Beyond a point — and the point comes quickly — you are no longer modelling the world. You are memorising your sample and mistaking recall for understanding. A model with as many parameters as data points does not explain the data; it merely repeats it back, the way a parrot that has learned a sentence has not learned the language.

And this is before we reach the subtler corruptions, the ones that do their work without a single dishonest intention. The first is data snooping, sometimes called data dredging: the practice of trying relationship after relationship against the same body of data until one of them, by the ordinary luck of large numbers, comes out looking significant. If you test twenty independent hypotheses at the usual threshold, you should expect one to clear the bar by chance alone, having nothing whatever to do with the world. Test a thousand, and you will harvest a field of impressive coincidences. The literature on calendar effects in stock returns — the January effect, the Monday effect, the turn-of-the-month effect — is in large part a monument to this: torture a long enough price series with enough candidate patterns and it will confess to anything, and the confession will not replicate.

The second is multiple testing’s near relation, the practice the field has lately learned to call p-hacking, and which Andrew Gelman has aptly named the garden of forking paths. Here the analyst does not even run a thousand explicit tests. He makes, instead, a long sequence of small choices — which observations to include, where to start the sample, how to bin the data, which outliers to discard as anomalous, which transformation to apply — and at each fork he takes, often in perfect good faith, the path that makes the picture a little cleaner. No single step is fraud. The cumulative effect is a result selected for its beauty rather than its truth, sculpted out of the degrees of freedom hidden in the analyst’s own discretion. The fitted curve that emerges from such a process is not a finding. It is the residue of a thousand quiet decisions, each made because it helped.

Put the pieces together and the moral is stark, and it runs exactly opposite to common intuition. The closer a model fits the history, the more suspicious we should be, not the less — for the closeness may be the symptom of disease rather than the sign of health. The more freedom a model is given, the more easily it conforms to the past; and the more easily it conforms to the past, the less its conformity can possibly tell us, because a model that could have fitted anything has told us nothing by fitting this. Flexibility and evidential value are not merely different. They are at war. Every dial you add to capture the past is a dial subtracted from your claim on the future. This is the arithmetic the curve-seller never shows you, and it is the whole of the matter.

IV. Power Laws: What They Are and Why People Misuse Them

Of all the curves fitted to history and sold as destiny, the power law has lately become the favourite, and it deserves a section of its own, because the abuse of it is so widespread and the genuine article so misunderstood.

A power law is a relationship of the form y equals a times x raised to a fixed power b — in symbols, y = ax^b. Its defining property is scale invariance: stretch the input by any factor and the output stretches by a constant factor of its own, regardless of where on the scale you began. This is the mathematical signature of a system with no characteristic size, no natural unit at which things change their behaviour. And it has a seductive visual fingerprint. Take logarithms of both axes and the power law becomes a straight line, its slope equal to the exponent. It is this last fact that has caused most of the trouble, for it means that anyone who can plot data on log-log axes and lay a ruler against it can persuade himself, and a credulous audience, that he has found a law of nature.

Power laws are real, and where they are real they are genuinely remarkable. The frequency of earthquakes falls off as a power of their magnitude — the Gutenberg–Richter relation — so that great quakes are rare and small ones common in a fixed mathematical proportion. The sizes of cities, ranked from largest to smallest, trace something close to a power law, the regularity Zipf made famous. Certain networks have degree distributions that follow power laws over a range. Some quantities in biology scale as powers of body mass. These are not numerological curiosities. They are clues to mechanism.

And that is the first thing the abusers forget: a power law is not a mystical universal that descends upon any sufficiently grand process. It is the fingerprint of a particular kind of generating mechanism, and only of that kind. Power laws arise from specific causes — from multiplicative growth checked by constraints; from preferential attachment, the rich-get-richer dynamic in which the already-large grow fastest; from systems poised at a critical point, balanced on the knife-edge between order and disorder; from certain optimisations under cost. Where one of these mechanisms operates, a power law is the expected and intelligible result. Where none of them has been shown to operate, a straight-ish line on a log-log plot is not evidence of a hidden law. It is a shape, and shapes are cheap.

This is the heart of the misuse. The market mystic reasons: the chart looks straight in log-log; therefore a power law governs the system; therefore the future is constrained to follow the line. Every step of that inference is unearned. The first error is empirical. A straight line on doubly logarithmic axes is one of the weakest patterns in all of statistics, because logarithms are merciful: they compress vast ranges into small spaces and flatter all manner of curves into apparent straightness. Over a limited span — and the price history of any modern speculative asset is a very limited span — a great many distributions masquerade convincingly as power laws. The log-normal distribution in particular, the natural result of any process built from many multiplied random factors, can imitate a power law over orders of magnitude so faithfully that the eye, and even a careless regression, cannot tell them apart. To distinguish a true power law from its impostors requires not a ruler but a great deal of data and a great deal of care.

It was precisely this care that Aaron Clauset, Cosma Shalizi, and Mark Newman brought to the question in their 2009 paper in SIAM Review, the most important corrective in the modern literature and one that ought to be read by everyone tempted to wave a log-log plot. Their argument was twofold and devastating. First, the method by which power laws are usually “found” — fitting a straight line to log-log data by least squares — is statistically unsound: it produces biased estimates of the exponent and, worse, gives no indication whatever of whether the data obey a power law at all, since it never asks the question. It will return a slope for data that are not remotely power-law distributed, and the slope will look just as authoritative as a real one. Second, when one does the work properly — estimating the exponent by maximum likelihood, measuring goodness of fit with a proper statistic, and, crucially, testing the power law against serious alternatives such as the log-normal and the stretched exponential — a great many of the power laws confidently claimed in the scientific literature simply evaporate. For dataset after dataset, the data are fit at least as well, and often better, by an alternative that no one had bothered to rule out. The lesson is not that power laws do not exist. It is that almost no one who claims one has earned the claim, because almost no one has done the test that could have refuted it.

The point has been made again and again by those who study these distributions for a living. Stumpf and Porter, writing in Science in 2012 under the deliberately blunt title “Critical Truths about Power Laws,” observed that a convincing demonstration of a power law requires far more than a plausible-looking straight line: it requires a statistically sound fit over a sufficient range and, ideally, an account of the mechanism that would produce it. Absent the mechanism, even a clean fit is merely descriptive — and the mechanisms behind several of the textbook examples remain contested to this day. Whether the famous “scale-free” networks are as common as a generation of enthusiasts believed is itself now seriously doubted. Even the canonical cases, in other words, are harder to establish than the popular accounts pretend. So when a forecaster takes a decade or so of prices, plots them on log-log axes, lays his ruler against the smear, and announces a law of the asset’s nature, he has not done the hard part. He has skipped it. He has performed the one procedure that the people who actually understand power laws identified, years ago, as the procedure that proves nothing.

V. The Bubble Problem

History has run the experiment we need, over and over, for four centuries, and it always returns the same verdict. During the inflation of a speculative bubble, prices rise with a smoothness and an apparent inevitability that invite exactly the mathematical regularities the modellers love; and after the bubble bursts, the regularity vanishes as though it had never been, because it never was a law — only the trajectory of a particular madness, fitted in flight.

Begin with the tulips, though we must begin honestly, for even here the story is cleaner than the facts. In the popular telling, the Dutch Republic of the sixteen-thirties lost its collective mind over flower bulbs, drove their prices to the cost of houses, and was ruined in the crash of 1637. The careful history — Anne Goldgar’s archival work is the modern correction — shows a mania more confined and a collapse less catastrophic than the legend insists; the ruin was largely a ruin of reputations and contracts, not of the economy. The qualification matters, and it matters in a way that serves our theme exactly: the smooth ascending curve of tulip prices is itself partly a retrofitted artefact, a tidy line drawn through a messier reality by people who wanted a tidy lesson. The bubble narrative, like the bubble model, is more orderly in the retelling than it ever was in the living.

The pattern, with that caution entered, recurs with grim reliability. In the railway mania of the eighteen-forties, British investors poured their savings into railway shares on the unanswerable logic that railways were the future — which they were — and that the shares must therefore rise without end — which they did not. During the boom the rise looked structural, almost a law of progress; in the bust it was revealed as an ordinary speculative overshoot. The South Sea Bubble of 1720 inflated the stock of a company whose actual trade never remotely justified its price, on a tide of official endorsement and contagious greed, and burst spectacularly. Isaac Newton is supposed to have lost a fortune in it and to have remarked that he could calculate the motions of the heavenly bodies but not the madness of men. The line is almost certainly apocryphal — which is itself a fitting detail, a manufactured quotation about manufactured value, passed down because it is too apt to discard. The Mississippi scheme convulsed France in the same years on the same principle. The dot-com boom at the turn of our own century bid the shares of companies with no earnings, and sometimes no revenue, to heights justified entirely by extrapolation of the recent past — until the extrapolation stopped. The housing bubble of the mid-2000s rested on the proposition, encoded in models of impressive sophistication, that house prices in aggregate did not fall: a proposition that was true in the sample to which the models had been fitted and false the moment it mattered.

Lay these episodes side by side and the common anatomy is unmistakable. During the expansion, the trend appears stable, the growth appears inevitable, and clean mathematical relationships emerge from the data and beg to be fitted. Analysts oblige. Lines are drawn, exponents estimated, channels and bands constructed, and each fits the run of prices to that point with gratifying precision. After the collapse, the relationship breaks, the line is overshot or undershot by a margin that no confidence band had contemplated, and the fitted law is quietly abandoned or, more often, quietly recalibrated to the wreckage so that it may be fitted afresh to the next ascent.

Here is the argument the whole section exists to establish, and it is short. The existence of a trend during a bubble does not establish a law. It describes the bubble. A bubble is, almost by definition, a stretch of faster-than-justified growth, and faster-than-justified growth will of course trace out a curve, and the curve can of course be fitted, and the fit will of course look impressive for as long as the bubble lasts. None of this is evidence of an underlying necessity. It is the mathematical shadow of a crowd inflating a price, and when the crowd reverses, the shadow goes with it. To fit a curve to the inflating phase and call it a law is to mistake the symptom for the cause, the trajectory for the engine, the weather of a particular folly for the climate of the world.

In fairness — and the fairness sharpens rather than blunts the point — there exists an honest way to bring power laws to bubbles, and it throws the dishonest way into relief. Didier Sornette and his collaborators have for years modelled bubbles as faster-than-exponential growth decorated with accelerating oscillations and tending toward a critical point at which the bubble becomes unstable. Whatever one makes of its track record, which is contested, the Sornette programme does the thing the mystics will not: it commits, in advance, to dated and falsifiable statements about when a regime is likely to end, and it accepts the scoreboard when the date arrives. That is the difference between a scientist using a power law and a salesman wielding one. The scientist sticks his neck out and lets the future cut it off if it will. The salesman draws his line through the past and, when the future declines to cooperate, draws a new one.

VI. Moving Goalposts and the Illusion of Validation

We come now to the mechanism by which a worthless model achieves immortality, and it is worth dwelling on, because it is the engine that keeps the whole enterprise running long after it should have died. A model that makes a fixed claim can be killed by a single contrary fact. A model whose claim is allowed to move can never be killed at all. Most market “laws” survive not because they are right but because their custodians retain, and freely exercise, the right to redefine what counts as success.

The methods of rescue are a small and predictable repertoire, and once you have seen them named you will see them everywhere. There is boundary adjustment: the price strays above the line, so the line acquires an upper band, and above the band a further band, until the “law” is no longer a line but a corridor wide enough to contain whatever the price does. There is confidence-interval inflation: the original prediction was a value, then it was a range, then it was a range so generous that no outcome could fall outside it, at which point the prediction predicts nothing while appearing to predict everything. There is parameter recalibration: the exponent that fitted the data through last year no longer fits, so the exponent is revised, and the revised model is presented as the same triumphant law rather than as the confession of failure that every recalibration actually is. There is start-date migration: begin the fit in 2011 and it works; when it stops working, begin it in 2012, or 2015, and announce that the “true” law was always the one anchored to the new origin. And there is the simple excision of the inconvenient: the observations that violate the model are reclassified as anomalies, special cases, exogenous shocks, this-time-is-different exceptions, and removed from the reckoning, so that the law is judged only on the data that flatter it.

Each of these moves has the same structure and the same cost. Each preserves the appearance of success, and each, in exactly the same motion, destroys whatever predictive content the model had. For the content of a prediction lies entirely in what it forbids. A claim that rules out nothing carries no information; a claim that rules out everything except what happened carries no information either, since it was fitted to what happened. Every widening of the band, every revision of the exponent, every shift of the origin, every quiet deletion of an embarrassing point, is a transfer of the model from the column of things that could have been wrong into the column of things that cannot be wrong — and a thing that cannot be wrong cannot be right, because in matters of fact rightness is just the property of having survived the chance of being wrong.

This is the insight that Popper made the cornerstone of his philosophy, and he came to it, by his own account, by watching exactly this kind of intellectual machinery at work. What struck him about the grand explanatory systems of his Vienna — he named Adler’s psychology, Marx’s theory of history, Freud’s psychoanalysis — was not that they explained too little but that they explained too much. Whatever a man did, the theory could account for it; the same act and its opposite were equally confirming; no conceivable observation was inconsistent with the doctrine. To its adherents this universal explanatory reach felt like overwhelming strength. Popper saw it for what it was: the absence of any empirical content at all. A theory compatible with every possible state of the world says nothing about which state obtains. The apparent strength was the disease. The power to explain anything is identical to the inability to predict something, and the two are not opposites but the same fact seen from two sides.

Imre Lakatos refined the diagnosis into a distinction we badly need here, between research programmes that are progressive and those that are degenerating. The test is not whether a programme ever meets an anomaly — all of them do — but what it does in response. A progressive programme answers an anomaly by proposing a modification that goes beyond the anomaly, that predicts some new and independent fact which then turns out to be true, so that the patch buys you more than it cost. A degenerating programme answers an anomaly by proposing a modification whose only function is to absorb that anomaly, that predicts nothing new, that exists solely to stop the bleeding. The serial recalibration of a fitted curve is the very type of the degenerating programme. Each adjustment is purely defensive. None of them ever yields a fresh, checkable prediction that pays its way. The model is not learning. It is retreating, in good order, forever.

State the principle in the bluntest possible form and nail it over the door of every forecasting shop: a model that cannot fail cannot succeed. The capacity to be refuted is not a weakness a theory must regrettably tolerate. It is the precise thing that makes a theory worth holding, the membership fee for the company of serious ideas. A model whose advocates have arranged matters so that no outcome can ever count against it has not been strengthened by their ingenuity. It has been quietly removed from the realm of knowledge and re-housed in the realm of faith, where it may comfort its believers but can no longer inform anyone.

VII. Why Backfitted Models Convince

If backfitted models are so empty, why are intelligent people so reliably taken in by them? Not because those people are stupid, but because the human apparatus of judgement is fitted with a series of biases that the fitted curve exploits with almost diabolical precision. The persuasiveness of these models is not an accident of presentation. It is the predictable output of well-documented flaws in how we weigh evidence.

Begin with confirmation bias, the deepest of them, the tendency to seek, notice, and remember what agrees with a belief we already hold and to overlook, discount, and forget what does not. Show a believer a year in which the curve held and you have given him a vivid confirmation he will treasure; show him a year in which it failed and he will find a reason the year does not count, and the reason will feel, to him, like rigour rather than rescue. The model is never tested in his mind, only celebrated, because the mind that holds it is built to gather evidence for the defence and to lose the evidence for the prosecution.

Compounding it is selection bias in the evidence itself, of which survivorship bias is the most instructive variety, and it carries a famous illustration. During the Second World War the statistician Abraham Wald was asked to advise on where to add armour to bombers, given a survey of the bullet holes in the planes that came back. The instinct was to armour where the holes clustered. Wald saw the trap: the survey covered only the survivors. The planes hit in those places had returned to be counted; the planes hit elsewhere — the engines, the cockpit — had not come back at all, and so left no holes in the data. The armour belonged precisely where the returning planes were unscathed. The lesson generalises to every forecasting record we are ever shown. We see the models and the modellers that survived, the calls that happened to come good, the curves still standing because they have not yet been overshot. The failures have crashed silently out of view. A presented track record is a record of survivors, and to reason from it as though it were the whole population is to armour the wings and lose the plane.

Authority bias does its part. A claim advanced by someone eminent, confident, and credentialed is granted a weight its content has not earned, and the curve-seller is typically all three. Availability does more: we judge the probability of a thing by how easily examples come to mind, and the vivid, oft-repeated success of a famous forecast is far more available than the diffuse, unremembered failures, so that the mind’s estimate of the model’s reliability is corrupted by the mere memorability of its hits — the bias Tversky and Kahneman identified and that Gigerenzer has spent a career teaching people to correct. We remember the prediction that landed and forget the hundred that did not, and from this asymmetric memory we construct a wholly false impression of a method’s power.

Over all of this presides what Taleb named the narrative fallacy: the irresistible human compulsion to wrap a sequence of events in a story, to convert the random into the inevitable by the simple act of telling it backward. After the fact, every crash has its clear cause, every boom its obvious logic, every price its reason — not because the causes were knowable in advance, but because the mind cannot tolerate a sequence without a plot and will supply one whether the world provides it or not. The fitted curve is the narrative fallacy rendered in mathematics. It tells the story of the past as a smooth necessity, and the smoothness is supplied by the teller.

And the mathematics is the final and most powerful seduction, for numbers wear the mask of objectivity. A story told in words invites argument; the same story told in equations, exponents, and goodness-of-fit statistics seems to have descended from some realm above human motive, impartial and exact. But a fitted curve is not an oracle. It is a rhetorical device with a lab coat on. The precision of its decimals says nothing about the truth of its claim; a meaningless quantity can be computed to ten places. When mathematics is used not to derive a consequence and submit it to test but to dress a backward-looking story in the costume of rigour, it has ceased to be a tool of inquiry and become an instrument of persuasion — and it is the more dangerous for being so widely mistaken for the opposite.

That none of this is idle theory was shown, at length and at cost, by Philip Tetlock’s long study of expert prediction, which followed the forecasts of hundreds of specialists across many years and many domains. The findings are humbling and ought to be famous. The average expert predicted the future about as well as a reasonable guess, and sometimes worse; and — the detail that bears directly on our subject — the experts who were most confident, most famous, and most committed to a single grand explanatory framework were among the least accurate of all, while doing the best business in attention. The market for forecasts, Tetlock’s work makes plain, does not reward being right. It rewards being vivid, confident, and memorable, which are exactly the qualities a well-fitted curve and a self-assured curve-seller possess in abundance, and exactly the qualities that have nothing to do with whether tomorrow will obey.

VIII. The Economics of Selling a Model

We must now be impolite and follow the money, because the persistence of empty models is not finally explained by cognitive bias alone. Bias explains why the audience is willing to be fooled. Incentive explains why there is always someone willing to do the fooling, and to keep doing it. A model is not only an instrument of analysis. It is a product, and behind every durable product stands an interest in its survival.

Consider what a successful model returns to the man who sells it. Followers, first — an audience that grows with each retelling. Attention, the currency of the age. Influence over the decisions of others, which is to say power. And beneath these, the harder rewards: consulting fees, subscription revenue, the capital that flows toward a confident forecaster, the speaking engagements, the standing of the seer. A model that has acquired a following has acquired, for its proprietor, an income and an identity. And no man with an income and an identity bound up in a thing is a disinterested judge of whether the thing is true.

Upton Sinclair put the principle once and for all: it is difficult to get a man to understand something when his salary depends on his not understanding it. The custodian of a popular model is in precisely that position with respect to the model’s falsity. Every instinct of self-interest pulls him toward preservation and away from refutation. When the data turn against the model, he is not moved, as a scientist is supposed to be moved, to ask whether the model is wrong; he is moved to find the adjustment that will save it, because the model is not merely his theory but his livelihood and his name. The recalibrations we catalogued earlier are not, at bottom, intellectual errors. They are the rational responses of a man defending an asset.

The structure is a textbook principal–agent problem. The forecaster, the agent, is rewarded for the appearance of insight, not for the substance of it, and the two come apart precisely when it matters. His audience, the principals, cannot easily tell a genuine edge from a lucky run or a flattering fit, and by the time the truth arrives the fees have been collected and the attention spent. The agent’s optimal strategy, given all this, is not to be right. It is to seem right for as long as possible and to have an exit ready when seeming fails — and a model whose success criteria can be quietly moved provides exactly that: a way to seem right indefinitely.

Reputational lock-in seals it. A forecaster who has staked his public identity on a model cannot abandon it without abandoning himself. The more loudly and the longer he has championed it, the more the model becomes load-bearing for his standing, and the more unthinkable its repudiation grows. He is chained to his own past confidence. To recant would be to confess that the followers were misled and the fees unearned, and few men will pay that price while any recalibration remains available to postpone it. So the model is defended past all reason, not in spite of the seller’s interests but in faithful service to them.

Keynes saw the general law behind this and stated it with his usual cool: worldly wisdom teaches that it is better for reputation to fail conventionally than to succeed unconventionally. The forecasting industry runs on that wisdom. The analyst who hedges, who issues the vague and unfalsifiable call, who stays with the herd, suffers nothing when he is wrong, because everyone was wrong together and no particular blame attaches. The analyst who commits to a sharp, checkable prediction and misses is exposed and remembered. The payoffs are therefore brutally asymmetric, and they select, with perfect efficiency, against precisely the behaviour that science requires. A bold, falsifiable forecast is a career risk with little private upside and large private downside; a fitted curve with movable goalposts is a career asset that pays whether or not the world cooperates. The rational forecaster, responding to these incentives, will make exactly the unfalsifiable, endlessly adjustable, retrospectively fitted claims we have been anatomising — not because he is a fool, but because he is not one. And because there is, in this industry, almost no penalty for being wrong — no licence revoked, no fee returned, no record kept by anyone but the obscure and the obsessive — the supply of confident curves is, and will remain, effectively unlimited. The wonder is not that so many worthless models are sold. The wonder is that anyone expected otherwise.

IX. A Bayesian Reckoning

It is time to make the argument precise, and the right instrument is Bayes’s theorem, because it forces into the open the question the curve-seller most wants to keep hidden: not “does the model fit?” but “how surprised should I be by this fit under each of the explanations available to me?” That is the only question that bears on belief, and the fit, taken alone, cannot answer it.

Bayesian reasoning, stripped to its bones, says that the credibility a piece of evidence lends a hypothesis depends on two quantities and their comparison across rivals. The first is the likelihood: how probable is this evidence if the hypothesis is true? The second is the prior: how probable was the hypothesis before this evidence arrived? And the decisive move is comparative — we do not ask how well the evidence fits our favoured hypothesis in isolation, which is the chronic error, but how much better or worse it fits ours than it fits the alternatives. Evidence supports a hypothesis only to the extent that it is more expected under that hypothesis than under its competitors. A fact equally expected under every explanation discriminates among none of them, however dramatic it looks.

Set the two relevant hypotheses against each other. Theory A holds that the asset’s price is governed by a genuine law, a stable mechanism that constrains its path and will go on constraining it. Theory B holds that the price is the output of a speculative process — a crowd of human beings buying because the price is rising and selling because it is falling, generating, as such crowds do, runs and trends and temporary regularities that persist for a while and then dissolve. Now ask the Bayesian question of the evidence on offer, which is a clean curve fitted to the price history. Under Theory A, such a fit is expected. But under Theory B it is also expected — entirely, unremarkably expected — because speculative processes routinely throw off clean-looking trends over their runs; producing a stretch of data that a curve fits well is one of the most ordinary things a bubble does. The likelihood of the observed fit is high under both hypotheses. The evidence is therefore almost worthless for telling them apart. The Bayes factor sits near one; the fit that the modeller presents as decisive is, in the only calculation that matters, nearly silent. If anything it tilts the wrong way, since over a short and recent window the mundane process has the easier time manufacturing a persuasive curve.

Then weigh the priors, and the case against the law steepens. Theory A is an extraordinary claim. It asserts that the price of a speculative asset — a quantity set, moment to moment, by the shifting beliefs and appetites of millions of people — is in the grip of a fixed mathematical necessity of the kind we otherwise find only in the deep regularities of physics. That is a very strong claim about the world, and it should carry a correspondingly low prior. Laplace stated the standard two centuries ago and Sagan made it a slogan: the weight of evidence required for an extraordinary claim must be proportioned to its strangeness; extraordinary claims demand extraordinary evidence. Historical fit is not extraordinary evidence. It is the most ordinary evidence there is — exactly what the unremarkable hypothesis predicts anyway. To overturn a low prior one needs evidence that the mundane alternative would be very unlikely to produce, and the fitted curve is the opposite of that: evidence the mundane alternative produces as a matter of course.

This is the same truth that Deborah Mayo has pressed under the banner of severe testing, and her formulation is the most useful single criterion in this whole essay. A hypothesis is corroborated by passing a test only if the test was severe — only if, were the hypothesis false, it would very probably have failed that test. A test the hypothesis would pass comfortably whether it were true or false is no test at all; clearing it tells us nothing. Fitting a curve to the data the curve was chosen to fit is the least severe procedure imaginable. The fit could not have failed, because the parameters were selected, after the fact, to guarantee it. A claim that has survived only such a non-test has survived nothing, and deserves not one degree of added belief. The proper response to “but look how well it fits” is therefore not admiration but a single question: under what competing account of this data would the fit have been surprising? If the answer is none — if the speculative, law-free hypothesis predicts the very same fit — then the fit, however lovely, is not evidence of a law. It is evidence that a curve can be drawn, which we never doubted.

X. What a Real Law Looks Like

It will sharpen everything if we set beside the counterfeit a portrait of the genuine article, so that the contrast can do its work. What does a real law look like — the kind of regularity that has earned the name, as against the kind that has merely borrowed it? A genuine law has, at least, the following marks.-

It is predictive. It tells you something about cases you have not yet seen, and it told you before you saw them.

-

It is falsifiable. There is a clearly specifiable observation that, if it occurred, would show the law to be false — and the believers can tell you what it is.

-

It is stable. Its content does not have to be revised every time a new observation arrives; the same law, unaltered, keeps accounting for fresh cases.

-

It is mechanistic. It is connected to an understanding of why the regularity holds, not merely that it has held, so that the regularity is intelligible and not just observed.

-

It is robust to new data. It goes on working outside the range and the era in which it was first found, on data that played no part in its construction.

-

It is independent of parameter manipulation. Its success does not depend on tuning dials to each new dataset; it holds with its parameters fixed, or with no free parameters at all.

Measure the great laws against this list and they pass without strain. Newtonian gravitation did not merely fit the planetary motions known when it was framed; it predicted motions no one had yet seen. When the orbit of Uranus misbehaved, the law was not patched — it was trusted, and the trust was specific: Le Verrier and Adams used it to compute where an unseen planet must be to cause the discrepancy, and when the telescopes were pointed at the calculated spot, Neptune was there. The same law had already told Halley that his comet would return, decades ahead, and it did. These are predictions in the strong sense: claims about the unseen, made in advance from a fixed theory, that the world then confirmed. No dials were turned. The law stuck out its neck a century at a time and kept its head.

The laws of thermodynamics are sterner still, and they show another mark of the genuine: they forbid. The second law does not fit a curve to anything; it prohibits a class of outcomes outright — no engine shall convert heat wholly into work, no machine shall run forever on nothing. Every attempt to build a perpetual-motion machine has failed, and each failure is not an awkward anomaly to be explained away but a fresh confirmation of the prohibition. This is the signature that inverts the behaviour of the fitted curve. A real law grows stronger every time reality is given the chance to break it and declines to. A backfitted model grows weaker every time reality is consulted, which is why its custodians work so hard to keep reality from being consulted cleanly.

Two of the examples often grouped with these deserve a more careful word, because intellectual honesty forbids lumping different things under one triumphant heading. Information theory — Shannon’s — is frequently and rightly cited as a body of deep, predictive law, and so in a sense it is: his coding theorems fix exact limits on how much information a channel can carry and how far a message can be compressed, and those limits bind every communication system ever built, predicting with certainty what no engineering cleverness can exceed. But they are, strictly, mathematical theorems — propositions proved from definitions — rather than empirical laws wrung from observation in the manner of gravitation or thermodynamics. Their predictive force is real, but it is the force of proof, not of induction, and the distinction is worth keeping rather than blurring. Game-theoretic equilibria deserve a sharper caution still. As mathematics, the existence of equilibrium in broad classes of games is a theorem, secure. As a description of what actual human beings do, equilibrium is a model of variable and often disappointing accuracy; whether real agents play the predicted strategies is an empirical question whose answer is frequently no. To wave at “game-theory equilibria” as though they were laws of nature on the footing of the second law of thermodynamics would be to commit, in miniature, the very confusion this essay condemns — to mistake a framework that can be fitted to behaviour for a necessity that governs it. The honest taxonomy keeps three things apart: empirical laws that forbid and predict, like gravitation and thermodynamics; mathematical theorems that bind with the certainty of proof, like Shannon’s limits; and modelling frameworks whose empirical standing must be earned case by case, like equilibrium in economics. Only the first two are “laws” in the sense the curve-seller is trying to borrow, and he has earned neither.

The deepest contrast is the one already glimpsed and now made general. Genuine laws become stronger when challenged. General relativity was a rigid theory that risked everything on advance predictions, and a century of attempts to break it has instead extended it: the bending of light, the slowing of clocks in gravity, and in our own time the direct detection of gravitational waves by LIGO in 2015 and the imaging of a black hole’s shadow in 2019 — each a place the theory could have been wrong and was not, each a tightening rather than a loosening of our confidence. That is what corroboration looks like: a fixed claim, repeatedly exposed to the chance of refutation, repeatedly surviving. The fitted market model exhibits the exact opposite biography. It does not risk and survive; it accommodates and retreats. Each new observation, instead of testing it, occasions another adjustment of it. Where the real law is strengthened by confrontation, the counterfeit must be shielded from it. That asymmetry is not a detail. It is the whole difference between a law and a story, written in the way each behaves when reality is finally allowed into the room.

XI. The Test of Scientific Honesty

All of this can be compressed into a handful of questions — a short interrogation to which any model claiming the dignity of a law should be subjected, and which it should be made to answer before it is believed, funded, or followed. The questions are simple. The discomfort they produce in the wrong sort of model is the point.-

Were the parameters fixed in advance, or chosen after the data were seen? A model whose dials were set before the fit risked something; a model whose dials were set to produce the fit risked nothing.

-

Was the hypothesis stated before the test, or assembled afterward to match what happened? Prediction registered ahead of the fact is evidence; explanation composed after it is decoration.

-

Can the model be falsified, and will its advocates say how? Ask them, plainly: what observation would prove this wrong? If no such observation exists, or none can be named, the model is not a claim about the world but an article of faith.

-

How many times have the parameters been changed? Each recalibration is a failure the model already suffered and survived only by being altered. A long history of revisions is a long confession.

-

How many failed versions preceded the one now on display? The graveyard is never shown. Ask to see it. A model is only as impressive as the count of its discarded predecessors allows it to be.

-

Has the model been tested out of sample — on genuinely held-back data, from a different period or domain, that played no part in fitting it? In-sample success is not evidence; it is the thing to be explained. Only out-of-sample success counts.

-

Can independent researchers, with no stake in the outcome, reproduce the result from the data and the stated method? A result that survives only in the hands of its proprietor is not a finding. It is a performance.

Where these questions can be answered well, a model has begun — only begun — to earn its standing. Where they cannot, the proper attitude is not hostility but a settled, unembarrassed scepticism. The burden lies on the claim, not on the doubter. A model that will not submit to this interrogation has told you, by its refusal, exactly what it is.

XII. A Curve Is Not a Law

The history of science is not, whatever the textbooks’ tidy diagrams suggest, the history of drawing lines through data. It is the history of discovering mechanisms — accounts of why the world behaves as it does — and then sending those accounts out, defenceless and specific, to be confronted by a reality with no obligation to be kind. The honour belongs not to the men who fitted the past most snugly but to the men who told the future most riskily and were not destroyed.

So when a fitted power law is laid across thirteen years of an asset’s prices and pronounced a law of its nature, an inevitability, a glimpse of the decades to come, the right response is to measure the claim against everything we have said. A power law fitted to thirteen years of observations is not evidence of inevitability. At the very most it is evidence that a mathematical curve can be drawn through thirteen years of observations — which was never in doubt, and which a great many false laws and a great many real bubbles would produce in exactly the same way. The fit establishes the drawability of the curve. It establishes nothing about the necessity of the path.

The decisive question is never whether a model explains yesterday. Every model explains yesterday; that is the one thing every model can do, and the worse the model the more effortlessly it does it, because flexibility and hindsight together can account for anything that has already occurred. The question is whether the model predicts tomorrow — and predicts it without revision, without the band quietly widening, the exponent quietly shifting, the start date quietly migrating, the inconvenient observation quietly removed. The moment a model survives only by adjusting its boundaries, recalibrating its parameters, redefining its criterion of success, or moving its point of origin, it has crossed the line that separates the two kinds of thing this essay has tried to keep apart. It has ceased to be a law and become a narrative — a story told backward, in the language of mathematics, about why what happened had to happen.

There lies the whole of it, and it can be put in a sentence. Science advances through prediction. Pseudoscience survives through explanation. The one stakes itself on the unknown and grows stronger when reality confirms it; the other feeds on the known and grows more elaborate each time reality refuses. They can look identical on the page — the same axes, the same exponents, the same confident decimals — and they are nonetheless opposites, divided by a question that has nothing to do with how the curve looks and everything to do with what it dared. Did it tell you the number before the number was known? Or did it wait, as the charlatan always waits, until the answer was safely on the table, and then draw its elegant line through the past, and ask you to mistake a memory for a prophecy?

A curve is not a law, any more than a portrait is a prophecy or an obituary is a life. It is a record of where the points have already fallen. Whether the next point will fall there too is a question the curve cannot answer and was never built to ask — and the surest sign that you are in the presence of a fraud rather than a finding is that the curve is offered as the answer all the same.


Sources and Further Reading

Box, G. E. P. (1976). “Science and Statistics.” Journal of the American Statistical Association, 71(356), 791–799. https://doi.org/10.1080/01621459.1976.10480949

Box, G. E. P., & Draper, N. R. (1987). Empirical Model-Building and Response Surfaces. New York: Wiley. (Source of “Essentially, all models are wrong, but some are useful,” p. 424; ISBN 0-471-81033-9.)

Broido, A. D., & Clauset, A. (2019). “Scale-free networks are rare.” Nature Communications, 10, 1017. https://doi.org/10.1038/s41467-019-08746-5

Clauset, A., Shalizi, C. R., & Newman, M. E. J. (2009). “Power-Law Distributions in Empirical Data.” SIAM Review, 51(4), 661–703. https://doi.org/10.1137/070710111

Dyson, F. J. (2004). “A Meeting with Enrico Fermi.” Nature, 427, 297. https://doi.org/10.1038/427297a (The source of von Neumann’s elephant remark, by way of Fermi.)

Feynman, R. P. (1974). “Cargo Cult Science.” Commencement address, California Institute of Technology; reprinted in Surely You’re Joking, Mr. Feynman! New York: W. W. Norton (1985).

Friedman, M. (1953). “The Methodology of Positive Economics.” In Essays in Positive Economics. Chicago: University of Chicago Press.

Gelman, A., & Loken, E. (2014). “The Statistical Crisis in Science.” American Scientist, 102(6), 460–465. https://doi.org/10.1511/2014.111.460

Gigerenzer, G. (2002). Reckoning with Risk: Learning to Live with Uncertainty. London: Penguin. (US edition: Calculated Risks: How to Know When Numbers Deceive You. New York: Simon & Schuster.)

Goldgar, A. (2007). Tulipmania: Money, Honor, and Knowledge in the Dutch Golden Age. Chicago: University of Chicago Press.

Keynes, J. M. (1936). The General Theory of Employment, Interest and Money. London: Macmillan. (The remark on failing conventionally is in ch. 12.)

Lakatos, I. (1970). “Falsification and the Methodology of Scientific Research Programmes.” In I. Lakatos & A. Musgrave (eds.), Criticism and the Growth of Knowledge (pp. 91–196). Cambridge: Cambridge University Press.

Mangel, M., & Samaniego, F. J. (1984). “Abraham Wald’s Work on Aircraft Survivability.” Journal of the American Statistical Association, 79(386), 259–267. https://doi.org/10.1080/01621459.1984.10478038

Mayer, J., Khairy, K., & Howard, J. (2010). “Drawing an Elephant with Four Complex Parameters.” American Journal of Physics, 78(6), 648–649. https://doi.org/10.1119/1.3254017

Mayo, D. G. (2018). Statistical Inference as Severe Testing: How to Get Beyond the Statistics Wars. Cambridge: Cambridge University Press.

Popper, K. R. (1959). The Logic of Scientific Discovery. London: Hutchinson. (German original Logik der Forschung, 1934.) See also Conjectures and Refutations. London: Routledge & Kegan Paul (1963).

Sagan, C. (1980). Cosmos. New York: Random House. (The “extraordinary claims require extraordinary evidence” formulation.)

Silver, N. (2012). The Signal and the Noise: Why So Many Predictions Fail — but Some Don’t. New York: Penguin Press.

Sinclair, U. (1935). I, Candidate for Governor: And How I Got Licked. (The remark on salary and understanding.)

Sornette, D. (2003). Why Stock Markets Crash: Critical Events in Complex Financial Systems. Princeton: Princeton University Press.

Stumpf, M. P. H., & Porter, M. A. (2012). “Critical Truths About Power Laws.” Science, 335(6069), 665–666. https://doi.org/10.1126/science.1216142

Taleb, N. N. (2001). Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere.

Taleb, N. N. (2007). The Black Swan: The Impact of the Highly Improbable. New York: Random House.

Tetlock, P. E. (2005). Expert Political Judgment: How Good Is It? How Can We Know? Princeton: Princeton University Press.

Tversky, A., & Kahneman, D. (1974). “Judgment under Uncertainty: Heuristics and Biases.” Science, 185(4157), 1124–1131. https://doi.org/10.1126/science.185.4157.1124


Tags: backfitting; retrodiction; power laws; statistical inference; data mining; overfitting; predictive models; scientific method; financial bubbles; market forecasting; falsifiability; Bayesian inference; econometrics; model risk; speculative assets; scientific realism; forecasting.


← Back to Substack Archive