The Information AI Cannot Learn
Foundation models do not merely inherit data. They inherit the economic mechanism that produced the data, including its missing counterfactuals.
Keywords
information economics, foundation models, AI allocation, mechanism design, administrative data, market signals, counterfactual bias, welfare loss, recursive training, platform dependence, provider lock-in, digital policy
Abstract
A foundation model can only learn from the information contained in its training data. That sounds obvious until the data are treated not as neutral records, but as outputs of economic mechanisms. Market mechanisms, administrative systems, procurement platforms, and firm-level provider arrangements generate different kinds of information. Some generate marginal-value signals. Some record only realised outcomes. The distinction matters because allocation requires more than predicting what happened before. It requires knowing what would happen if resources were moved somewhere else. This post explains the core idea behind “Information Acquisition, Mechanism-Generated Data, and Counterfactual Allocation Bias in Foundation-Model Systems”: foundation models cannot recover counterfactual marginal values that the data-generating mechanism never acquired. The issue is not mainly model size, sample size, or compute. It is the information structure generated by the mechanism. A market mechanism with local bid elicitation can generate counterfactual marginal-value information. An administrative record of realised allocations and realised utilities usually cannot. The result is an irreducible counterfactual-allocation bias. This has consequences for public-sector AI, platform dependence, procurement policy, and firm strategy in foundation-model markets.
The Information AI Cannot Learn
Most discussions of artificial intelligence begin with the model. How many parameters does it have? How large was the training corpus? How accurate is it on benchmarks? How well does it generalise? Those are useful questions, but they are not the first economic questions.
The first economic question is different: what information did the data-generating mechanism actually acquire?
That question changes the debate. It moves the focus away from the model as a magical extractor of hidden truth and towards the institutional process that produced the observations in the first place. A foundation model does not train on economic reality directly. It trains on records. Those records are generated by markets, platforms, administrative agencies, firms, procurement systems, recommendation engines, surveys, payment systems, pricing rules, and reporting processes. Each of these mechanisms selects what is observed, what is left unobserved, what incentives agents face when producing the signal, and what counterfactual variation is created.
The crucial point is simple. Data are not neutral. Data are mechanism-generated.
A market price is not merely a number. It is the result of a bidding process, an opportunity-cost structure, scarcity, substitution, budget constraints, and strategic participation. A government administrative record is also not merely a number. It is the result of a policy rule, eligibility criteria, observed outcomes, reporting systems, compliance incentives, and the particular allocation chosen by the administrative authority. A platform record is different again. It may show clicks, conversions, churn, usage intensity, subscription status, and model calls, but it may not reveal the counterfactual demand that would have appeared under a different platform, a different price, a different ranking, or a different provider.
The paper behind this post develops that distinction into an information-economics argument. It asks what happens when foundation-model systems are used for allocation: pricing, recommendation, procurement, supply-chain adjustment, labour matching, credit allocation, resource prioritisation, and administrative decision-making. These tasks are not merely prediction tasks. They are counterfactual allocation tasks. The planner does not only need to know what happened. The planner needs to know what would happen if resources moved elsewhere.
That is where the difficulty begins.
A model trained on observed administrative outcomes may be very good at predicting outcomes inside the historical policy path. It may reproduce patterns in the data. It may rank cases in ways that look plausible. It may forecast the next step under the existing regime. But if the historical regime never generated information about off-path allocations, the model cannot infer that missing marginal information by scale alone. Larger data do not necessarily solve the problem. More parameters do not necessarily solve the problem. More compute does not necessarily solve the problem.
The model inherits the information structure of the mechanism that generated the data.
What the paper is about
The paper’s central claim is that foundation-model allocation systems should be analysed as systems of information acquisition. The relevant question is not only whether a model predicts well. The relevant question is whether the economic mechanism that produced the training data acquired the counterfactual marginal-value information needed for allocation.
This matters because allocation requires gradients. In economic terms, a planner needs to know not just the realised value of an allocation, but how value changes when the allocation is perturbed. If a resource is moved from one use to another, does welfare rise or fall? If a firm switches providers, does operational risk decline? If a platform introduces an alternative procurement channel, does the resulting competitive signal reveal a lower cost or a different product-market fit? If an administrative authority allocates services differently, does the marginal recipient gain more than the marginal recipient who loses access?
Those are marginal questions. They are not answered by observing a single realised allocation.
The paper compares two stylised mechanisms. The first is a local market mechanism. It does not need to reveal the entire value function over every possible allocation. That would be too strong and too unrealistic. Instead, it requires local counterfactual information: bid signals, perturbation menus, or tangent-space reports around the relevant allocation. If agents truthfully report values at nearby perturbations, the mechanism can recover marginal-value information, up to a finite-difference approximation error. With sufficiently local perturbation, the gradient becomes identifiable.
The second is an administrative record. The administrative system observes the allocation it made and the realised utility or outcome at that allocation. It may observe those outcomes accurately. It may observe many cases. It may have a large dataset. But unless it generates variation around the allocation, it still sees only realised outcomes under the policy path. It observes what happened at the chosen point. It does not necessarily observe how welfare would change if the allocation moved elsewhere.
That is the identification gap.
The paper does not claim that markets are always better than administrative systems. It does not claim that administrative systems can never generate useful counterfactual information. It does not make a blanket ideological ranking. The claim is narrower and stronger: counterfactual-support generation is a property of the mechanism. Where a mechanism generates marginal-value information, a model may learn it. Where a mechanism does not generate that information, a model cannot recover it merely by statistical scale.
This distinction is the basis of the paper’s formal results.
Why prediction is not enough
The economics of artificial intelligence is often framed through prediction. Agrawal, Gans and Goldfarb made this framing prominent: AI reduces the cost of prediction, and cheaper prediction increases the value of complementary judgement. That is a useful starting point. But allocation is not only prediction. Allocation asks what should be done, and that requires knowing the opportunity cost of the alternatives.
A prediction system trained on administrative data may answer a question such as: given the historical policy rule, what outcome is likely for this case? That is useful within the existing policy support. It does not automatically answer: what would happen if the allocation rule changed?
The difference is not semantic. It is structural.
Suppose an agency has historically allocated a service to one class of applicants and denied it to another. The record will contain detailed outcomes for those who received the service and perhaps some outcomes for those who did not. But unless the agency created exogenous variation, randomised access, local experiments, instrumental shifts, or some other support-generating mechanism, the data may not contain enough information to identify the marginal value of reallocating the service. The model can become highly accurate at reproducing the administrative regime. That does not make it reliable for counterfactual redesign.
The same problem appears in platform markets. A firm may use one foundation-model provider because that provider was easiest to integrate at founding. Its internal data will then reflect the performance, costs, outages, latency, interface constraints, and product possibilities of that provider. It may not contain information about the counterfactual performance of other providers. The firm sees its own path. It does not observe the cross-provider opportunity-cost surface unless it deliberately creates a mechanism to acquire that information: multi-provider sourcing, benchmarking, switching trials, modular architecture, procurement competition, or controlled technical experiments.
The same problem appears in labour allocation. If a platform has historically routed certain workers to certain tasks, a model can learn the historical routing pattern. That does not necessarily reveal whether a different routing would have produced higher output, higher welfare, or better matching. The counterfactual is missing unless the routing system generated variation.
This is why the paper treats foundation models as signal-conditional estimators. They do not stand outside the economic process. They transform a signal into an estimated value function. If the signal lacks counterfactual marginal information, the estimator inherits that absence.
The mechanism-generated data idea
The core concept is mechanism-generated data. A dataset is not only a statistical object. It is an economic object produced by a mechanism. The mechanism determines what agents reveal, what they conceal, what they are incentivised to report, where variation occurs, and what remains off-path.
Hayek’s argument about the price system was that prices communicate dispersed information that no central planner could easily collect. Hurwicz formalised the informational requirements of allocation processes. Blackwell gave economics a way to compare information structures. Grossman and Stiglitz showed that information in markets is costly and that full informational efficiency creates incentive problems. Mechanism design then added a central insight: institutions do not merely use information; they induce it.
The paper extends that line of thought to foundation-model systems. It asks: when an AI system is trained on data generated by a mechanism, what information from that mechanism is actually available for learning?
The answer depends on the signal. A market-like mechanism that elicits local bid schedules or perturbation values can generate information about the slope of the value function. That slope is the gradient. It tells the planner how welfare changes when the allocation moves in each direction. An administrative record, by contrast, may identify only the value at the realised allocation. It tells the planner what happened there. It does not identify the local slope unless the record also contains variation around that point.
The difference is easiest to see with a simple example. Imagine a hill in the dark. An administrative record tells you the height at the place where you are standing. That is a value level. A market-like perturbation mechanism tells you the height a short distance north, south, east, and west. That gives you the slope. From those nearby readings, you can infer which direction rises or falls. Without them, knowing the height where you stand does not tell you which way the hill slopes.
A foundation model trained only on height-at-the-standing-point records may become very good at describing those points. But allocation requires knowing slopes.
That is the missing information AI cannot learn.
The formal argument in plain language
The first formal result is the identification gap. Under a local market mechanism, the system acquires counterfactual information near the realised allocation. If the mechanism elicits truthful values at small perturbations around an allocation, then finite differences recover marginal values up to a small error. With a smooth value function, that error shrinks as the perturbation becomes smaller. In economic terms, the mechanism identifies the direction in which value changes.
Under an administrative record, the signal is weaker. It contains the realised allocation and the realised utility or outcome at that allocation. Two different value functions can agree exactly at the realised allocation while having different slopes away from it. They generate the same administrative record but imply different counterfactual allocation decisions. The administrative record does not distinguish them.
This is not a small-sample problem. It is not a modelling problem. It is not a problem of weak neural networks. It is a non-identification problem. More observations at the same kind of point reduce noise around the realised value. They do not create the missing off-path gradient.
The second result quantifies the worst-case bias. If the value function belongs to a regular class with bounded gradients, then any estimator trained only on administrative records faces a lower bound on off-allocation value error. The bound is proportional to distance from the realised allocation. In the paper, the two-point minimax construction gives a lower bound of Gd/2, where G is the gradient bound and d is the distance from the realised allocation.
The intuition is direct. If the administrative record pins down only the value at one point, then two admissible value functions can pass through that same point and slope in different directions. The further away you move from the observed allocation, the larger the possible value difference becomes. The error grows with distance because the missing information is the gradient.
The third result concerns recursive training. Foundation models are increasingly trained in environments where model-generated content feeds future models. A growing literature, including Shumailov and co-authors, studies model collapse under recursive training. The paper separates a population-level mechanism from finite-sample drift. If the true data-generating distribution lies outside the model’s hypothesis class, then the first projection into that class loses the features that the class cannot represent. Once lost, those features do not return under further population projection. Finite-sample recursion can add stochastic drift, but the fundamental loss begins with misspecification.
This is relevant to economic allocation because the missing features may be precisely the counterfactual structures that matter: multimodality, tail risk, local complementarities, substitution possibilities, or off-path welfare gradients. A model can preserve some moments while losing the structure that matters for decisions.
The fourth result translates the identification problem into welfare loss. In a quadratic environment, if the market-data planner can identify the relevant curvature but the administrative-data planner uses a default curvature, the welfare loss has a closed-form expression. It depends on the mismatch between the true curvature and the planner’s default, and on the variance of preference shocks. Under a uniform spectral-gap condition, the expected welfare loss scales with allocation dimensionality and preference-shock variance.
That is not just formal neatness. It means the problem becomes more serious in high-dimensional allocation settings. Foundation-model systems are attractive partly because they operate across many dimensions. But high dimensionality also increases the number of possible margins on which counterfactual information can be missing.
The simulation evidence
The paper includes a fully reproducible simulation to show the magnitudes implied by the formal results. The simulation is not the source of the theorem. It is an illustration of what the theorem implies under calibrated conditions.
The first simulation visualises the Gd/2 bias bound. As the allocation moves further from the realised point, the worst-case lower bound rises linearly. That is exactly what the theory predicts. The slope depends on the gradient bound. The policy meaning is straightforward: the further a proposed AI-guided policy move is from the support of the training mechanism, the less confidence one should place in conventional predictive precision.
The second simulation illustrates recursive projection loss. A bimodal distribution is projected into a single-Gaussian class. The variance can remain plausible, but the modal structure is lost. That is important because many diagnostics focus on moments or aggregate fit. A model can look stable while losing the structural features that matter for counterfactual decision-making.
The third simulation examines welfare scaling. Welfare loss rises with allocation dimension and with preference-shock variance, consistent with the paper’s quadratic welfare result. This is an important point for applied modelling. The cost of missing marginal information does not remain fixed as systems become more complex. It grows when allocation becomes more multidimensional and when preferences are more variable.
The fourth simulation examines convergence to a bias floor. Increasing administrative sample size reduces estimation variance, but it does not remove the missing-support bias. In the reported calibrated environment, error flattens quickly. That pattern is the empirical signature of an identification problem rather than an ordinary data-shortage problem.
The fifth simulation considers hybrid mechanisms. Real systems are rarely pure markets or pure administration. A platform may use bids in some dimensions and administrative allocation in others. A procurement system may create competition for some services and use fixed internal assignment for others. A public agency may run experiments in one part of the allocation space while relying on historical records elsewhere.
The hybrid result is important because it gives the policy argument a practical form. The answer is not “always use markets” or “never use administrative data.” The answer is: identify where counterfactual marginal information is missing, and design mechanisms to acquire it where it matters most. If cross-dimensional complementarities are strong, a small amount of targeted bid coverage or perturbation can have large value. If dimensions are separable, the design problem becomes more like a dimension-by-dimension cost comparison.
That is the practical policy content of the theory.
The firm-level empirical illustration
The paper also includes a firm-level empirical section. It uses a de-identified dataset of 127 AI-native firms and 179 firm-episode observations. The empirical question is whether single-provider dependence is associated with elevated failure risk. The theory predicts that it should be, not because one provider is necessarily bad, but because dependence on a single provider restricts cross-provider counterfactual information. A firm locked into one foundation-model provider observes its own path through that provider. It does not observe the substitution possibilities, performance gradients, and operational trade-offs across providers unless it creates mechanisms to acquire that information.
The empirical evidence is deliberately secondary. It is observational. It does not identify a causal effect. It contains near-complete separation: 22 of 23 failures occur among single-provider firms. That makes unpenalised estimates unstable. The paper therefore treats the empirical section as external-consistency evidence, not as validation of the theory.
The baseline Cox model reports a large association between single-provider dependence and failure hazard. The ridge-penalised Cox estimate is more defensible than the unpenalised estimate and gives a single-provider hazard ratio of about 6.53. The revised replication package also includes a Firth rare-event robustness result, with a positive single-provider association. The precise magnitude is descriptive. The sign is the relevant point: the observed pattern is consistent with the model’s qualitative prediction that single-provider dependence creates information-acquisition fragility.
This matters because provider dependence is often discussed as an operational risk: outages, pricing power, API changes, contractual vulnerability, or switching costs. Those are real risks. The paper adds another layer. Provider dependence is also an information risk. A firm that never compares providers, never benchmarks alternatives, and never maintains a modular architecture cannot acquire the counterfactual information needed to know what it is giving up.
That is a different way to understand platform lock-in. Lock-in is not merely that switching is costly. Lock-in is that the firm may not even know the relevant substitution gradient.
Policy implications
The policy implications follow directly from the information-acquisition framing.
The first implication is that AI allocation systems should not be certified only on predictive accuracy. A system can predict accurately inside the historical support of the data and still be unreliable for counterfactual reallocation. Auditors should ask whether the training mechanism generated support for the decisions the model is now being asked to make. If a model is being used to recommend policy changes, budget reallocations, procurement shifts, or eligibility redesign, the relevant diagnostic is not only error on historical cases. It is counterfactual-support coverage.
That means a serious audit should ask: where in the allocation space did the system actually observe variation? Were there local experiments? Were bids elicited? Were alternatives benchmarked? Were shadow prices estimated from actual trade-offs, or were they inferred from realised outcomes under a fixed policy? Which dimensions have support, and which are extrapolation?
The second implication is that administrative AI systems need designed variation. Administrative records are valuable, but they are often path-dependent. If the agency only records outcomes under one policy, the record may be weak for evaluating another. The repair is not necessarily full marketisation. It may be randomised pilots, local perturbations, phase-ins, eligibility thresholds, shadow pricing exercises, controlled procurement comparisons, or structured appeal processes that reveal marginal valuations. The policy question is not whether administration or markets are morally superior. The question is whether the mechanism generates the information needed for the decision.
The third implication concerns procurement. Public and private buyers of foundation-model services should treat multi-provider exposure as an information-acquisition tool. Multi-provider sourcing is not only redundancy. It is a mechanism for discovering performance gradients, substitution possibilities, and hidden costs. A buyer that uses only one provider may know that provider well but know little about the counterfactual. A buyer that maintains credible alternatives acquires information through comparison.
The fourth implication concerns platform regulation and competition policy. Digital-market concentration is often analysed through prices, market shares, entry barriers, and control over users. The information-acquisition view adds another dimension: concentration can reduce the generation of counterfactual signals. If downstream firms all adapt to one dominant provider, the market may lose information about alternative technical paths. That loss may not appear immediately in prices. It appears in reduced exploration, weaker substitution signals, and greater fragility when the dominant provider changes terms.
The fifth implication concerns AI safety and model governance. Many governance frameworks focus on outputs: hallucination, bias, harmful content, fairness metrics, or benchmark performance. Those are necessary concerns. But for allocation systems, governance also needs input-side mechanism analysis. What generated the data? Which incentives shaped it? Which counterfactuals were never observed? Where does the model extrapolate from administrative artefacts rather than economic information?
Why this matters now
This argument matters now because foundation-model systems are moving from content generation into decision infrastructure. They are no longer used only to write text or summarise documents. They are being embedded in enterprise workflows, public administration, product design, procurement, pricing, education, compliance, legal services, software engineering, medical triage, finance, and logistics.
The danger is that statistical confidence will be mistaken for economic knowledge.
A model may be trained on vast amounts of data, but the question is not only how much data exists. The question is what kind of information the data contain. A billion observations generated under one administrative rule may still be weak for evaluating a different rule. A platform’s internal data may be rich inside the platform and poor outside it. A provider’s logs may reveal usage patterns while concealing the opportunity cost of alternative providers. A firm’s historical path may contain evidence of survival under one architecture but not the welfare value of another.
Economic decision-making often turns on the unseen margin. That is exactly where mechanism-generated data can fail.
The paper’s central message is therefore conservative in the technical sense. It says: do not assume that models can infer information the mechanism did not acquire. Do not treat administrative records as if they were local market signals. Do not treat realised outcomes as if they identify counterfactual gradients. Do not treat scale as a substitute for support. Do not treat prediction accuracy as allocation reliability.
The implication is not anti-AI. It is pro-mechanism design.
Foundation models can be valuable allocation tools if they are embedded in systems that generate the right information. They can help process signals, detect patterns, estimate local effects, and support decision-making. But they cannot replace the institutional process that acquires counterfactual information. If the information is not generated, it cannot be learned.
What should be done
The practical answer is to design information acquisition into AI allocation systems.
For administrative systems, this means building controlled variation into policy. Pilot programmes, randomised margins, phased rollouts, threshold designs, and structured appeals can all generate counterfactual information. The goal is not experimentation for its own sake. The goal is to acquire marginal-value information where the model would otherwise extrapolate.
For procurement systems, this means maintaining credible comparison. Multi-provider trials, benchmark contracts, modular integration, and periodic re-tendering generate information. They also discipline provider dependence. The value is not only lower prices. It is better knowledge of the substitution surface.
For firms using foundation models, this means avoiding architectures that make counterfactual learning impossible. A firm that integrates so deeply with one provider that it cannot benchmark alternatives has sacrificed information. The firm may still be efficient in the short term, but it becomes blind to off-path possibilities. In fast-moving AI markets, that blindness is costly.
For regulators, this means asking how market structure affects information production. A concentrated market may reduce the diversity of observed technical and commercial paths. It may make downstream firms more homogeneous. It may weaken the empirical basis for evaluating alternatives. Competition policy should therefore consider not only price and output, but also the production of counterfactual information.
For researchers, this means distinguishing prediction problems from allocation problems. A model that performs well on held-out historical data has not necessarily solved the allocation problem. Researchers need to ask whether the training data contain support for the counterfactuals being evaluated. If not, the correct conclusion is not “the model needs more data.” The correct conclusion may be “the mechanism needs to generate different data.”
The broader point
The broader point is that information economics remains central in the age of foundation models. AI does not abolish the old problem of dispersed knowledge. It changes the machinery through which that problem appears.
Hayek’s price system, Hurwicz’s mechanism design, Blackwell’s information comparisons, and modern machine learning may look like different intellectual traditions. In this setting, they meet. A foundation model is a statistical estimator sitting on top of an economic information structure. Its limits are partly statistical, partly computational, and partly institutional. The institutional part is the one most often missed.
The paper therefore does not ask whether AI can optimise the economy. It asks a prior question: what information did the economy’s mechanisms generate for the AI to learn from?
That question is harder and more useful.
If a market mechanism generates local bid information, then a model may learn marginal valuations around the allocation. If an administrative mechanism records only realised outcomes, then a model may learn the historical policy path without learning the off-path gradient. If a platform concentrates downstream activity into one provider ecosystem, firms may lose cross-provider counterfactual information. If recursive training projects rich distributions into misspecified classes, structural features can be lost even when some aggregate diagnostics remain stable.
These are not failures of intelligence. They are failures of information acquisition.
Conclusion
The information AI cannot learn is the information the mechanism never generated.
That is the central lesson. Foundation models are powerful, but they are not exempt from information economics. They inherit the support, incentives, omissions, and counterfactual gaps of the systems that produced their training data. Treating data as neutral obscures that fact. Treating models as substitutes for mechanisms makes the mistake worse.
The practical implication is clear. Before using AI for allocation, ask what mechanism generated the training data. Ask what counterfactual variation exists. Ask whether marginal values are identified or merely imputed. Ask whether the model is being used inside support or outside it. Ask whether the system is learning from market signals, administrative records, experiments, provider comparisons, or realised-outcome artefacts.
The future of AI allocation will not be determined only by larger models. It will also be determined by better mechanisms for acquiring the information those models need.
References
Acemoglu, D., and Restrepo, P. (2019). Automation and new tasks: How technology displaces and reinstates labour. Journal of Economic Perspectives.
Acemoglu, D., and Restrepo, P. (2020). Robots and jobs: Evidence from US labour markets. Journal of Political Economy.
Afriat, S. N. (1967). The construction of utility functions from expenditure data. International Economic Review.
Agrawal, A., Gans, J., and Goldfarb, A. (2018). Prediction Machines. Harvard Business Review Press.
Agrawal, A., Gans, J., and Goldfarb, A. (2019). Exploring the impact of artificial intelligence: Prediction versus judgment. Information Economics and Policy.
Arrow, K. J. (1962). Economic welfare and the allocation of resources for invention.
Babina, T., Fedyk, A., He, A., and Hodson, J. (2024). Artificial intelligence, firm growth, and product innovation. Journal of Financial Economics.
Blackwell, D. (1951). Comparison of experiments.
Blackwell, D. (1953). Equivalent comparisons of experiments. Annals of Mathematical Statistics.
Bommasani, R., et al. (2021). On the opportunities and risks of foundation models.
Clarke, E. H. (1971). Multipart pricing of public goods. Public Choice.
Cox, D. R. (1972). Regression models and life-tables. Journal of the Royal Statistical Society, Series B.
Grossman, S. J., and Stiglitz, J. E. (1980). On the impossibility of informationally efficient markets. American Economic Review.
Groves, T. (1973). Incentives in teams. Econometrica.
Hayek, F. A. (1945). The use of knowledge in society. American Economic Review.
Hurwicz, L. (1972). On informationally decentralised systems.
Jordan, J. S. (1982). The competitive allocation process is informationally efficient. Journal of Economic Theory.
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature.
Stiglitz, J. E. (2000). The contributions of the economics of information to twentieth century economics. Quarterly Journal of Economics.
Vickrey, W. (1961). Counterspeculation, auctions, and competitive sealed tenders. Journal of Finance.