The Machine That Read Everything
The copyist’s problem returns at industrial scale — when a model trains on the whole corpus, is it the parasite the abolitionist warned about, or the reader the law has always let learn freely?
The copyist’s problem returns at industrial scale — when a model trains on the whole corpus, is it the parasite the abolitionist warned about, or the reader the law has always let learn freely? The honest answer is: it depends, and the slogans on both sides are lying to you
Keywords: intellectual property, copyright, AI training, generative models, fair use, text and data mining, TDM, appropriability, public goods, fixed costs, marginal cost, idea–expression dichotomy, market harm, Bartz v Anthropic, Kadrey v Meta, Thomson Reuters v Ross, EU Copyright Directive, opt-out, licensing.
Abstract
The argument that abolishing intellectual property would licence “appropriation by scale” — the creator bearing the fixed cost of producing a work while a better-capitalised copier reproduces it at near-zero marginal cost — has a new and literal defendant: the generative model trained on the whole corpus. The structure is the same as the single copyist, only aggregated. Thousands of authors, journalists, researchers, illustrators, and coders each bear a fixed cost to produce one work; a training run ingests them all and yields a system that can reproduce the capability of the corpus — summaries, style, substitutes — at marginal cost approaching zero. This is, on its face, the appropriability problem of Arrow and of Landes and Posner, scaled to the level of an entire creative economy and concentrated in the hands of whoever owns the compute and the distribution. But honesty forbids the polemic that the abolitionist’s critics would most enjoy, because the case against AI training is genuinely harder than the case against the photocopier, and for two reasons the slogans on the creator side ignore. First, a trained model is not a verbatim copy: it ingests expression but, when working as designed, outputs new strings, and the idea/expression dichotomy together with Feist‘s rule that facts and learning are free cut both ways — what a model extracts is often exactly the unprotected layer. Second, copyright protects expression, not the act of reading or learning, and whether ingestion is “copying” or “use” is the live, unsettled legal question, not a settled wrong — a question the 2025 American decisions answered in conflicting ways (training held transformative in Bartz v. Anthropic and Kadrey v. Meta; market-substitution fatal to fair use in Thomson Reuters v. Ross; a $1.5 billion settlement turning not on training but on how the books were obtained), and which the EU has answered with a text-and-data-mining exception subject to rightsholder opt-out. The disciplined position, carried over from the companion essays, is the bounded-form test: training that reproduces a work’s protected expression, or produces a market substitute for it, is appropriation of the bounded form and is the copyist’s problem at scale; training that takes only the unprotected layer — facts, ideas, style, the statistical shape of language — and outputs new expression is the reader learning, which the law has always permitted. The hard cases sit between: ingestion makes intermediate copies even when the output infringes nothing, and a model that copies no single work may still erode the market for an entire class of works — appropriability harm without verbatim copying. The “you still have your file” defence fails here exactly as it failed for the photocopier; but so does “all training is theft,” and so does “all training is fair use.” The remedy is neither abolishing intellectual property nor banning the models — both are slogans — but pricing the appropriation of the form while keeping the learning free: licensing, opt-out and text-and-data-mining regimes, collective bargaining, and carve-outs. The principle is the one the companion essays defended. Only the machine is new.
I. The oldest problem, the newest defendant
The companion essays made a structural argument. Creative and inventive goods carry a high fixed cost of creation and a near-zero marginal cost of reproduction; in the absence of any right to exclude, the creator bears the cost of resolving the uncertainty — writing the book, discovering the invention, compiling the data — while a better-capitalised copier waits, reproduces the finished product at marginal cost, undersells the originator who alone paid the fixed cost, and captures the return. That is the appropriability problem, identified by Kenneth Arrow in 1962 and translated into the law of copyright by Landes and Posner in 1989, and the companion essays argued that abolishing intellectual property does not free the market but hands it to whoever owns the means of reproduction — appropriation by scale.
That argument was written as if its defendant were a publisher with a printing press. It now has a defendant that is far larger and far stranger: the generative model trained on the corpus of nearly everything that has been written, drawn, photographed, or coded. And the structure is identical, only aggregated. Where the single copyist took one finished book, the training run takes the finished output of a whole creative economy. Where the copyist reproduced one work at marginal cost, the model reproduces the capability of the corpus — the ability to summarise it, to write in its styles, to produce substitutes for the kinds of thing it contains — at a marginal cost that, once the model exists, approaches zero. The fixed cost is borne, as before, by the creators: thousands of them, each paying in time and risk to produce one work before any return. The capability is captured, as before, by the party with the complementary assets — here the compute, the data pipeline, and the distribution. The copyist’s problem did not go away. It went to industrial scale and acquired a new owner.
Figure 1. The copyist’s problem at industrial scale. The asymmetry is the same as the single copyist’s; the two disanalogies in amber are why the case is harder than the photocopier’s, and why the polemic must not skip them.
If the analysis stopped there, the conclusion would write itself: the model is the parasite the abolitionist warned about, only enormous, and the creators are the producers being stripped at the point of value. That conclusion is half-true, and a polemic that stopped there would be lying by omission — because the case against the machine is genuinely harder than the case against the press, and the next two sections are about why.
II. Why this is harder than the photocopier
The photocopier makes a copy. That is the whole of what it does, and it is why the copyist case is easy: the output is the input, reproduced. The trained model does something the photocopier does not, and the difference is not a technicality — it goes to the core of what copyright protects.
The first disanalogy is that a trained model is not a verbatim copy, and when it works as designed it does not output one. It ingests an enormous quantity of expression and, from it, learns a statistical structure — which words tend to follow which, how arguments are shaped, what the styles and registers of human writing are — and then it generates new strings that did not appear in any input. Copyright has a precise and ancient way of describing what is happening here, and it cuts against the creators as often as for them. The idea/expression dichotomy protects the specific expression and withholds protection from the idea, the method, the system, and the fact; and Feist holds that facts and the products of effort are free, that copyright rewards originality of expression and not labour or learning. What a model extracts from a text is, very often, exactly the unprotected layer — the facts it states, the ideas it advances, the statistical shape of the language it is written in — none of which copyright has ever given anyone the right to control. A human who reads a thousand novels and learns how to write one infringes nothing, because what he took was the unprotectable lesson, not the protectable form. The model’s defenders say it is doing the same thing at scale, and on the cases where the output is genuinely new expression, they are not obviously wrong. The dichotomy that the companion essays used against the abolitionist — IP protects the bounded form, not the idea — here cuts partly for the machine, because learning the idea is exactly what the law leaves free.
The second disanalogy is that copyright protects expression, not the act of reading or learning, and whether machine ingestion counts as “copying” or as “use” is a genuinely contested legal question rather than a settled wrong. When a person reads a book to learn from it, no copyright event occurs, because reading is not one of the exclusive rights. When a machine “reads” a book, it makes copies in the technical sense — it loads, tokenises, and processes the text, producing intermediate reproductions — and whether those copies are the actionable act, or whether they are an unactionable incident of a non-expressive use, is precisely the question the courts and legislatures are now fighting over. This is not a question the creator side gets to assume away by calling ingestion “theft,” and it is not a question the developer side gets to assume away by calling it “just reading.” It is open, and the honest essay says so.
These two disanalogies are why “you still have your file” — the slogan the companion essays demolished when a publisher used it to excuse reproducing a book — does not, by itself, resolve the AI case. It still fails as a defence: the fact that the author retains the manuscript is as irrelevant to a model that has appropriated the bounded form as it was to the publisher who reprinted it. But it no longer decides the case, because the prior question — did the model appropriate the bounded form at all, or only the unprotected layer the law leaves free? — is now live, and the answer is not the same for every model or every use.
III. The bounded-form test, applied to the machine
The companion essays’ organising idea was the bounded form: the wrong of copying is the appropriation of the determinate expressive or inventive form a producer made, not the taking of its value and not the failure to compensate effort. That test does real work here, because it sorts the AI cases that the slogans mash together.
Apply it as a question about the output, not a slogan about the technology. Does the model reproduce the work’s protected expression — verbatim or near-verbatim — or produce a market substitute for the specific work? If so, that is appropriation of the bounded form. The model is doing what the copyist did: displacing the work’s market, capturing the value of its form, occupying the demand the author created. The fact that it arrived at the reproduction through a training process rather than a photocopier changes nothing, because the bounded-form test looks at what was taken, not at the mechanism of taking. This is the copyist’s problem, at scale, and the appropriability argument applies with full force.
Or does the model take only the unprotected layer — the facts, the ideas, the style, the statistical shape of language — and output genuinely new expression? If so, that is, in principle, the reader learning. What was taken is what the law has always left free; the idea/expression dichotomy and Feist are not obstacles to this use but descriptions of why it is permitted. A model that has read widely and writes newly is doing what every author who has ever read widely and written newly has done, and the bounded-form test acquits it for the same reason it acquits the well-read novelist: the learning, not the form, is what was taken, and reading is not copying.
Figure 2. The bounded-form test applied to training. The two clear branches are easy; the value of the test is that it identifies the genuinely hard cases instead of pretending they are easy.
The point of the test is not that it makes every case easy. It is that it makes the easy cases easy and isolates the hard ones honestly, instead of the two slogans — “all training is theft” and “all training is fair use” — each of which simply asserts that one branch swallows the other. It does not. Some training reproduces the form; some takes only the lesson; and the interesting law, and the genuine moral difficulty, live in the cases that sit between.
IV. The hard case in the middle, named honestly
Two features of model training make the middle of that tree genuinely hard, and the disciplined position has to name them rather than wish them away.
The first is that ingestion makes intermediate copies even when the output infringes nothing. A model that never reproduces a single protected sentence still, in the course of training, loaded and processed the protected works — and whether that intermediate copying is itself the actionable act, or an unactionable incident of a transformative, non-expressive purpose, is the precise question on which the American courts split in 2025 and on which the EU and UK have legislated. In Bartz v. Anthropic, Judge Alsup held that using lawfully acquired books to train an LLM was “exceedingly” transformative fair use, on the reasoning that authors cannot exclude others from using works to learn and that the training produced something new rather than a substitute — but he treated each step separately and held that downloading and retaining pirated copies to build a permanent library was a distinct, non-transformative act, and it was that — the acquisition, not the training — that drove a settlement reported at roughly $1.5 billion. Two days later in Kadrey v. Meta, Judge Chhabria reached the same fair-use result on the record before him but by a different route, treating downloading and training as one integrated process and stressing that the plaintiffs had failed to prove market harm — while pointedly warning that stronger evidence of market substitution could change the outcome in a future case. The lesson is not that training is settled law; it is that two judges in the same courthouse in the same week agreed on the result and disagreed on almost everything about how to get there.
The second hard feature is the one the creator side states best and the developer side most wants to avoid: a model that copies no single work may still erode the market for an entire class of works. Even if no output reproduces any particular author’s expression, a system that can produce an endless supply of competent substitutes for the kind of thing a class of authors produces may compete away the return to producing it — appropriability harm without verbatim copying. This is exactly the level-versus-margin point the companion essays insisted on: the question is not whether any one work was copied but whether the incentive to produce the class of works survives. The contrast in the 2025 cases is instructive. Where the use was a direct market substitute, fair use failed: in Thomson Reuters v. Ross, a company that took Westlaw’s copyrightable headnotes to build a competing legal-research tool was held to have infringed, the court resting heavily on the fourth fair-use factor — the effect on the work’s potential market — because the product was a substitute for the original. Where the plaintiffs could not show that substitution, fair use succeeded, but the courts signalled that the market-harm question was the one that would decide the next round. The bounded-form test and the appropriability argument meet exactly here: market substitution is the signature of appropriation, and it is the thing the law is converging on as the real question, whether or not any individual work was reproduced.
So the “you still have it” defence fails in the middle of the tree too — the author’s retained file no more answers the erosion of his market than it answered the publisher’s reprint. But it fails on both sides of the slogan war. “All training is theft” is false, because much training takes only the unprotected layer and outputs new expression, which the law has always permitted. “All training is fair use” is false, because training that reproduces the form or floods the market with substitutes is the copyist’s appropriation at scale. The case turns on design and testing — what the model ingests, what it outputs, whether it substitutes — and not on a slogan in either direction.
V. Where the value moves, and what to do about it
Strip the technology away and look at the distribution, because that is where the companion essays’ first argument — the distributive face — reappears at platform scale.
On one side stand the bearers of the fixed cost: the authors, journalists, researchers, illustrators, coders, and photographers, in their thousands, each having paid in time and risk to create one work before any return arrived. On the other stands the party that captures the return: the firm with the compute, the data pipeline, the trained model, and the distribution, which reproduces the corpus’s capability at marginal cost and serves it at scale. The value moves from the first to the second, and it moves for the same reason it moved from the author to the publisher in the companion essay’s production-house scenario — because the party that bore the fixed cost has no way to capture the return once the capability has been extracted and can be served at marginal cost. This is appropriation by scale with a new and much larger defendant.
Figure 3. Where the value moves when the corpus trains the model. The distributive face of the companion essay, at platform scale.
But naming the harm is not the same as prescribing the cure, and the two cures the slogans offer are both wrong. “Abolish intellectual property” — the abolitionist’s answer — would not free the creators; it would remove the only lever they have and complete the transfer the companion essays described, leaving the model’s owner in undisputed possession of a capability built entirely from work it never had to pay for. “Ban the models” — the maximalist’s answer — would forgo the genuine public good of systems that can do what learning from the corpus enables, and would punish the legitimate branch of the tree (the reader that learns and writes newly) to reach the illegitimate one. Both answers fail because both ignore the bounded-form test: they treat training as all-appropriation or all-learning, when it is some of each.
The cure that follows from the test is to price the appropriation of the form while keeping the learning free, and the institutional materials for it already exist and are being built. The European Union’s approach is the clearest template: its text-and-data-mining regime permits reproductions for mining and model-building but subjects the general, commercial exception to a rightsholder opt-out — the right to reserve a work against training by machine-readable means — and the AI Act ties general-purpose model providers to respecting those reservations. The United Kingdom, in its December 2024 consultation, proposed an EU-style opt-out exception underpinned by transparency about training sources, with collective licensing expected to fill the gap. The American cases point the same way from the litigation side: train on lawfully acquired works and take the unprotected layer, and you are likely safe; build your corpus from piracy, or produce market substitutes, and you are not — a roadmap that prices the appropriation (through liability and settlement) while leaving the learning (through the transformative-use holding) free. None of these regimes is finished, and each has real defects — the opt-out puts the burden on the smallest and least-resourced rightsholders, transparency obligations are thin, and the licensing markets are immature. But they are all instances of the same correct move: not the abolition of intellectual property, and not the prohibition of the technology, but the construction of institutions that distinguish the appropriation of the form from the learning of the lesson, and charge for the first while permitting the second.
VI. Conclusion: the principle is old, the machine is new
The model that read everything is not a refutation of the companion essays’ argument; it is its largest confirmation. The appropriability problem they described — fixed cost borne by the creator, near-zero marginal cost captured by the better-capitalised reproducer — is exactly the problem that training on the corpus presents, scaled from one book to a creative economy and concentrated in the hands of whoever owns the compute. The “you still have your file” defence fails against the machine for the same reason it failed against the press: the author’s retained copy was never the point, and the appropriation of the form and its market is.
But the machine is genuinely harder than the press, and the honesty the companion essays demanded of the abolitionist is owed here in the other direction too. A trained model is not a photocopier; it ingests expression and, when it works as designed, outputs new expression, taking the unprotected layer — facts, ideas, style — that the law has always left free. Copyright protects the form, not the act of reading, and whether machine ingestion is copying or use is the live question the courts split on in 2025 and the legislatures are now answering. The slogans on both sides are lying: “all training is theft” ignores the legitimate branch where only the lesson is taken, and “all training is fair use” ignores the illegitimate branch where the form is reproduced and the market is flooded with substitutes.
The bounded-form test cuts the knot the slogans cannot. Training that reproduces the protected form, or produces a market substitute, is the copyist’s appropriation at scale, and the appropriability argument applies with full force. Training that takes only the unprotected layer and outputs new expression is the reader learning, and the law has always permitted it. The hard cases — intermediate copies that infringe no output, models that copy no single work but erode the market for a class of works — are real, and the disciplined position names them instead of pretending either branch swallows the other. And the remedy is the one the companion essays implied: not the abolition of intellectual property and not the banning of the models, but the pricing of the appropriation of the form while the learning is kept free — through licensing, opt-out and text-and-data-mining regimes, collective bargaining, and carve-outs.
The creator who writes and the firm that trains on the writing without paying for it is the author and the production house again, at a scale the production house could never have reached. The issue is not that the creator “still has his file.” The issue is that the capability his work made possible has been appropriated and served back to the market by the party that bore none of his cost. That is the oldest problem in the economics of creation. The machine that read everything is only its newest, and largest, instance — and the answer is the old one: charge for the form, free the lesson, and build the institutions that can tell them apart.
References
Economic foundations-
Arrow, Kenneth J. “Economic Welfare and the Allocation of Resources for Invention.” In The Rate and Direction of Inventive Activity, 609–626. Princeton: Princeton University Press (NBER), 1962.
-
Landes, William M., and Richard A. Posner. “An Economic Analysis of Copyright Law.” Journal of Legal Studies 18, no. 2 (1989): 325–363.
United States case law (2025)-
Bartz v. Anthropic PBC, No. 3:23-cv-04768 (N.D. Cal. June 23, 2025) (training on lawfully acquired books held “exceedingly” transformative fair use; pirated-library acquisition a distinct, non-transformative act; subsequent class certification and a settlement reported at ~$1.5 billion).
-
Kadrey v. Meta Platforms, Inc., No. 3:23-cv-03417 (N.D. Cal. June 25, 2025) (training held fair use on the record; downloading and training treated as one integrated process; plaintiffs failed to prove market harm; court warned stronger market-harm evidence could change the result).
-
Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 1:20-cv-00613 (D. Del. Feb. 11, 2025) (Westlaw headnotes protectable; use to build a competing tool not fair use; decided on the fourth factor, market substitution; on appeal to the Third Circuit).
Statutory / legislative anchors-
17 U.S.C. § 107 (fair use; the four factors, including effect on the potential market).
-
17 U.S.C. § 102(b) (idea/expression; exclusion of ideas, procedures, methods, facts).
-
Directive (EU) 2019/790 (Copyright in the Digital Single Market), Arts. 3–4 (text-and-data-mining exceptions; Art. 4(3) rightsholder opt-out by machine-readable reservation).
-
Regulation (EU) 2024/1689 (AI Act), Art. 53(1)(c) and Recital 106 (general-purpose model providers must respect Art. 4(3) reservations).
-
UK Intellectual Property Office, Copyright and Artificial Intelligence consultation (17 December 2024) (proposing an EU-style TDM exception with opt-out and transparency).
Doctrinal anchors (developed in the companion essays)-
Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340 (1991) (originality, not effort; facts free).
-
Harper & Row, Publishers, Inc. v. Nation Enterprises, 471 U.S. 539 (1985) (the bounded form and its exclusivity; first-publication right).
Note on method and scope. All case holdings, settlement figures, and statutory provisions above are drawn from current reporting and the instruments themselves and were verified at the page level before writing; none rests on an abstract. The 2025 American decisions are district-court rulings — one settled, one in continuing litigation, one on appeal — and are described as non-binding and unsettled, which they are; nothing here states that the law of AI training is fixed. Two propositions are flagged as contested rather than settled: whether intermediate copying in training is itself actionable, and whether non-substitutive training nonetheless inflicts cognisable market harm on a class of works. The essay defends the bounded-form principle and the appropriability analysis; it does not endorse any particular litigant, statute, or licensing scheme, and it states plainly that the opt-out and TDM regimes it points to have real and unresolved defects.