The Slow Decline of Academia: When Methodology Became a Dropdown Menu
How mass higher education, managerial templates, publication metrics and AI can turn scholarship into procedural compliance
Keywords: academia, higher education, methodology, research methods, managerialism, audit culture, publication pressure, artificial intelligence, AI agents, doctoral education, research integrity, reproducibility
Abstract
Universities remain among the most useful institutions ever created. They preserve specialised knowledge, train researchers, maintain archives and laboratories, create communities around difficult problems, and still offer forms of intellectual apprenticeship that are hard to reproduce elsewhere. That is precisely why institutional drift inside academia matters. By “decline” I do not mean that universities are uniformly worse than they were, nor do I claim a single time-series measure showing that scholarship everywhere has deteriorated. The claim is narrower and more defensible: in parts of higher education, the combination of expansion, managerial governance, metric-driven evaluation and scalable teaching has increased the reward for procedural conformity at the very point where research requires judgement. Methodology can then be reduced from an argument about how a claim is warranted to a template specifying which approved technique, package, structure and vocabulary should appear. I have completed six doctorates and am working on a seventh and eighth, alongside more master’s degrees than is probably evidence of good judgement. That experience supplies motivating cases, not population estimates. I have seen a more rigorous approach rejected because it sat outside a departmental template, and I have been criticised for using Python rather than SPSS when the actual analysis was substantially more sophisticated. The important issue is not that Python is intrinsically superior to SPSS—it is not—but that software familiarity can be mistaken for methodological quality. This essay argues that AI will not create that error so much as make it harder to hide. As generative systems and computer-use agents reduce the cost of writing code, operating software, generating tables, searching literature and drafting prose, universities will face an awkward question: if the assessed skill was merely following the procedure, what exactly was the human being learning?
Academia is not dead. That would be simpler.
Decline is an awkward word because it invites melodrama. Universities have not collapsed. Libraries still contain books. Laboratories still contain people who know what they are doing. Good supervisors still exist. Brilliant students still appear, often despite the systems designed to process them. Important research is produced every week. I continue to enrol in universities because I continue to get something from them, which would be an odd habit if I thought the entire institution worthless.
The problem is subtler. Academia has become exceptionally good at expanding the machinery of scholarship while sometimes weakening the distinction between scholarship and its machinery.
That distinction matters. A research method is not the same thing as a research-methods form. Statistical reasoning is not the same thing as operating statistical software. A literature review is not the same thing as filling a table with citations. Peer review is not the same thing as acquiring two reports and ticking a workflow box. Doctoral training is not the same thing as producing a thesis that resembles the theses produced last year. Yet large institutions naturally favour procedures that can be standardised, audited, taught to cohorts, defended to regulators and verified by administrators. What is administratively legible gradually becomes confused with what is intellectually sound.
I have watched that process for a long time. I have completed six doctorates and am now working on a seventh and eighth. I have also accumulated more master’s degrees than any sensible person needs. This is not an appeal to authority; credentials do not make an argument correct. It does, however, mean I have spent an absurd amount of time inside different academic systems, across different disciplines, institutions and periods. The contrast is difficult to miss.
When I first began postgraduate work, methodology was more often treated as something one had to defend. The point was not that an approved package had produced an approved output. The point was that the researcher understood why a method was appropriate, what assumptions made the inference possible, what alternative specifications existed, what the data could not establish, and what would falsify or weaken the claim. There was bureaucracy, certainly. Universities have never suffered from a shortage of paper. But the paper was not yet so often mistaken for the thought.
At one university, I left shortly before another degree was to be awarded. One reason was methodological. I had proposed a more rigorous approach than the prescribed model, but it sat outside the departmental template. That made it suspect. At another institution, a supervisor told me off for not using SPSS. I had used Python instead and conducted what I could defend, on methodological grounds, as a substantially more rigorous analysis. The objection I received was not that the model was wrong, that the code was faulty, that the estimand was incoherent, that the assumptions were violated or that the inference failed. It was a tool-familiarity objection. I cannot infer another person’s private understanding from that episode, but I can identify the epistemic failure in the objection: the software brand was being used as a proxy for methodological adequacy.
There is something magnificently comic about treating an analysis as suspect because it was not performed in software the examiner recognises. It is rather like condemning a proof because the mathematician used a fountain pen instead of the departmental pencil.
The comedy stops when such behaviour becomes institutional practice.
Figure 1. Inquiry-driven research and template-driven compliance. The diagram is deliberately stylised: real inquiry is iterative and often loops back from analysis or interpretation to design and data. The lower pathway represents an institutional failure mode, not a claim about all departments or all research.
What I am—and am not—calling decline
Before blaming an entire sector, it helps to specify the dependent variable. “Academia” is too large to decline in one dimension at one rate. Universities can improve access while worsening supervision ratios. They can produce better computational tools while teaching weaker conceptual understanding. A field can become more transparent while simultaneously becoming more metric-driven. A department can be excellent while the institution around it becomes increasingly procedural. A period that is better for one student can be worse for another.
So this is not an argument for a golden age. There was no golden age. Older universities could be exclusionary, arbitrary, parochial and spectacularly confident in bad ideas. Small elite systems often relied on opaque judgement and social gatekeeping. More access, clearer standards, better statistical software, stronger ethics procedures and open-science reforms are genuine improvements. Schofer and Meyer (2005) document the extraordinary worldwide expansion of higher education during the twentieth century; Trow (2006) explains why movement from elite to mass and then universal access changes institutional form. Neither result implies decay. Expansion is an achievement.
My claim is about a particular failure mode that becomes easier to produce at scale: the substitution of legible procedure for difficult judgement. The evidence for that failure mode comes from several literatures rather than from one grand “decline index”: research on audit culture and rankings, studies of metric reactivity, publication pressure and hypercompetition, work on reproducibility and incentive selection, and the emerging literature on generative and agentic AI. Those literatures support mechanisms. They do not prove that every university has moved monotonically in the same direction.
That distinction matters because otherwise the argument commits the very methodological sin it is criticising: taking a vivid observation, giving it an impressive label, and pretending that the label did the identification.
Massification was a success, but success has consequences
Any serious criticism of modern higher education has to begin by conceding what nostalgic accounts usually omit. The expansion of university participation was, in many respects, an extraordinary achievement. Higher education moved from an elite system serving a small segment of the population toward mass and, in some countries, near-universal participation. Martin Trow’s classic analysis described the transition from elite to mass to universal higher education as a structural transformation, not merely an increase in student numbers (Trow, 2006). As participation expands, the purposes, governance, curriculum, assessment and relationship between institutions and society change with it.
There are obvious benefits. More people can study medicine, engineering, economics, history, computing, law and the sciences. Talent that would once have been excluded by class, geography, gender or institutional gatekeeping can enter academic life. Economies gain a larger pool of specialised labour. Mature students can retrain. Researchers can move between disciplines. Knowledge that once belonged to narrow professional circles becomes more widely accessible.
It would be foolish to argue that universities were better simply because fewer people were allowed into them. Scarcity is not a synonym for excellence.
Scale creates a governance problem, but it does not mechanically cause intellectual decline. An institution educating hundreds can rely more heavily on apprenticeship, tacit judgement and personal knowledge. An institution educating tens of thousands needs more explicit coordination: rubrics, modules, learning outcomes, progression rules, ethics forms, quality-assurance procedures, submission systems and moderation protocols. That is an organisational response to scale, not evidence that mass participation itself is defective.
Some of this standardisation is necessary and beneficial. If ten supervisors assess ten doctoral students according to ten entirely private standards, arbitrary treatment becomes likely. Standards can protect students as well as constrain them. Replicable procedures can reduce prejudice. Explicit criteria can improve transparency. The enemy is not standardisation itself, and massification is neither necessary nor sufficient for the failure I am describing.
The problem appears when governance mechanisms designed to establish a procedural floor become an intellectual ceiling—especially when they interact with metrics, risk aversion, completion targets and evaluator familiarity.
The expansion of higher education also changed universities economically. Slaughter and Rhoades (2004) described the emergence of “academic capitalism”, in which universities increasingly engage markets, revenue generation, competitive positioning and commercially oriented networks. One need not accept every element of that framework to recognise the broader shift. Universities now compete for students, rankings, grants, citations, international visibility, industry links and measurable “impact”. These pressures produce administrative rationality. What can be counted becomes easier to manage; what can be compared becomes easier to rank; what can be templated becomes easier to scale.
The institution therefore develops an understandable preference for outputs that are legible from a distance.
A paper count is legible. A citation score is legible. A completion rate is legible. A standardised methods section is legible. Whether a student genuinely understands identification in a causal model is less legible. Whether a historian has spent five years learning to distinguish a genuine archival pattern from an artefact of record survival is less legible. Whether an econometric specification is conceptually appropriate rather than merely executable is less legible. Whether a mathematical argument contains an original idea rather than fifty pages of polished derivation is less legible.
Large organisations do what large organisations do. They measure the visible thing and then, over time, optimise around it.
The audit culture and the genius of measuring the proxy
Campbell’s law is often paraphrased more aggressively than Campbell wrote it, but the underlying insight remains important. When quantitative social indicators become targets for decision-making, they become vulnerable to corruption pressures and can distort the processes they were meant to monitor (Campbell, 1979). The modern university is an almost laboratory-perfect environment for observing this phenomenon.
Shore and Wright (2015) describe the spread of audit culture, rankings and numerical governance across universities and other institutions. Espeland and Sauder (2007), studying law-school rankings, show how public measures do not merely describe organisations; they can change them. Rankings create reactivity. Institutions alter behaviour in response to being measured, and the measure begins to reshape the world it claims merely to represent.
Power’s (1997) account of the “audit society” supplies the broader institutional logic. Verification practices can become rituals with organisational lives of their own: the presence of an auditable process can begin to stand in for confidence in the underlying activity. Universities did not invent this tendency. They have merely become unusually accomplished practitioners of it.
Research evaluation behaves similarly. De Rijcke et al. (2016), reviewing the effects of indicator use, document strategic responses, gaming and behavioural changes associated with research assessment systems. The Leiden Manifesto was written precisely because research metrics had become sufficiently influential that Hicks et al. (2015) thought it necessary to restate what should have been obvious: quantitative evaluation should support expert judgement, not replace it, and indicators need to be interpreted in relation to context, field and mission.
One might think this was unnecessary advice. One would be wrong.
Universities have managed the impressive feat of employing thousands of experts and then constructing systems designed to reduce the amount of expert judgement required. This is not because administrators are stupid. It is because judgement is expensive, contestable and slow. A metric is cheap, defensible and wonderfully indifferent to context. It also produces a spreadsheet, which gives the comforting appearance that somebody is in control.
Once a metric matters, academics adapt. If promotion depends on publication count, salami slicing becomes rational. If journal rank matters, topic selection shifts toward what prestigious journals are likely to publish. If citation counts matter, researchers have incentives to work in large, fashionable literatures. If student satisfaction matters too strongly, academic challenge can become a customer-service problem. If completion time matters, difficult methodological detours become administrative inconveniences. None of these responses requires fraud. People simply respond to the environment they inhabit.
That is the crucial point. Decline does not require villains.
It requires incentives.
Publish or perish eventually becomes publish and perish
Publication pressure is hardly new. Academics have complained about “publish or perish” for generations, usually while publishing another paper about it. Yet the scale and formalisation of output pressure matter. Haven et al. (2019) found negative attitudes toward the publication climate across academic ranks and fields in their survey of Amsterdam researchers. Edwards and Roy (2017) argue that competition for funding, quantitative performance metrics and changes in the university business model have created a climate of hypercompetition and perverse incentives. The point is not that every researcher becomes dishonest. The point is that systems can select for behaviours that increase output without proportionately increasing knowledge.
Smaldino and McElreath (2016) make the argument unusually sharply. In their model of scientific communities, selection for high output can favour poorer research methods even without conscious misconduct. Laboratories using practices that generate more publishable results can reproduce institutionally because publication success carries career advantages. The paper requires an important bibliographic qualification that is easy to miss: an independent replication later identified a coding error in one part of the original implementation, while reporting that the central conclusions were otherwise reproduced (Kohrt et al., 2023). Smaldino and McElreath (2023) then published a formal correction. That history is not an embarrassment to hide; it is almost an advertisement for the process the essay is defending. Models should be inspectable, errors should be discoverable, and claims should survive correction rather than prestige.
This is an unpleasant result because it removes the comforting explanation that bad science is mostly produced by bad people.
A bad system does not need bad people. It merely needs normal people responding sensibly to badly designed incentives.
Ioannidis (2005) famously argued that many published findings can be false under conditions involving low power, flexible analysis, multiple testing, small effects and selective publication. His paper generated substantial debate, including criticism of some of its stronger mathematical claims (Goodman & Greenland, 2007), and a 2022 correction fixed a missing set of parentheses in one equation in Table 2 (Ioannidis, 2022). The productive lesson is not the bumper-sticker version that “most science is wrong”. Science is a process designed to correct error, and many fields have strong cumulative records. The narrower lesson is that publication is not a magical purification ritual. Peer review does not convert a weak design into a strong one. A statistically significant coefficient does not acquire moral character when typeset by a respectable publisher.
The reproducibility debate made this impossible to ignore. In a Nature survey of 1,576 researchers, more than 70 per cent reported having tried and failed to reproduce another scientist’s experiment, while more than half reported failure to reproduce their own work (Baker, 2016). That survey is not a representative census of all science and should not be treated as one. It is nonetheless an extraordinary warning about the distance that can exist between the appearance of settled method and the actual reliability of results.
Open-science reforms attempt to narrow that distance. The Transparency and Openness Promotion guidelines, for example, seek to align journal practices with data transparency, analytic transparency, preregistration and reproducibility (Nosek et al., 2015). Such initiatives are valuable. Yet even reform can become another checklist if adopted without understanding. A preregistration form filled in mechanically is not the same thing as a well-designed study. A repository full of incomprehensible code is not transparency in any useful sense. An “open data” badge does not rescue bad measurement.
Academia has a remarkable ability to respond to one bureaucratic failure by inventing another form.
Methodology is an argument, not a software package
This brings us back to SPSS.
There is nothing wrong with SPSS. It is useful software. So are Stata, R, SAS, MATLAB, Mathematica, Maple, Julia, Python and a long list of specialist packages. I used computational methods decades ago when the machines were slower, the interfaces uglier and the waiting time longer. The existence of software-assisted analysis is not a modern corruption of scholarship. Researchers have always used tools. Nor is Python intrinsically more rigorous than SPSS. Bad analysis can be written beautifully in Python; excellent analysis can be executed in SPSS. Rigour belongs to the design, assumptions, implementation, diagnostics, validation and inferential argument—not to the logo on the splash screen.
The corruption occurs when operation of the tool substitutes for understanding the analysis.
A student can learn to click “regression”, select variables, request robust standard errors and produce a table without understanding what is being estimated. Another can paste a prompt into an AI system and obtain Python code for a fixed-effects model without understanding the identifying variation. A third can run a difference-in-differences package with all the correct syntax while remaining unaware that staggered treatment timing and heterogeneous effects can make a conventional two-way fixed-effects coefficient difficult to interpret causally. That is not a software superstition; the econometric problem is explicit in the modern literature (Callaway & Sant’Anna, 2021; Goodman-Bacon, 2021). A fourth can generate a structural-equation model because the software returned a pleasing collection of fit indices.
All four can produce output that looks academic.
That is precisely why appearance is a poor standard.
Methodology begins before software. What is the research question? What is the object of inference? What is measured and what is merely proxied? What variation identifies the parameter? Which assumptions are required? Which assumptions are testable? What selection process generated the sample? What data-generating process would produce the same pattern under an alternative explanation? What uncertainty is represented by the reported interval, and what uncertainty is omitted? What decisions were made after seeing the data? How sensitive is the conclusion to specification, coding, missingness, prior choice, bandwidth, functional form or measurement error?
A competent researcher should be able to explain these things before clicking anything.
The peculiar danger of template education is that it reverses the sequence. Instead of beginning with the question and deriving an appropriate method, the student begins with a menu of authorised methods and searches for a question that fits one of them. The template is pedagogically convenient because it lets a department teach methods at scale. The student can be told, for example, that quantitative work follows a recognised positivist structure, qualitative work follows an approved interpretive structure, mixed methods requires the correct diagram, and each has a prescribed vocabulary that signals methodological literacy.
Some standard vocabulary is useful. Shared language makes disciplines possible. But when vocabulary becomes a password system, students learn the semiotics of competence rather than competence itself.
They learn that certain words open doors.
They learn to say “triangulation”, “robustness”, “reflexivity”, “saturation”, “validity”, “reliability”, “ontology” and “epistemology” at the correct ceremonial moments. In quantitative work they learn that a p-value below a threshold is pleasing. In qualitative work they learn that thematic saturation is something one eventually announces. In mixed methods they learn that drawing two arrows between boxes can make a project look integrated. Every field has its rituals.
Rituals are not necessarily useless. They become dangerous when nobody remembers what they were for.
Doctoral education and the production of acceptable sameness
The doctorate ought to be where this tendency is weakest. A PhD is supposed to demonstrate the ability to conduct original research. Original research cannot be fully specified in advance because, if it could, it would not be original in any interesting sense.
Yet doctoral education is also where institutional risk aversion is strongest. Universities want predictable supervision, predictable milestones, predictable ethics approval, predictable examination and predictable completion. Students want to finish. Supervisors want to avoid disasters. Graduate schools want procedures that can be applied across departments. Everyone has a rational reason to prefer a project that fits known categories.
The result can be intellectual conservatism disguised as methodological discipline.
A standard approach is “rigorous” because the department knows how to examine it. A novel approach is “unclear” because the department does not. A familiar software package is “appropriate” because supervisors can support it. A custom computational pipeline is “risky” because fewer people can audit it. An established theoretical lens is “well grounded” because the literature review already exists. A genuinely interdisciplinary project is “too broad” because it does not sit comfortably inside the administrative boxes that allocate expertise.
Again, not every institution behaves this way. Strong departments often do the opposite. Good supervisors protect unconventional work. Some graduate programmes require serious technical competence and reward methodological innovation. But the structural pressure runs toward legibility.
This is where the personal examples are useful, provided they are not asked to prove more than anecdotes can prove. Being criticised for using Python instead of SPSS did not establish the prevalence of the problem across academia. It did, however, illustrate the mechanism cleanly: an objection about evaluator familiarity can be recoded as an objection about research quality. The approved tool then becomes a proxy for auditable competence.
That is institutional epistemology at its most convenient: what fits the recognised workflow is easy to certify; what does not fit it acquires an administrative burden of proof.
The more universities expand, the more tempting this becomes. If thousands of dissertations must be supervised, assessed and completed, then templates lower transaction costs. Cookie-cutter work is easier to manage because everybody already knows which drawer it belongs in.
The irony is that universities often advertise “critical thinking” while building systems that punish departures from the anticipated route.
Critical thinking, apparently, is excellent so long as it uses the supplied headings.
The coming AI problem is not plagiarism. It is procedural abundance.
Most university discussion of generative AI began at the shallow end: plagiarism, cheating, detection, authorship and whether students should be allowed to use ChatGPT to write an essay. Those questions matter, but they are not the most important ones.
The deeper issue is that AI sharply reduces the cost of many tasks that once created procedural scarcity.
For decades, one hidden assumption of academic assessment was that producing the artefact required at least some command of the underlying process. If a student submitted a 10,000-word literature review, somebody probably had to read a substantial amount. If a researcher produced a complicated regression table, somebody probably had to know enough software to specify it. If a doctoral candidate supplied code, figures, robustness tests and a formatted bibliography, producing that collection of outputs took time and specialist labour.
Generative AI weakens the inference from output to competence.
Geng and Trotta (2024), examining a large corpus of arXiv abstracts, found evidence consistent with substantial adoption of LLM-assisted academic writing, particularly in computer science. The exact fraction depends on their modelling assumptions and should not be treated as a direct count of AI-authored papers. The broader conclusion is less controversial: generative models are already changing academic production.
But writing assistance is the trivial case.
Agentic systems move from text generation toward workflow execution. Lu et al. (2024) presented an “AI Scientist” framework designed to generate ideas, write code, execute experiments, visualise results, draft papers and conduct an automated review process in machine-learning domains. Schmidgall et al. (2025) developed Agent Laboratory, an LLM-agent framework coordinating literature review, experimentation and report writing, with human feedback improving output quality. Computer-use research makes the broader point even more directly. OSWorld evaluates multimodal agents on open-ended tasks in real operating-system environments, including ordinary desktop applications; its original results also showed that agents remained dramatically below human performance on the benchmark (Xie et al., 2024). The intellectually honest conclusion is therefore neither “agents cannot operate software” nor “agents can autonomously replace researchers”. They can increasingly execute multi-step computer workflows, while remaining brittle enough that verification is indispensable.
The agent no longer has to stop at telling you which button to click. Depending on the system and permissions, it can interact with the interface, write a script, install a package, clean data, run a model, inspect an error, revise code, rerun the model, generate a figure, format a table and draft the paragraph explaining what the coefficient supposedly means. Some of those actions will fail. Some will be wrong. That is precisely why methodological understanding matters more, not less.
There is nothing conceptually protecting SPSS from automation merely because it has menus. A computer-use agent need not possess a mystical understanding of SPSS; it needs sufficient interface control to operate it, or it can call equivalent statistical routines elsewhere. Telling a student merely to “use SPSS” therefore has a rapidly diminishing claim to count as an intellectual learning objective.
That is not an argument for banning AI. Banning calculators did not save arithmetic. Banning symbolic algebra systems would not improve mathematics. Banning statistical software would not improve statistics. Miao and Holmes (2023), in UNESCO’s guidance on generative AI in education and research, emphasise human agency and institutional capacity rather than treating the technology as a simple yes-or-no problem. That is the right direction.
The question is what universities choose to assess once the tool can perform the procedure.
Figure 2. Procedural automation and persistent epistemic duties. Conceptual schematic, not an empirical scale. Automation can reduce the cost of execution; problem formulation, validity, identification, verification and responsibility remain distinct requirements.
AI will make bad methodology look better
The immediate effect of AI will not necessarily be a flood of obviously dreadful papers. It may be worse. It will make weak work look competent.
Bad academic writing used to contain useful diagnostic information. A confused methods section often sounded confused. A student who did not understand a model might describe it incoherently. Sloppy citations, inconsistent terminology and malformed equations were clues that the underlying work required inspection. Generative models can remove some of those surface defects without repairing the underlying inference.
They can give a weak idea excellent grammar.
They can give a conventional idea the rhetorical appearance of a sophisticated literature review.
They can give a poorly identified model beautifully commented code.
They can give an invalid inference a paragraph full of caveats.
There is empirical reason to treat that as a risk rather than merely a joke. Gao et al. (2023), using a late-2022 version of ChatGPT to generate medical research abstracts, found that blinded human reviewers misclassified 32 per cent of generated abstracts as real; the generated abstracts were plausible in form while containing invented data. Walters and Wilder (2023), studying 636 citations produced by GPT-3.5 and GPT-4 in generated literature reviews, found substantial rates of fabricated references and errors even among real references. Those studies are historical snapshots of particular models, not estimates of the performance of 2026 systems, and model capabilities have changed rapidly. Their durable lesson is narrower: polished scientific form is not evidence that the underlying claims, data or citations have been verified.
This creates an assessment problem. Fluency has always been an imperfect proxy for knowledge, but institutions often relied on it because producing polished technical prose required effort and practice. AI further separates polish from understanding. The prose can be immaculate while the researcher has no idea why the method works.
The correct institutional response is not nostalgia for ugly writing. It is better examination.
Ask the researcher to explain the causal structure without slides. Ask why the standard error changes under a different dependence assumption. Ask what variable would destroy the interpretation if omitted. Ask what happens if the treatment effect is heterogeneous. Ask why the coding scheme corresponds to the historical claim. Ask which archival absences could generate the observed pattern. Ask what the theorem depends upon. Ask which step in the proof fails if an assumption is relaxed. Ask what the model predicts outside the observed sample and why anybody should believe it.
AI can answer those questions too, of course. That means assessment increasingly has to become interactive, adversarial and specific to the work actually submitted. A viva matters more in an AI-rich environment, not less. So do code walkthroughs, replication exercises, oral defence, specification changes performed in real time and questions that require the candidate to connect formal output to substantive meaning.
The old model assessed possession of the procedure.
The new model must assess command of the reasoning.
The danger of agents running agents
There is another step that academic policy has barely begun to absorb. AI systems can now orchestrate other tools and other models. An agent can ask one model to search literature, another to write code, a third to critique the output and a fourth to rewrite the paper. It can call statistical packages, solvers, databases and browsers. It can repeat the cycle until a stopping condition is reached.
At that point, the procedural parts of research become cheap enough to generate at scale.
This is often described as democratisation, and in one sense it is. A researcher without extensive programming experience can perform sophisticated computational tasks. A scholar working in a second language can produce cleaner prose. A small laboratory can automate repetitive analysis. A historian can use machine assistance to organise a large corpus. A mathematician can test conjectures computationally before attempting proof. These are real benefits.
The danger lies in confusing reduced cost with increased understanding.
Suppose an agent can run 500 model variants overnight. That may improve robustness, or it may automate specification searching. Suppose it can generate 50 research questions from a dataset. That may stimulate creativity, or it may industrialise HARKing with better grammar. Suppose it can review 1,000 papers. That may expand coverage, or it may create a synthetic literature review in which nobody has checked whether the cited claims survive contact with the actual papers. Suppose it can produce a journal-formatted manuscript in an afternoon. That may free time for thought, or it may simply increase the number of manuscripts.
Under a sane incentive system, productivity gains should allow researchers to spend more time thinking.
Under a publish-or-perish system, there is a plausible risk that productivity gains become new output expectations. I want to label that correctly: it is an incentive-based forecast, not a measured law of academic labour. The reason to take it seriously is that the literature on metric reactivity and hypercompetition already shows that researchers and institutions respond strategically to rewarded outputs (Edwards & Roy, 2017; Espeland & Sauder, 2007; de Rijcke et al., 2016).
If an institution currently rewards four papers, and AI materially lowers the labour required to produce them, nothing in the metric itself tells the institution to spend the saved time on deeper thought. It may instead raise the expected volume. That outcome is not inevitable; it is a governance choice.
Metrics plus AI could therefore create an academic treadmill in which automated literature production feeds increasingly automated evaluation systems, while humans are retained to certify that machine-generated paperwork satisfies machine-readable criteria. The sentence is deliberately conditional. The future is not identified by a sarcastic paragraph. But the incentive mechanism is coherent enough that universities should design against it before discovering that they have automated the least valuable part of scholarship and doubled its quota.
One hesitates to call this scholarship, but it will certainly have excellent formatting.
What a university is still for
It would be easy to end with a denunciation. That would also be lazy.
Universities still provide things AI systems do not replace. They create durable institutions around knowledge. They preserve archives and laboratories. They assemble communities with competing expertise. They impose deadlines, standards and exposure to criticism. They can force a student to encounter work they would never have chosen voluntarily. They provide supervision in which an experienced researcher notices that the question itself is malformed. They can support long projects whose value is not immediately obvious. They can create reputational consequences for error and misconduct. They can certify competence when the certification is based on real examination rather than paperwork.
The university is most valuable where it is least like a content-delivery platform.
Its comparative advantage is judgement.
That is why excessive templating is so destructive. A university cannot out-template an AI system. It cannot out-format a language model. It cannot win a competition in which the task is to turn a familiar research question into a correctly structured 8,000-word article using standard methods and an approved citation style. Machines are becoming very good at that sort of thing because that sort of thing is, in significant part, pattern completion.
The human value lies elsewhere: choosing important questions, recognising when the standard model is wrong, understanding why a result is surprising, detecting that a dataset does not measure what everybody says it measures, seeing an archival connection others missed, identifying a hidden assumption, inventing a better formalisation, rejecting a fashionable explanation, designing a decisive test, or knowing that the software output is nonsense.
Those are difficult things to rubric.
That does not make them optional.
What should change
If universities want to survive the transition from procedural scarcity to procedural abundance, several changes follow.
First, research-methods education needs to move upward from software operation to inferential reasoning. Students should certainly learn tools, but the software should be interchangeable. If a candidate can run an analysis only in the package used during the module, the candidate has learned a user interface. A methods course should make it possible to explain the same model in equations, code, assumptions and substantive language.
Second, supervisors and examiners need to be willing to say “I do not understand this method; we need someone who does.” That sentence is a mark of professionalism, not weakness. The alternative is to force work back into the evaluator’s comfort zone and then call the result quality assurance.
Third, doctoral programmes should distinguish minimum standards from methodological templates. Ethical requirements, transparency, data management and clear reporting can be standardised. Intellectual strategy cannot be standardised to the same degree without defeating the purpose of doctoral research.
Fourth, assessment needs more defence and less artefact worship. In an AI-rich environment, the polished document is weaker evidence of competence. Oral examination, live methodological questioning, code review, replication, derivation, source criticism and adversarial testing become more important.
Fifth, research evaluation should reduce dependence on simple output metrics. Hicks et al. (2015) did not argue that measurement is useless; they argued for contextual, field-sensitive and judgement-led evaluation. That becomes more urgent when AI can reduce the marginal cost of producing superficially acceptable papers.
Sixth, institutions should reward fewer, better contributions where appropriate. This is not a romantic call for everybody to spend twenty years on a monograph. Different fields have different production functions. It is a call to stop treating volume as a universal proxy for intellectual value.
Seventh, AI use should be disclosed in ways that matter epistemically. “AI was used” is nearly useless as a statement. Used for what? Literature discovery? Translation? Code generation? Model selection? Data cleaning? Statistical execution? Interpretation? Drafting? Figure production? An institution concerned with integrity should care about where human verification occurred and who takes responsibility for the result.
Finally, universities should stop pretending that familiarity is rigour.
A method is not rigorous because it appears in the departmental handbook. A journal is not rigorous because its impact factor is high. A piece of software is not rigorous because the supervisor recognises the menus. A paper is not rigorous because it survived peer review. A thesis is not rigorous because it contains the approved chapter sequence. These things may correlate with quality. None defines it.
The uncomfortable test AI is about to impose
AI will perform a useful service if it forces universities to clarify what they thought they were teaching.
If a student can ask an agent to run SPSS, produce the tables, draft the interpretation, format the references and write the methods chapter, then “ability to produce a methods chapter” is no longer a meaningful endpoint. The institution has to ask a harder question: does the student understand what happened?
If the answer is yes, AI is a tool.
If the answer is no, the student is not the only one with a problem.
The programme has been assessing the wrong thing.
This is why I am less worried about AI destroying academia than I am about AI revealing how much of academia was already procedural. Machines are extremely good at exposing tasks that were mistaken for expertise merely because they were laborious. Once labour disappears, the epistemic content becomes easier to see.
The person who understands statistics has little to fear from an agent that runs statistical software. The person whose status depended on being the only one in the room who knew which menu to open has a more immediate concern.
The same applies to literature reviewing, formatting, coding and routine synthesis. Expertise survives automation when expertise consists of judgement. It looks far more fragile when expertise consists of procedural gatekeeping.
Conclusion: the university after the template
The decline of academia is not a story of stupid students, lazy academics or malevolent administrators. Those caricatures are emotionally satisfying and analytically useless. The more plausible story is institutional.
Higher education expanded, and that expansion brought enormous benefits as well as coordination problems. Institutions responded with more explicit standards and administrative procedures. In some settings those procedures then interacted with competition, rankings, funding systems, completion targets and publication incentives. Metrics made performance legible; once consequential, they also became objects of strategic response. Software made complex procedures easier to execute. Reproducibility debates exposed weaknesses in the assumption that procedural conformity guarantees epistemic reliability. Generative AI and agents are now making parts of procedural execution cheaper still.
That is a mechanism, not a proof of universal historical decline. The causal links are contingent, and good institutions can interrupt them at several points.
Every step is understandable.
The combination is dangerous when the institution mistakes the proxy for the purpose.
Universities remain valuable, but they need to remember what cannot be reduced to a template. Methodology is not the approved sequence of steps. It is the reasoning that justifies those steps. Statistics is not the software output. It is the logic connecting data, assumptions and inference. Scholarship is not the document. It is the intellectual work that makes the document worth reading.
I have spent enough time in universities to know that the good version still exists. I have also spent enough time in them to know how often the system rewards the imitation.
AI will make that imitation cheaper, faster and more polished.
That is not a reason to fear AI.
It is a reason to demand more from academia.
If an agent can now click the SPSS buttons, write the Python, generate the table and draft the paragraph, then perhaps the university can finally stop congratulating people for pressing the buttons and start asking whether they understand the result.
It would be a modest reform.
One might even call it education.
References
Baker, M. (2016). 1,500 scientists lift the lid on reproducibility. Nature, 533, 452–454. https://doi.org/10.1038/533452a
Callaway, B., & Sant’Anna, P. H. C. (2021). Difference-in-differences with multiple time periods. Journal of Econometrics, 225(2), 200–230. https://doi.org/10.1016/j.jeconom.2020.12.001
Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning, 2(1), 67–90. https://doi.org/10.1016/0149-7189(79)90048-X90048-X)
de Rijcke, S., Wouters, P. F., Rushforth, A. D., Franssen, T. P., & Hammarfelt, B. (2016). Evaluation practices and effects of indicator use—A literature review. Research Evaluation, 25(2), 161–169. https://doi.org/10.1093/reseval/rvv038
Edwards, M. A., & Roy, S. (2017). Academic research in the 21st century: Maintaining scientific integrity in a climate of perverse incentives and hypercompetition. Environmental Engineering Science, 34(1), 51–61. https://doi.org/10.1089/ees.2016.0223
Espeland, W. N., & Sauder, M. (2007). Rankings and reactivity: How public measures recreate social worlds. American Journal of Sociology, 113(1), 1–40. https://doi.org/10.1086/517897
Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. npj Digital Medicine, 6, 75. https://doi.org/10.1038/s41746-023-00819-6
Geng, M., & Trotta, R. (2024). Is ChatGPT transforming academics’ writing style? arXiv. https://doi.org/10.48550/arXiv.2404.08627
Goodman, S., & Greenland, S. (2007). Why most published research findings are false: Problems in the analysis. PLOS Medicine, 4(4), e168. https://doi.org/10.1371/journal.pmed.0040168
Goodman-Bacon, A. (2021). Difference-in-differences with variation in treatment timing. Journal of Econometrics, 225(2), 254–277. https://doi.org/10.1016/j.jeconom.2021.03.014
Haven, T. L., Bouter, L. M., Smulders, Y. M., & Tijdink, J. K. (2019). Perceived publication pressure in Amsterdam: Survey of all disciplinary fields and academic ranks. PLOS ONE, 14(6), e0217931. https://doi.org/10.1371/journal.pone.0217931
Hicks, D., Wouters, P., Waltman, L., de Rijcke, S., & Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature, 520, 429–431. https://doi.org/10.1038/520429a
Ioannidis, J. P. A. (2005). Why most published research findings are false. PLOS Medicine, 2(8), e124. https://doi.org/10.1371/journal.pmed.0020124
Ioannidis, J. P. A. (2022). Correction: Why most published research findings are false. PLOS Medicine, 19(8), e1004085. https://doi.org/10.1371/journal.pmed.1004085
Kohrt, F., Smaldino, P. E., McElreath, R., & Schönbrodt, F. (2023). Replication of the natural selection of bad science. Royal Society Open Science, 10(2), 221306. https://doi.org/10.1098/rsos.221306
Lu, C., Lu, C., Lange, R. T., Foerster, J., Clune, J., & Ha, D. (2024). The AI Scientist: Towards fully automated open-ended scientific discovery. arXiv. https://doi.org/10.48550/arXiv.2408.06292
Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
Nosek, B. A., Alter, G., Banks, G. C., Borsboom, D., Bowman, S. D., Breckler, S. J., Buck, S., Chambers, C. D., Chin, G., Christensen, G., Contestabile, M., Dafoe, A., Eich, E., Freese, J., Glennerster, R., Goroff, D., Green, D. P., Hesse, B., Humphreys, M., ... Yarkoni, T. (2015). Promoting an open research culture. Science, 348(6242), 1422–1425. https://doi.org/10.1126/science.aab2374
Power, M. (1997). The audit society: Rituals of verification. Oxford University Press.
Schmidgall, S., Su, Y., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Moor, M., Liu, Z., & Barsoum, E. (2025). Agent Laboratory: Using LLM agents as research assistants. In C. Christodoulopoulos, T. Chakraborty, C. Rose, & V. Peng (Eds.), Findings of the Association for Computational Linguistics: EMNLP 2025 (pp. 5977–6043). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-emnlp.320
Schofer, E., & Meyer, J. W. (2005). The worldwide expansion of higher education in the twentieth century. American Sociological Review, 70(6), 898–920. https://doi.org/10.1177/000312240507000602
Shore, C., & Wright, S. (2015). Governing by numbers: Audit culture, rankings and the new world order. Social Anthropology, 23(1), 22–28. https://doi.org/10.1111/1469-8676.12098
Slaughter, S., & Rhoades, G. (2004). Academic capitalism and the new economy: Markets, state, and higher education. Johns Hopkins University Press.
Smaldino, P. E., & McElreath, R. (2016). The natural selection of bad science. Royal Society Open Science, 3(9), 160384. https://doi.org/10.1098/rsos.160384
Smaldino, P. E., & McElreath, R. (2023). Correction to: ‘The natural selection of bad science’ (2016) by Paul E. Smaldino and Richard McElreath. Royal Society Open Science, 10(9), 231026. https://doi.org/10.1098/rsos.231026
Trow, M. (2006). Reflections on the transition from elite to mass to universal access: Forms and phases of higher education in modern societies since WWII. In J. J. F. Forest & P. G. Altbach (Eds.), International handbook of higher education (pp. 243–280). Springer. https://doi.org/10.1007/978-1-4020-4012-2_13
Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. https://doi.org/10.1038/s41598-023-41032-5
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., & Yu, T. (2024). OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments. arXiv. https://doi.org/10.48550/arXiv.2404.07972