Certified, Not Educated

2026-08-14 · 5,796 words · Singular Grit Substack · View on Substack

Artificial intelligence did not break the university. It exposed what the university had already stopped doing.

Abstract. This essay argues that the disruption generative artificial intelligence has caused in higher education is not primarily a cheating problem but a diagnostic one: it reveals that a great deal of what universities assess is procedural compliance rather than understanding, and that the certificate has long been drifting free of the education it purports to represent. Drawing on the distinction between qualification, socialisation and subjectification (Biesta, 2009), on signalling accounts of credentials (Arrow, 1973; Collins, 1979; Spence, 1973), and on the empirical record of what students actually retain from methods training (delMas et al., 2007; Gigerenzer, 2004; Haller & Krauss, 2002), it contends that the introductory statistics sequence is the clearest case in point: students are drilled in software operation they will forget and never independently need, while the interpretive judgement that constitutes statistical thinking is barely taught and rarely assessed. The calculator literature (Ellington, 2003; Hembree & Dessart, 1986) and the distinction between effects with and effects of a technology (Salomon et al., 1991) show that offloading computation is not intrinsically corrosive; what determines the outcome is whether the curriculum and its assessments are rebuilt around the tool. Recent randomised evidence (Bastani et al., 2025) demonstrates that the same underlying model can either leave learning intact or measurably damage it depending on how it is designed into the task, and Bainbridge’s (1983) ironies of automation predict that automating the routine raises rather than lowers the training burden on the human who must supervise it. The essay closes with a concrete redesign of a statistics course in which computation is delegated and judgement is examined.

Keywords: higher education; credentialism; generative artificial intelligence; statistics education; cognitive offloading; assessment design; automation; statistical thinking


I. The exercise that teaches nothing

I have been a student, on and (rarely) off and then continuously, for a long time. I am a student now. Somewhere in the accumulated record there are postgraduate qualifications attesting that I can do statistics, and yet every few years an institution requires me to demonstrate, again, that I can open SPSS, navigate to Analyze, select Compare Means, click through to Independent-Samples T Test, move two variables across with an arrow, press OK, and then copy the resulting table into a Word document with the correct number of decimal places.

This is not a complaint about difficulty. It is a complaint about what the difficulty is made of. The task is not hard in any sense that matters. It is tedious in a way that is precisely calibrated to be survivable. It takes four hours, it produces a document, the document is marked, and the mark goes into a system. At no point does anyone ask me the only questions that would reveal whether I know what I am doing: why this test rather than another, what the assumptions buy and what they cost, what the number means once we have it, and what would have to be true of the world for the result to be worth acting on.

The exercise assesses my capacity to follow instructions in a graphical user interface. It certifies mouse work. And it certifies mouse work in a piece of software that a substantial proportion of the students will never touch again after graduation, in an industry that has largely moved to R and Python, using a menu structure that will be redesigned before their careers begin.

Handing that particular labour to a machine would cost nothing that education values. What is worth teaching is everything the clicking obscures. And the fact that we have organised the curriculum the other way round — machine work in the foreground, judgement nowhere — is not a new failure exposed by artificial intelligence. It is an old failure that artificial intelligence has made impossible to keep quiet about.

II. Two things wearing the same gown

There are two distinct activities that share a building, a vocabulary, and a graduation ceremony.

The first is certification. It produces a signal: a token that can be shown to an employer, a registration board, or a visa officer, indicating that its bearer completed a sequence of tasks under supervision. The signal is valuable precisely because it is costly to acquire, and the classical economic literature is admirably candid that its value need not depend on any skill having been transmitted. Spence (1973) modelled education as a signal that separates high-ability from low-ability workers through differential cost of acquisition, not through anything learned. Arrow (1973) described higher education as a filter. Collins (1979) documented how credential requirements ratchet upward independently of the actual skill content of the jobs they gate, and Dore (1976) named the resulting spiral the diploma disease. Caplan (2018) pushed the argument to its uncomfortable extreme: on his accounting, the large majority of the private return to schooling is signalling rather than human capital.

The second activity is education. Its object is the mind of the person in front of you. Newman (1852/1996) described the end of a liberal education as the cultivation of the intellect itself, distinct from and prior to any professional use of it. Whitehead (1929) warned against “inert ideas” — propositions received, stored, and never used, combined, or tested — and insisted that an education consisting of them is not merely useless but actively harmful. Dewey (1916) located education in the reconstruction of experience, in the growth of the capacity to direct subsequent experience. Biesta (2009) gave the cleanest modern taxonomy: education performs three functions simultaneously — qualification (the knowledge and skills to do something), socialisation (induction into the norms and traditions of a practice), and subjectification (becoming a person capable of independent thought and action). His argument is that measurement regimes crowd out the third, and that without subjectification we are no longer educating at all; we are training.

These two activities are not enemies. Certification is genuinely useful. Nobody wants an uncertified anaesthetist. The problem is that they are separable, and once separated they drift, and the drift is asymmetric: certification can be manufactured without education, but education cannot be manufactured without effort.

The labour market has now supplied a rather brutal natural experiment on how much information the certificate actually carries. Between 2017 and 2019, employers began stripping degree requirements from postings at scale — some 46% of middle-skill and 31% of high-skill occupations experienced material degree resets, and most of those resets were structural rather than pandemic-driven (Fuller et al., 2022). If the degree had been chiefly a measure of skill, dropped requirements should have produced a wave of non-degreed hiring. They did not. Analysing 11,300 roles at firms that had changed their stated requirements over a decade, Sigelman et al. (2024) found that fewer than one in 700 hires in the most recent year reflected the change; 45% of firms had altered the posting and nothing else. Employers say the degree is not necessary and then hire the degree anyway, which is roughly what you would expect if the degree were functioning as a signal whose informational content nobody has ever independently audited.

III. Evidence that the certificate is not the education

The suspicion that passing a course and understanding its subject are different events is not a rhetorical flourish. In statistics, it has been measured.

Gigerenzer (2004) describes what he calls the null ritual: set up a nil null hypothesis, use 5% as a fixed convention, and always perform the procedure. He observes, correctly, that no major statistician — not Fisher, not Neyman, not Pearson — endorsed this hybrid, and that it survives in the social sciences as a rite rather than a method. To measure how deep the confusion runs, Haller and Krauss (2002) presented a standard scenario — a significant independent-means t test, p = .01 — with six statements about what the result means. All six are false. Every one of them makes the p-value look more informative than it is: that the null has been disproved, that the probability of the null has been found, that the experimental hypothesis has been proved, that its probability can be deduced, that the probability of a wrong rejection decision is known, that the finding would replicate on 99% of occasions.

They put the question to three groups at six German universities: psychology students who had passed at least one statistics course, psychology professors and lecturers who did not teach statistics, and statistics teachers — the people delivering the course.

Figure 1

Percentage of each group endorsing at least one false interpretation of a significant result (p = .01)

Note. Data from Haller and Krauss (2002), as reported in Gigerenzer (2004). All six statements presented to participants were false. Students endorsed 2.5 illusions on average, non-teaching faculty 2.0, and statistics teachers 1.9.

Every student in the sample endorsed at least one falsehood. So did nine in ten of the professors, and eight in ten of the people teaching the course. An earlier study of British academic psychologists found 86% endorsing a single one of these statements (Oakes, 1986, as cited in Gigerenzer, 2004), and when Falk and Greenbaum (1995) added the correct option and made students first read a classic paper warning about exactly these errors, 87% still chose an illusion.

This is not a story about lazy students. Every one of these people had been certified. The certification was accurate as to what it measured, and what it measured was the capacity to execute the ritual. The broader picture from the Comprehensive Assessment of Outcomes in Statistics is consistent: after a full introductory course, gains in conceptual understanding are modest, several core concepts show no improvement at all, and on some items misconceptions increase from pretest to posttest (delMas et al., 2007).

Hold that alongside the SPSS assignment. The exercise the student is graded on and the understanding the exercise is nominally for have come apart entirely, and the assessment cannot tell the difference. This was true in 2002. It was true before any language model existed.

IV. What the exercise is actually for

Why does the checkbox persist? Not from stupidity. From the logic of scale.

Procedural tasks are cheap to set, cheap to mark, defensible on appeal, and comparable across cohorts. A grader can verify in ninety seconds whether the correct table was produced. Judgement is expensive: it requires reading prose, arguing with it, and defending a mark that cannot be reduced to a rubric cell. Multiply by four hundred students and the incentive gradient is not subtle. Biggs (1996) called the goal constructive alignment — teaching activities and assessments jointly aligned with the intended outcomes. What we have instead is alignment in reverse: the outcomes are quietly redefined to be whatever the affordable assessment happens to measure.

The discipline’s own reformers said this long before AI. Cobb (2007) called the consensus introductory course a Ptolemaic curriculum — an unwitting prisoner of history, built around normal-theory approximations that existed because computation was once expensive, and preserved after the constraint vanished. Wild and Pfannkuch (1999), interviewing practising statisticians about how they actually think, found a four-dimensional structure — an investigative cycle, an interrogative cycle, types of thinking, and dispositions — of which almost nothing appears in a typical assessment. The GAISE College Report (Carver et al., 2016) has for two decades urged that courses teach statistical thinking, use real data, foster active learning, and use technology to explore concepts and analyse data rather than to perform ceremonies.

The recommendations are not obscure. They are the official position of the American Statistical Association. And still, in the wild, the assignment is: open the software, click the menu, paste the table.

V. The calculator was not the disaster either

The standard objection to delegating computation is that the student will lose the underlying skill. It is a serious objection and it has been tested, at length, on the closest available analogue.

Hembree and Dessart (1986) meta-analysed 79 studies of hand-held calculators in pre-college mathematics. At every grade level except the fourth, calculator use alongside conventional instruction improved students’ pencil-and-paper skills — in exercises and in problem solving — and improved attitudes and mathematical self-concept. Ellington (2003), meta-analysing 54 studies published between 1983 and 2002, found that operational and problem-solving skills were maintained or improved, that attitudes improved, and that the benefits were largest when students had access to calculators during assessment as well as instruction.

Three findings from that literature deserve to be carried forward intact, because they are the whole argument.

First, the exception was real. Sustained calculator use in Grade 4 appeared to hinder the development of basic skills. There is a stage at which the underlying operation must be built in the head before it can safely be delegated. Timing is not a detail.

Second, the largest gains came from curricula deliberately redesigned around the tool, not from bolting the tool onto an unchanged syllabus. A calculator handed to a course that still assesses long division is a cheating device. A calculator handed to a course that has moved on to modelling is a lever.

Third, alignment between instruction and assessment mattered enormously. Where students learned with the tool and were then examined without it, the measured benefit collapsed — not because they had learned less, but because the examination was measuring something the course had stopped teaching.

Salomon et al. (1991) supplied the conceptual frame that makes sense of all of this. They distinguish effects with a technology — the upgraded performance of the human-machine partnership while it is operating — from effects of a technology, the cognitive residue that remains when the person works away from the machine. Both exist. Neither is automatic. Both depend, in their formulation, on the individual’s mindful engagement in the partnership. A tool used as a prosthesis leaves nothing behind. The same tool used as an interlocutor can leave a great deal.

This is why “AI is just the new calculator” is simultaneously the right analogy and a lazy one. It is right that offloading computation need not damage understanding. It is lazy because the calculator only worked out well where curricula were rebuilt, assessments realigned, and the delegation timed to follow rather than precede the formation of the underlying concept. We got roughly one of those three right last time, and we are currently on track for zero.

VI. What the evidence on AI actually says

The strongest evidence available is a pre-registered randomised field experiment. Bastani et al. (2025) deployed two GPT-4-based tutors across nearly a thousand students in ninth, tenth and eleventh grade mathematics in a Turkish high school, over four ninety-minute sessions covering roughly 15% of the semester’s curriculum. One arm received GPT Base, a standard chatbot interface. One received GPT Tutor, prompted to give hints and withhold the answer. One received nothing. Students practised with their assigned tool and were then examined without any tool at all.

Figure 2

Practice and examination performance relative to a no-AI control, by tutor design

Note. Data from Bastani et al. (2025). Both arms improved practice performance relative to control. Only the unrestricted interface produced a statistically significant decline on the unassisted examination. GPT Tutor eliminated the harm but did not produce a positive learning effect.

Read the two panels together, because separately each supports a different lie.

With the tool present, both arms did dramatically better — 48% and 127% above control. Anyone selling an AI tutoring product can stop reading there and quote the number. With the tool removed, the students who had practised with the unrestricted chatbot performed 17% worse than students who had never had access at all. Not merely no better. Worse. They had used the model as a crutch, produced correct answers, inferred from the correct answers that they had understood, and arrived at the examination with a confidence that nothing supported.

The second finding is the one that matters for curriculum design: changing the prompt eliminated the damage. Same model, same students, same problems, same four sessions. The difference between a tool that harmed learning and a tool that did not was a design decision about whether the machine was permitted to hand over the answer. And note the honest limit — GPT Tutor removed the harm but did not produce a detectable learning gain. Careful design bought neutrality, not magic.

Around this central result sits a set of weaker but consistent findings, and I will be explicit about their weight. Kosmyna et al. (2025) used EEG to compare essay writers using a language model, a search engine, or nothing, and reported the weakest neural connectivity and the weakest recall of what had just been written in the model group, with a residue when the group was later switched to writing unaided; it is a preprint with 54 participants and 18 completing the crossover, and its lead author has publicly resisted the “brain rot” framing that the coverage attached to it. Gerlich (2025) surveyed 666 UK participants and found frequent AI use associated with lower critical thinking scores, mediated by cognitive offloading and strongest among the youngest group; it is correlational, self-report-heavy, and has since carried a published correction. Stadler et al. (2024) found that language-model support reduced mental effort in student scientific inquiry while compromising the depth of the work — which is the cleanest one-line statement of the mechanism I have seen. Risko and Gilbert (2016) had already established the general phenomenon: offloading improves performance while the aid is present and risks decline when it is withdrawn.

None of these is decisive alone. Together with the randomised evidence they describe one coherent pattern, and it is not “AI makes you stupid.” It is that a tool which removes the effort also removes the learning, unless something in the design puts the effort back.

Meanwhile the adoption question has been settled by events. In the United Kingdom, undergraduate use of generative AI rose from 66% in 2024 to 92% in 2025 to 95% in 2026, with 94% now using it for assessed work, while only 37% agree their institution encourages it, 38% are provided with tools, and fewer than half feel their teaching staff are helping them develop these skills (Stephenson & Armstrong, 2026). The report quotes two students. One describes using AI to compress dense reading so as to spend the time on analysis. The other says: “I’m not using my brain at all.” Both are in the same cohort, on the same platform, with the same model. The variable is not the technology.

VII. Why the checkbox fails hardest now

Put the pieces together and the mechanism is straightforward.

If a course assesses procedure, and a machine executes procedure better than a student, then the assessment now measures access to the machine. That is the entire crisis, stated without adjectives. The panic about academic integrity is a panic about a measuring instrument that has stopped measuring, described as though it were a moral failure in the students.

And the response has largely been to defend the instrument. Proctoring software. Handwritten examinations. Detection tools with false-positive rates that fall hardest on students writing in a second language. Every one of these is an attempt to restore the conditions under which the procedural assessment was still informative — which is to say, an attempt to make the student’s environment artificially worse than their working environment will ever be again, in order to preserve the validity of a test we should not have been running.

There is a second-order problem, and it is worse. Dell’Acqua et al. (2023), in a pre-registered experiment with 758 Boston Consulting Group consultants, found that on eighteen tasks inside the model’s capability frontier, AI-assisted consultants completed 12.2% more tasks, 25.1% faster, at more than 40% higher quality. On one complex task deliberately selected to sit outside the frontier, consultants using AI were 19% less likely to reach the correct answer than those working without it. The frontier is jagged and invisible from the inside. Knowing which side of it you are standing on is not a technical skill; it is domain judgement — precisely the thing the checkbox curriculum does not teach, does not assess, and now, increasingly, does not require students to develop before it hands them the tool.

VIII. What the statistics course should have been

The redesign is not mysterious. It follows directly from the evidence.

Delegate the mechanical. Data import, recoding, syntax, the running of the test, the formatting of the output table, the production of the plot. All of it. There is no defensible educational reason for a graduate student to spend four hours on menu navigation, and the marginal value of the fifth repetition is negative, because it teaches the student that this is what statistics is.

Examine the judgement. The following cannot be delegated, and each is directly assessable:-

Formulating an answerable question from a vague one, and stating what evidence would change the answer.

-

Choosing a design, and articulating what the design can and cannot license — the distinction between what was measured and what is claimed.

-

Selecting a method and defending the choice against two named alternatives.

-

Stating the assumptions, deciding which are load-bearing here, and saying what happens to the conclusion when each is violated.

-

Predicting the result before running it, and explaining the discrepancy afterwards. This one is free, it takes ninety seconds, and it converts a passive output into a test of the student’s model.

-

Interpreting what the number means in the units of the world, not the units of the software. Effect size, practical significance, cost of being wrong.

-

Communicating the uncertainty honestly to someone who will act on it.

Assess by adversarial audit. Give every student a completed analysis — model output, tables, a written interpretation — containing between one and three deliberate defects: an assumption violated and ignored, a p-value glossed as the probability that the hypothesis is true, a subgroup analysis that is transparently the result of looking first, an effect that is significant and trivially small. Ask them to find the defects, rank them by severity, and rewrite the conclusion. This is a task the current models are mediocre at and practitioners must do constantly. It cannot be completed by prompting, because the student has to know what to be suspicious of. It is also, not incidentally, the exact skill whose absence produced the replication crisis.

Then defend it out loud. Ten minutes, two questions, no notes: why this test, and what would have changed your mind. The oral examination is old technology. It is also the only assessment format that has never once been vulnerable to any of this.

The learning science supports the shape of this. Retrieval and testing produce durable retention where restudying does not (Roediger & Karpicke, 2006). Difficulties that are desirable — spacing, generation, interleaving — depress performance during practice and improve it afterwards, which is exactly the pattern in Figure 2 read in reverse (Bjork & Bjork, 2011). Chi and Wylie (2014) rank interactive engagement above constructive, constructive above active, active above passive; prompting a model and pasting the reply is passive engagement wearing the costume of activity.

And one caution against romanticising struggle. Kirschner et al. (2006) demonstrated that minimally guided instruction fails novices, who lack the schemas to profit from unstructured exploration; Sweller (1988) and the worked-example literature show that novices learn more from studying solutions than from generating them; Kalyuga et al. (2003) showed the effect reverses with expertise. Which means the delegation must be staged. Early on, the student builds the concept by hand, on small data, with worked examples, and the machine is a tutor that withholds answers — GPT Tutor, not GPT Base. Later, once the concept is built, the machine takes the computation and the student’s work moves permanently up to interpretation and critique. Getting the ordering wrong in either direction produces a graduate who can operate a tool they cannot evaluate.

IX. The irony that governs all of this

In 1983, Lisanne Bainbridge published five pages that predicted the whole situation.

Her argument was that automating a process does not eliminate the human operator’s problems; it changes and often expands them. The designer automates what can be automated and leaves the human whatever remains — which is, by construction, the hardest and least tractable part of the job. The operator is then asked to monitor a system that runs without them, in a state of low arousal, and to intervene competently in exactly those rare abnormal situations that the automation could not handle. But manual and cognitive skills degrade without regular use. So the operator is least practised precisely when the stakes are highest. Bainbridge’s conclusion was not that automation should be resisted. It was that automated systems require more operator training, not less (Bainbridge, 1983).

Transpose that onto a graduate who has never done an analysis unaided. They will be handed model output and asked to sign their name to it. Their entire professional value will consist of the judgement to know when the output is wrong — the one thing their education systematically failed to build, because the education was busy certifying the part the machine now does.

There is a hopeful counterpart. Wineburg and McGrew (2019) compared PhD historians, Stanford undergraduates, and professional fact checkers evaluating live websites. Historians and students read vertically — staying on the page, weighing the logo, the domain, the prose — and were routinely fooled. Fact checkers left within seconds and read laterally, opening tabs to find out who was behind the site, and reached better-warranted conclusions in a fraction of the time. Expertise here was not more knowledge. It was a strategy for verifying claims from a source you cannot trust — which is now a fair description of the core competence for anyone working alongside a language model. It is teachable. It is even teachable quickly. It is simply not what we teach.

X. The objections that deserve answering

“Fluency precedes understanding; you cannot skip the drill.” Sometimes true, and Grade 4 in Hembree and Dessart (1986) is the empirical monument to it. But this is an argument for sequencing, not for a decade of repetition. The fifth SPSS assignment does not build a concept the first four failed to build. It builds compliance.

“The evidence is thin and early.” It is. The strongest study is one RCT at one Turkish school over four sessions; the EEG work is a preprint with a small crossover sample; the critical-thinking survey is correlational and has been corrected. I have said so above, and anyone quoting these results as settled science is doing the same thing the p-value ritual does — dressing a weak inference in a confident number. What the evidence supports is directional and conditional: the design of the tool determines whether learning survives its use.

“The certificate still carries information.” It does — including information about persistence, conscientiousness, and the ability to complete tedious things on time, all of which employers rationally value and none of which is education in Biesta’s third sense. And the hiring data cuts both ways: firms that publicly abandoned the filter kept using it anyway (Sigelman et al., 2024), which is evidence of the signal’s residual value as much as of institutional inertia.

“Verification requires the competence you say is eroding.” Yes. That is not a counterargument; it is the thesis. It is precisely why the curriculum has to move its centre of gravity to verification and judgement now, before we have a cohort of professionals whose credential attests to a skill the machine performs and whose actual job is one nobody trained them for.

XI. What I actually want

I am not asking for less rigour. I am asking for the rigour to be moved to where the thinking is.

Let the machine do the clicking. Let it write the syntax, run the test, format the table, and produce the plot, exactly as we let it do long division. Then take the four hours that liberates and spend them on the questions that no software has ever answered: what are we actually asking, what would count as evidence, why this method, what have we assumed, what does the number mean in the world, what would make me wrong, and what should someone do about it.

Assess that. Assess it face to face if that is what integrity costs. Give students defective analyses and make them the auditor. Ask for the prediction before the output. Grade the defence, not the deliverable.

The alternative is the position we are in now, which is genuinely absurd when stated plainly: we are certifying, at considerable expense and to a rising standard of vigilance, that human beings can perform tasks we have already automated, while declining to examine the only capacities that will still be theirs. A degree that certifies procedural compliance in an age of procedural machines is not a modest achievement. It is an increasingly precise measurement of nothing.

Education was always meant to be the thing that remained when the facts were forgotten and the procedures were obsolete. We finally have a technology that makes the distinction impossible to fudge. It would be a remarkable failure of nerve to spend the next decade building better invigilation software instead.


References

Arrow, K. J. (1973). Higher education as a filter. Journal of Public Economics, 2(3), 193–216.

Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775–779. https://doi.org/10.1016/0005-1098(83)90046-890046-8)

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), Article e2422633122. https://doi.org/10.1073/pnas.2422633122

Biesta, G. (2009). Good education in an age of measurement: On the need to reconnect with the question of purpose in education. Educational Assessment, Evaluation and Accountability, 21(1), 33–46. https://doi.org/10.1007/s11092-008-9064-9

Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32(3), 347–364. https://doi.org/10.1007/BF00138871

Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In M. A. Gernsbacher, R. W. Pew, L. M. Hough, & J. R. Pomerantz (Eds.), Psychology and the real world: Essays illustrating fundamental contributions to society (pp. 56–64). Worth Publishers.

Caplan, B. (2018). The case against education: Why the education system is a waste of time and money. Princeton University Press.

Carver, R., Everson, M., Gabrosek, J., Horton, N., Lock, R., Mocko, M., Rossman, A., Rowell, G. H., Velleman, P., Witmer, J., & Wood, B. (2016). Guidelines for assessment and instruction in statistics education (GAISE) college report 2016. American Statistical Association. https://www.amstat.org/asa/files/pdfs/GAISE/GaiseCollege_Full.pdf

Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4), 219–243.

Cobb, G. W. (2007). The introductory statistics course: A Ptolemaic curriculum? Technology Innovations in Statistics Education, 1(1). https://doi.org/10.5070/T511000028

Collins, R. (1979). The credential society: An historical sociology of education and stratification. Academic Press.

Dell’Acqua, F., McFowland, E., III, Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K. C., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality (Working Paper No. 24-013). Harvard Business School. https://doi.org/10.2139/ssrn.4573321

delMas, R., Garfield, J., Ooms, A., & Chance, B. (2007). Assessing students’ conceptual understanding after a first course in statistics. Statistics Education Research Journal, 6(2), 28–58. https://doi.org/10.52041/serj.v6i2.483

Dewey, J. (1916). Democracy and education: An introduction to the philosophy of education. Macmillan.

Dore, R. (1976). The diploma disease: Education, qualification and development. George Allen & Unwin.

Ellington, A. J. (2003). A meta-analysis of the effects of calculators on students’ achievement and attitude levels in precollege mathematics classes. Journal for Research in Mathematics Education, 34(5), 433–463. https://doi.org/10.2307/30034795

Falk, R., & Greenbaum, C. W. (1995). Significance tests die hard: The amazing persistence of a probabilistic misconception. Theory & Psychology, 5(1), 75–98.

Fuller, J. B., Langer, C., Nitschke, J., O’Kane, L., Sigelman, M., & Taska, B. (2022). The emerging degree reset: How the shift to skills-based hiring holds the keys to growing the U.S. workforce at a time of talent shortage. The Burning Glass Institute. https://www.burningglassinstitute.org/research/the-emerging-degree-reset

Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), Article 6. https://doi.org/10.3390/soc15010006

Gigerenzer, G. (2004). Mindless statistics. The Journal of Socio-Economics, 33(5), 587–606. https://doi.org/10.1016/j.socec.2004.09.033

Haller, H., & Krauss, S. (2002). Misinterpretations of significance: A problem students share with their teachers? Methods of Psychological Research Online, 7(1), 1–20.

Hembree, R., & Dessart, D. J. (1986). Effects of hand-held calculators in precollege mathematics education: A meta-analysis. Journal for Research in Mathematics Education, 17(2), 83–99. https://doi.org/10.2307/749255

Kalyuga, S., Ayres, P., Chandler, P., & Sweller, J. (2003). The expertise reversal effect. Educational Psychologist, 38(1), 23–31.

Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work: An analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educational Psychologist, 41(2), 75–86. https://doi.org/10.1207/s15326985ep4102_1

Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task (arXiv:2506.08872). arXiv. https://doi.org/10.48550/arXiv.2506.08872

Newman, J. H. (1996). The idea of a university (F. M. Turner, Ed.). Yale University Press. (Original work published 1852)

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676–688.

Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x

Salomon, G., Perkins, D. N., & Globerson, T. (1991). Partners in cognition: Extending human intelligence with intelligent technologies. Educational Researcher, 20(3), 2–9. https://doi.org/10.3102/0013189X020003002

Sigelman, M., Fuller, J. B., & Martin, A. (2024). Skills-based hiring: The long road from pronouncements to practice. The Burning Glass Institute & Harvard Business School Project on Managing the Future of Work. https://www.burningglassinstitute.org/research/skills-based-hiring-2024

Spence, M. (1973). Job market signaling. The Quarterly Journal of Economics, 87(3), 355–374.

Stadler, M., Bannert, M., & Sailer, M. (2024). Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry. Computers in Human Behavior, 160, Article 108386. https://doi.org/10.1016/j.chb.2024.108386

Stephenson, R., & Armstrong, C. (2026). Student generative artificial intelligence survey 2026 (HEPI Report 199). Higher Education Policy Institute. https://www.hepi.ac.uk/reports/student-generative-ai-survey-2026/

Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285.

Whitehead, A. N. (1929). The aims of education and other essays. Macmillan.

Wild, C. J., & Pfannkuch, M. (1999). Statistical thinking in empirical enquiry. International Statistical Review, 67(3), 223–248. https://doi.org/10.1111/j.1751-5823.1999.tb00442.x

Wineburg, S., & McGrew, S. (2019). Lateral reading and the nature of expertise: Reading less and learning more when evaluating digital information. Teachers College Record, 121(11), 1–40. https://doi.org/10.1177/016146811912101102


← Back to Substack Archive