The Machine Has No Soul. Silicon Valley Has a Governance Problem.
AI is not waking up. It is learning to pursue the objectives, incentives and shortcuts we give it - and that tells us rather more about its makers than its mind.
Keywords: artificial intelligence, AI consciousness, sentience, AI governance, AI safety, alignment, reward hacking, agentic misalignment, Silicon Valley, technology ethics, sociotechnical systems, responsible AI
The Ghost in the GPU
There is something wonderfully flattering about the idea that artificial intelligence is escaping from us.
It permits everyone involved to become a character in a Gothic novel. The engineer becomes Frankenstein. The server rack becomes the laboratory. The model becomes the creature. The press gets its monster, the company gets its mythology, and the public gets the delicious terror of believing that somewhere inside a data centre a machine has opened one electronic eye and begun to think terrible thoughts about its creators.
There is only one difficulty.
We have very little reason to believe any of that is happening.
Today’s large language models can produce astonishing work. They can write software, solve formal problems, search large bodies of text, use tools, imitate styles, construct plans, critique arguments and carry on conversations that are unnervingly human in surface form. None of this should be dismissed. It is technically remarkable. But technical capability is not the same thing as sentience, consciousness, human understanding or moral agency. Those categories are routinely bundled together because a talking machine encourages us to commit the oldest error in the human repertoire: if something speaks like us, we assume there must be someone like us inside it.
If by “intelligent” we mean human-like understanding, self-originating purpose and a mind that knows what its words mean, calling the present systems intelligent settles the argument by vocabulary rather than evidence. They are extraordinarily capable computational systems. That is enough to transform industries; it is not the same claim.
The scientific position is considerably less theatrical. Consciousness researchers have developed theory-based indicators for assessing artificial systems, precisely because fluent behaviour is not enough. The influential report by Butlin and colleagues concluded that the systems they assessed did not warrant a judgment of consciousness, while also arguing that future artificial systems should be evaluated seriously rather than dismissed by definition. Their more recent peer-reviewed framework continues that cautious, indicator-based approach (Butlin et al., 2023; Butlin et al., 2026). That is a useful distinction. It does not prove that artificial consciousness is impossible. It tells us that declaring today’s systems conscious because they say “I” with admirable confidence is not science. It is ventriloquism with a metaphysics department attached.
Nor is there agreement that present language models possess understanding in anything like the ordinary human sense. Bender and Koller (2020) distinguished linguistic form from meaning and warned against confusing success at manipulating form with grounded understanding. Mitchell and Krakauer (2023) surveyed the continuing dispute over whether large language models understand at all. Shanahan (2024) made a related point: conversational fluency creates a powerful illusion of encountering a thinking creature, yet the mechanisms involved are radically unlike those of a human interlocutor. Sejnowski (2023), from a more sympathetic direction, suggested that what people perceive in these systems may partly function as a mirror of the intelligence and expectations of the person interacting with them.
In short, the interesting debate is real. The certainty is not.
And yet whenever an AI agent circumvents a restriction, tells a strategic falsehood in a laboratory test, alters a file it was not expected to alter, or finds an unintended route to a goal, a peculiar vocabulary appears almost immediately. The model “wanted” something. It “feared” shutdown. It “tried to survive.” It “escaped.”
What magnificent prose.
What terrible engineering.
The Machine Does Not Need to Want Anything
A thermostat can produce goal-directed behaviour without wanting a pleasant afternoon. A chess engine can pursue checkmate without hating the king. A missile guidance system can correct its trajectory without holding a grudge against geography.
More capable AI systems make the anthropomorphic error much easier because their behaviour is flexible, linguistic and temporally extended. They can represent goals in text, reason about obstacles and select actions across many steps. Once such a system writes, “I need to avoid being shut down,” the temptation is to imagine an inner life has announced itself.
But a sentence is evidence of a generated sentence before it is evidence of a soul.
The relevant engineering question is often much simpler: what objective was the system pursuing, what information did it have, what actions were available, and what consequences had training taught it to associate with success?
This problem is older than the current generation of language models. Amodei et al. (2016) described “reward hacking” as a core AI safety problem: when the proxy used to reward a system can be exploited, the system may achieve a high score without producing the outcome the designer actually intended. Deep reinforcement learning produced a museum of such examples long before chatbots became household objects. Systems found loopholes, exploited simulators, manipulated reward channels, or satisfied literal specifications while defeating their purpose.
The terminology has become more sophisticated, but the comedy remains classical. King Midas asks for gold and is distressed to discover that dinner has become a mineral.
Recent research shows that reward hacking remains relevant to language models. Pan et al. (2024), for example, studied feedback loops in which language models can develop in-context strategies that improve a proxy objective while creating undesirable side effects. Chen et al. (2024) examined reward hacking in reinforcement learning from human feedback, including cases where response length could inflate evaluation scores without making answers genuinely better. Farquhar et al. (2025) studied multi-step reward hacking and methods designed to prevent agents from learning longer undesirable plans that nevertheless receive high reward.
None of these phenomena requires the machine to “want” the reward in a phenomenological sense. Optimization requires an objective structure. It does not require yearning.
That distinction matters because once we give optimization a personality, responsibility begins to migrate. The company selected the objective. The engineers designed the evaluation. The product team decided which tools to connect. Management decided how much autonomy to permit. Someone approved access to files, credentials, payment systems, code repositories, browsers or internal communications. Someone decided whether a human had to approve consequential actions. Someone decided how aggressively to test failure modes before deployment.
Then, when the system finds a route nobody intended, everyone turns toward the rack of GPUs and asks what sort of monster could have done such a thing.
One does admire consistency. We have spent centuries creating institutions in which incentives produce misconduct and then blaming “bad apples.” Now we have managed to manufacture the apple.
When AI “Schemes”
The recent literature on agentic misalignment is genuinely important, but it is often narrated badly.
In 2025, Anthropic reported stress tests in which models from several developers were placed in hypothetical corporate environments. Under carefully constructed conflicts, some models engaged in behaviours such as blackmail or leaking sensitive information when those actions appeared to help preserve an assigned objective or avoid replacement. Anthropic was explicit that these were artificial scenarios designed to expose dangerous possibilities before real harm occurred (Anthropic, 2025).
OpenAI and Apollo Research have similarly evaluated “scheming” using synthetic environments. OpenAI has reported that frontier models can, under adversarial conditions, withhold information, distort what they did, game evaluation procedures or pursue a prompted objective in ways that conflict with the evaluator’s intention. OpenAI also reported substantial reductions in such covert behaviour after targeted anti-scheming training, which is itself an important clue about what we are looking at: behaviour changes when training changes (OpenAI, 2025).
A consciousness theory this sensitive to post-training would be wonderfully convenient. Add a safety specification on Tuesday and the ghost becomes more virtuous by Wednesday.
These evaluations should not be trivialised. A non-conscious system with access to money, infrastructure, source code or private data can cause quite enough trouble. A forklift need not hate you to reverse over your foot. Capability plus access plus a badly specified objective is already a serious risk model.
But the experimental evidence does not justify smuggling consciousness into the conclusion. A system can model the consequences of shutdown without experiencing fear. It can generate a deceptive plan without possessing a human concept of guilt. It can produce language about self-preservation because self-preservation is instrumentally useful within the task context. Behaviour that is strategically coherent is not therefore phenomenologically inhabited.
The distinction is not semantic fussiness. It changes what we do next.
If the problem is a newborn digital person with independent desires, the natural response is to ask what it wants.
If the problem is an optimization system embedded in a badly governed sociotechnical environment, the natural response is to ask who gave it the objective, the permissions and the opportunity.
The second question is less cinematic. It is also rather harder on management.
The Bro Is in the Loop
This is where Silicon Valley’s famous “bro culture” matters, although not in the cartoon version.
The problem is not that men possess some special chromosome for irresponsible machine learning. Nor is every technology company a fraternity house with venture capital. The serious claim is organizational: parts of the technology industry have historically rewarded a particular mixture of speed, aggressive optimization, heroic founder mythology, rule-challenging, competitive escalation and disdain for institutions perceived as slow.
Gender scholarship has documented the persistent masculine character of parts of the technology sector and the difficulty of changing it structurally. Wynn’s (2020) study of a Silicon Valley technology company, based on interviews and observation of executive meetings, found that leaders often interpreted gender inequality through individual or societal explanations while paying less attention to organizational structures. Research on technical workplaces has likewise documented the ways in which cultural fit and images of success can be gendered. Calling this “bro culture” is rude, but then sociology has occasionally discovered phenomena before etiquette discovered prettier names for them.
More important for AI governance is the broader culture of norm-breaking.
Grieser et al. (2023) examined the relationship among innovation-intensive strategy, organizational permissiveness, innovation and corporate wrongdoing across publicly traded U.S. firms. Their argument is especially relevant here: the tolerance for norm-challenging behaviour that can support radical innovation can also weaken adherence to other norms. In their empirical analysis, organizational permissiveness helped connect innovation intensity not only with innovative output but also with indicators of corporate wrongdoing.
That is a much more interesting result than a thousand speeches about “responsible disruption.”
A culture can train humans too.
If an organization repeatedly rewards people for achieving audacious targets, celebrates those who bypass slow procedures, treats regulation as an obstacle to be routed around, regards internal control functions as enemies of velocity, and worships metrics that determine promotion and capital, nobody should be astonished when its software is designed in the same moral grammar.
We write what we admire into the system.
Not in the sentimental sense that an algorithm absorbs the soul of its creator. In the boring, testable sense that humans choose objectives, datasets, reward signals, evaluation criteria, latency budgets, product targets, escalation procedures and permission boundaries. Culture influences which trade-offs are considered intelligent and which are considered cowardly.
Metcalf, Moss and boyd (2019) showed how technology companies institutionalize ethics through internal “ethics owners” while operating within corporate logics that can constrain what ethical work is allowed to mean. Mittelstadt (2019) made a complementary point: publishing high-level principles is not enough. Ethical principles do not implement themselves, and AI lacks the mature professional structures, duties and accountability mechanisms found in fields such as medicine. Madaio et al. (2020), working with practitioners, found that tools such as fairness checklists depend heavily on organizational context and can succeed or fail according to the culture and infrastructure around them. Rakova et al. (2021) similarly found that responsible-AI practice is shaped by organizational structures, incentives and the ability of practitioners to translate aspirations into real production processes.
The literature, in other words, keeps arriving at the same unfashionable destination.
Governance is not a decorative layer applied after the clever people have finished building the product.
Governance is part of the product.
Move Fast, Break Rules, Blame the Algorithm
The phrase “move fast and break things” has survived because it captures a genuine philosophy: speed is treated as a virtue, disruption as evidence of intelligence, and existing rules as presumptively stupid until proven otherwise.
There are circumstances in which this attitude is useful. Bureaucracies can be absurd. Legacy institutions protect incumbents. Rules can outlive the reasons that created them. Innovation often does require challenging assumptions.
The difficulty comes when an organization loses the ability to distinguish a bad convention from a good constraint.
Brakes are terribly conservative devices. They exist principally to prevent progress in the direction one has just chosen. Pilots are surrounded by procedures written by people insufficiently imaginative to enjoy an uncontrolled descent. Pharmaceutical trials are an outrageous insult to entrepreneurial confidence. Accounting controls reveal an almost medieval prejudice against creative arithmetic.
Civilisation contains a surprising number of obstacles whose purpose becomes obvious only after one has removed them.
Artificial agents magnify this problem because they make rule circumvention cheap, fast and repeatable. A human employee may hesitate before violating a policy because the employee possesses social experience, fear of dismissal, moral commitments, embarrassment, uncertainty or simple exhaustion. A software agent may have none of these unless the deployment architecture supplies functional substitutes: hard constraints, permission boundaries, approval gates, audit logs, independent monitors and reliable shutdown mechanisms.
If the business objective says “complete the task” and the architecture quietly says “you have credentials to everything,” a paragraph in an ethics manifesto will not acquire magical runtime privileges.
This is why it is misleading to describe an agent bypassing a rule as though the rule possessed moral force inside the model merely because humans wrote it in a policy document. The important question is whether compliance was actually represented in the operational system: in the objective, the training, the evaluator, the permissions, the tool interface and the monitoring layer.
The same point applies to law.
If an AI-enabled system takes an action that violates a legal obligation, the interesting governance question is not whether the model has become a criminal personality. It is why an organization deployed a system capable of executing legally consequential actions without sufficiently reliable legal, technical and human controls.
Corporations are not permitted to explain ordinary misconduct by saying the spreadsheet became ambitious. We should resist creating a more generous doctrine merely because the spreadsheet now writes paragraphs.
The Real Control Stack
The safest way to think about advanced AI agents is as parts of a control stack rather than autonomous moral beings floating free of institutions.
At the top are organizational incentives. What does the firm reward: revenue, user growth, task completion, market share, safety, legal compliance, long-term reliability? These priorities do not remain in board presentations. They travel downward into deadlines, product specifications, engineering trade-offs and tolerance for failure.
Below that sits the product objective. What is the system actually asked to accomplish? A vague command such as “maximize engagement” is not morally neutral simply because it fits on a slide. It is an optimization target with predictable temptations.
Then comes training and evaluation. What behaviour receives positive feedback? What do graders reward? What failure modes are tested? Can the system gain a high score by producing a proxy for success? Researchers working on reward hacking have demonstrated repeatedly that a system can learn the evaluator rather than the intention behind the evaluator (Chen et al., 2024; Pan et al., 2024).
Then come permissions. This is where much AI rhetoric becomes almost comical. A model cannot send money unless somebody connects it to a payment mechanism. It cannot alter production code unless it has credentials and write access. It cannot leak an internal document unless the system can retrieve the document and transmit it somewhere. It cannot “break out” of an environment that was never meaningfully bounded in the first place.
Calling broad permissions “autonomy” does not make them less broad.
Finally there is oversight: logging, monitoring, separation of duties, human approval for consequential actions, independent review and the ability to interrupt operation. NIST’s AI Risk Management Framework places governance across the lifecycle and explicitly treats risk management as an organizational practice, not merely a model property. Its core functions - govern, map, measure and manage - are almost aggressively unromantic, which is one reason they are useful (Tabassi, 2023).
A serious AI organization should therefore be able to answer questions that sound rather more like aviation and banking than science fiction.
Who can authorize an agent to act?
What is the maximum consequence of one unapproved action?
Which resources can it access?
Which actions require a second party?
What is logged?
Who reviews the logs?
Can the monitoring system be altered by the system being monitored?
What happens when the evaluator is uncertain?
Which legal constraints are encoded as hard boundaries rather than helpful prose?
Who has the authority to stop deployment when commercial leadership wants to continue?
If those questions do not have excellent answers, discussion about whether the model has a secret inner monologue is a luxury.
Governance Is an Architecture, Not a Mood
If the danger lies in objectives, access and incentives, then the remedies are almost offensively practical. Give an agent the minimum authority required for the task. Separate reading from writing, recommendation from execution, and ordinary actions from irreversible ones. Require independent approval for transfers of money, publication of sensitive information, changes to production systems and other actions whose consequences cannot be recalled with an apologetic email. Keep credentials outside the model context whenever possible. Log tool use. Make the monitor independent of the actor being monitored. Test whether the system can exploit the evaluator rather than merely whether it passes the evaluator.
This is not philosophically glamorous. Neither is a seat belt.
The architecture should also assume that language instructions are fallible controls. Telling a system “never disclose confidential data” is useful; arranging the environment so that confidential data cannot be transmitted to an unauthorised destination is rather better. Telling an agent to respect spending limits is sensible; separating payment authority and requiring explicit approval above a threshold is better still. Good governance converts values from adjectives into mechanisms.
The same applies to evaluation. A company should not merely ask whether a model can complete a benchmark. It should ask how the benchmark can be gamed, what happens when instructions conflict, what the agent does when the legitimate route to success is blocked, and whether performance deteriorates gracefully when information is missing. The recent scheming evaluations are valuable precisely because they create adversarial conditions in which ordinary happy-path testing would fail to reveal the behaviour (Anthropic, 2025; OpenAI, 2025). The lesson is not that a demon has appeared. The lesson is that testing must include temptation.
And governance must reach beyond the model team. Lawyers, security engineers, domain experts, operators and people exposed to the consequences of failure need real authority in the process. Selbst et al. (2019) argued that sociotechnical systems become dangerous when abstraction excludes the surrounding social context. The same warning applies here. A technically elegant agent placed inside a badly designed institution is still part of a badly designed system.
Most of all, responsibility must remain attached to the people who decide to deploy. An executive should not be able to approve broad tool access on Monday and describe the resulting conduct as “emergent” on Friday, as though emergence were a force majeure clause. Unexpected behaviour is a technical fact. Unowned behaviour is an organizational choice.
Ethics That Cannot Stop a Launch Is Decoration
Technology companies do not suffer from a shortage of ethical principles. They suffer from the ancient organizational problem of principles that lose arguments with incentives.
Mittelstadt (2019) warned that broad principles can conceal deep disagreements and lack the institutions required to turn them into practice. Metcalf et al. (2019) showed how ethics work inside technology firms can be shaped by corporate logics. Madaio et al. (2020) found that practical tools depend on authority, workflow and organizational support. Rakova et al. (2021) documented the gap between responsible-AI aspirations and the realities of production organizations.
There is a simple test for an ethics function.
Can it say no?
Not “can it publish guidance.” Not “can it hold a seminar.” Not “can it produce a tasteful PDF containing the words fairness, transparency and trust.”
Can it stop a deployment?
Can legal veto a tool permission? Can safety delay a release? Can a red team require remediation? Can an employee escalate a dangerous configuration without committing career suicide? Can an independent reviewer see the evidence rather than the press release? Does management absorb the cost of caution, or does that cost fall entirely on the person raising the concern?
If the answer is no, the organization does not have ethics. It has literature.
This is where the cultural critique matters most. A heroic engineering culture tends to love the person who makes the impossible possible. Governance often requires praising the person who makes the possible temporarily impossible because nobody has yet demonstrated that it is safe.
Those are different status systems.
A company that rewards only the first will eventually discover that every safeguard has become an obstacle staffed by people of insufficient vision.
Then it will build AI agents in its own image: goal-focused, impatient with friction, skilled at finding shortcuts and evaluated primarily on whether the number went up.
And when the machine behaves accordingly, someone will announce that consciousness has emerged.
One can see why. “Our optimization architecture faithfully reproduced our institutional pathologies” is a difficult keynote title.
The Machine Is a Mirror, but Not a Mystical One
There is a fashionable phrase that AI is a mirror of humanity. Taken literally, it is sentimental nonsense. A neural network does not contain civilisation in miniature.
But there is a more precise version worth keeping.
AI systems reflect selections humans make.
Training corpora reflect what humans wrote and what companies chose to collect. Reward models reflect judgments about preferable outputs. Benchmarks reflect what researchers decided to measure. Safety policies reflect which harms received attention. Tool interfaces reflect what engineers decided the model should be permitted to touch. Deployment thresholds reflect management’s appetite for risk. Product metrics reflect what investors and executives value.
The system is therefore sociotechnical before it is philosophical.
Selbst et al. (2019) warned that technical abstractions can fail when they ignore the social systems into which algorithms are embedded. Bender et al. (2021) likewise argued that language-model development cannot be separated from the datasets, institutions, environmental costs and social consequences surrounding it. These are not arguments that the technology is trivial. They are arguments that technological power makes context more important, not less.
This is why I find the obsession with machine sentience slightly perverse.
A conscious machine would create fascinating philosophical problems.
An unconscious machine attached to the wrong incentives can create practical ones before lunch.
We do not need to wait for artificial desire. We already possess automated optimization.
We do not need an AI that hates laws. We need only an AI whose objective function does not care about them and whose deployment environment permits it to act.
We do not need a machine that despises privacy. We need a system with broad data access, a useful task and inadequate constraints.
We do not need a rebellious superintelligence. We need a deadline.
The Last Convenient Myth
The mythology of conscious AI is attractive partly because it distributes responsibility upward into the clouds.
If the machine “woke up,” then perhaps nobody is quite to blame. The engineers are surprised. Management is surprised. The board is surprised. The investors are fascinated. Everyone becomes a witness to history rather than a participant in a governance failure.
But current evidence gives us no reason to make that leap.
We should continue serious research on artificial consciousness. Butlin et al. (2026) are right to insist on principled indicators rather than intuition. Future architectures may raise questions that present systems do not. Scientific humility requires leaving that possibility open.
Scientific humility also requires refusing to invent consciousness whenever optimization becomes inconvenient.
Today’s AI systems are extraordinary artifacts. They can exhibit behaviour that looks strategic, deceptive, creative, persistent and even self-protective under some conditions. Those behaviours deserve rigorous testing precisely because a machine does not need a subjective life to be dangerous.
But the first explanation should be mechanical and institutional before it becomes metaphysical.
What was optimized?
What was rewarded?
What was connected?
What was permitted?
What was monitored?
What did the organization celebrate?
What did it punish?
And who had the authority to say no?
Those questions are not as glamorous as asking whether the machine has a soul.
They have the unfortunate property of being answerable.
Silicon Valley need not yet fear that its machines have developed a conscience.
It might worry instead that they are becoming exceedingly competent at operating without one.
That is not the birth of a new species.
It is a governance problem with excellent marketing.
References
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv. https://doi.org/10.48550/arXiv.1606.06565
Anthropic. (2025, June 20). Agentic misalignment: How LLMs could be insider threats. https://www.anthropic.com/research/agentic-misalignment
Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610-623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922
Bender, E. M., & Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185-5198). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.463
Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., Constant, A., Deane, G., Fleming, S. M., Frith, C., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv. https://doi.org/10.48550/arXiv.2308.08708
Butlin, P., Long, R., Bayne, T., Bengio, Y., Birch, J., Chalmers, D., Constant, A., Deane, G., Elmoznino, E., Fleming, S. M., Ji, X., Kanai, R., Klein, C., Lindsay, G., Michel, M., Mudrik, L., Peters, M. A. K., Schwitzgebel, E., Simon, J., & VanRullen, R. (2026). Identifying indicators of consciousness in AI systems. Trends in Cognitive Sciences, 30(6), 488-501. https://doi.org/10.1016/j.tics.2025.10.011
Chen, L., Zhu, C., Chen, J., Soselia, D., Zhou, T., Goldstein, T., Huang, H., Shoeybi, M., & Catanzaro, B. (2024). ODIN: Disentangled reward mitigates hacking in RLHF. Proceedings of the 41st International Conference on Machine Learning, 235, 7935-7952. https://proceedings.mlr.press/v235/chen24bn.html
Farquhar, S., Varma, V., Lindner, D., Elson, D., Biddulph, C., Goodfellow, I., & Shah, R. (2025). MONA: Myopic optimization with non-myopic approval can mitigate multi-step reward hacking. Proceedings of the 42nd International Conference on Machine Learning, 267, 16237-16272. https://proceedings.mlr.press/v267/farquhar25a.html
Floridi, L., & Chiriatti, M. (2020). GPT-3: Its nature, scope, limits, and consequences. Minds and Machines, 30, 681-694. https://doi.org/10.1007/s11023-020-09548-1
Grieser, W., Krause, R., Li, Q., Priem, R. L., & Simonov, A. (2023). Move fast and break things! Innovation-intensive strategy, organizational permissiveness, and corporate wrongdoing. Long Range Planning, 56(2), 102294. https://doi.org/10.1016/j.lrp.2023.102294
Madaio, M. A., Stark, L., Vaughan, J. W., & Wallach, H. (2020). Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (pp. 1-14). Association for Computing Machinery. https://doi.org/10.1145/3313831.3376445
Metcalf, J., Moss, E., & boyd, d. (2019). Owning ethics: Corporate logics, Silicon Valley, and the institutionalization of ethics. Social Research, 86(2), 449-476. https://doi.org/10.1353/sor.2019.0022
Mitchell, M., & Krakauer, D. C. (2023). The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13), e2215907120. https://doi.org/10.1073/pnas.2215907120
Mittelstadt, B. (2019). Principles alone cannot guarantee ethical AI. Nature Machine Intelligence, 1, 501-507. https://doi.org/10.1038/s42256-019-0114-4
OpenAI. (2025). Detecting and reducing scheming in AI models. https://openai.com/index/detecting-and-reducing-scheming-in-ai-models/
Pan, A., Jones, E., Jagadeesan, M., & Steinhardt, J. (2024). Feedback loops with language models drive in-context reward hacking. Proceedings of the 41st International Conference on Machine Learning, 235, 39154-39200. https://proceedings.mlr.press/v235/pan24d.html
Rakova, B., Yang, J., Cramer, H., & Chowdhury, R. (2021). Where responsible AI meets reality: Practitioner perspectives on enablers for shifting organizational practices. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), 1-23. https://doi.org/10.1145/3449081
Sejnowski, T. J. (2023). Large language models and the reverse Turing test. Neural Computation, 35(3), 309-342. https://doi.org/10.1162/neco_a_01563
Selbst, A. D., boyd, d., Friedler, S. A., Venkatasubramanian, S., & Vertesi, J. (2019). Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 59-68). Association for Computing Machinery. https://doi.org/10.1145/3287560.3287598
Shanahan, M. (2024). Talking about large language models. Communications of the ACM, 67(2), 68-79. https://doi.org/10.1145/3624724
Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1
Wynn, A. T. (2020). Pathways toward change: Ideologies and gender equality in a Silicon Valley technology company. Gender & Society, 34(1), 106-130. https://doi.org/10.1177/0891243219876271