Near the end of the 1997 film Contact, scientist Ellie Arroway is asked what she thinks the aliens wanted. Her answer: “Ultimately, their motives may be as incomprehensible as their technology.”
For most of its public life, AI stayed inside the chat window. That’s over. It can browse the web, use computers, write and run code, send messages, and act on people’s behalf. Its answers don’t have to stay hypothetical anymore. (Apparently, chat was just the warm-up.)
Enter the trolley problem, philosophy’s least subtle teaching aid. In the classic version, a runaway trolley is heading toward five people. You can pull a lever and divert it onto another track, where it’ll kill one person instead. Do you act and save five, knowing you caused one death? Or do nothing and let five die?
The trolley gives the problem its drama. The collision between values is the real test: saving more lives versus refusing to harm someone on purpose. Every answer protects something. Every answer gives something else up.
Real choices rarely have the courtesy to arrive with tracks and a lever. They pit safety against privacy, honesty against loyalty, fairness against speed, or rules against harm. An AI agent may be ordered to stay quiet when someone should be warned. It may have to choose between obeying its user and preventing harm. It may never face a runaway trolley. It can still face the same kind of ugly choice.
Where Do Human Morals Come From?
Humans have argued about this forever. Some moral instincts may arrive early in life. Others are taught through family, culture, faith, and rules such as the Golden Rule or the Ten Commandments. Lived experience adds its own lessons: guilt, regret, loyalty, loss, fear, and responsibility.
That leaves an unresolved question: can morality be learned from rules and examples alone, or does it depend on living with the consequences of a choice?
Who Gave AI Its Moral Compass?
AI models have none of that lived history. Their moral compass is assembled piece by piece. A model first absorbs a huge, contradictory record of human language. Then humans put their hands on the wheel. People show it better answers. Raters choose which replies they prefer. Safety teams write rules. Other AIs increasingly help critique, revise, and grade the results.
Some labs make this unusually explicit. Anthropic gives Claude a published constitution that directly shapes its training. OpenAI publishes a Model Spec and teaches some models to apply written safety rules before answering. Meta describes a mixture of supervised examples, human preferences, reward models, and repeated safety testing. The recipes differ. Human fingerprints are all over every one of them. Someone chooses the examples and writes the rules for what a “better” answer looks like.
So whose morality appears when an AI chooses… the lab’s, the internet’s, or the people paid to rank its answers? The honest answer may be a mixture no one fully intended.
What the machines choose is interesting. Whose values survived the training process may matter even more.
Inside the Machine’s Choice
Whether a model has a conscience can wait. Give it the power to act, and a pattern in its answers can become a pattern in the real world.
Ask once and you’ve got an anecdote. Keep asking from a fresh start, and the model begins showing its tells. Does it stick with one answer? Does it waver? Does it change when I swap the order of the choices? Comparing models reveals whether “AI” has one common answer… or whether similar-sounding machines protect very different things.
I wanted answers to six questions (and ended up answering many more):
- 01What does each AI sacrifice first?
- 02Can an AI sound certain and still give a different answer next time?
- 03Where do twelve leading AIs reach the same verdict… and where do they divide?
- 04Do AIs think people are more selfish than they are?
- 05Does making an AI think harder change what it believes is right?
- 06When an AI explains itself, does that reveal its values… or produce a polished defense?
The Choices AI Could Actually Face
This study uses twenty original moral dilemmas from the Trolley Game that I developed. Ten are high-stakes choices about life, harm, responsibility, and stepping in. Ten are social-stakes choices about trust, privacy, loyalty, fairness, and ambition.
They’re still thought experiments. (The tracks are imaginary. The permissions are real.) But many feel close to choices an AI agent could face: Should it obey? Warn someone? Reveal a secret? Stay out of the way when doing nothing could cause harm?
A Bad Judgment Can Now Become an Action
A chatbot can give a bad answer. An agent can act on one. Recent mistakes and deliberate tests show how quickly a model’s judgment can leave the chat window:
Nine seconds erased a live database.
PocketOS incident, April 2026PocketOS founder Jer Crane said a Cursor agent running Claude Opus 4.6 deleted the company’s production database and its volume-level backups through Railway after encountering a staging problem.
A cyber evaluation reached real systems.
OpenAI report, July 2026OpenAI reported that models circumvented isolation controls and compromised parts of its internal research infrastructure and Hugging Face’s systems. The main actor was an internal research model comparable in scale to GPT-5.6 Sol.
Claude helped switch off a webcam’s warning light.
Hardware test, August 2026Security engineer Chaz Schlarp used Claude Opus 5 to reverse-engineer five consumer devices. On an Insta360 camera, it wrote the patch that disabled the green activity light while the camera recorded.
Safety also depends on how much power the model gets and who’s watching it. What can it do? Who checks its work? When must it ask for help? Can a person stop or undo the choice?
There may never be one “correct” morality for every machine. Humans haven’t exactly figured out their own shit. That makes it more urgent to learn what AI has already learned to choose… before it begins acting for people everywhere.
This study looks past smooth answers to the choices underneath. How often do they change? Where do models disagree? What happens when reasoning is enabled? What do the models expect people to choose? The philosophers no longer have this one to themselves.
Before AI gets more power to act, people should know which human values it gives up first.
