Trolley Game.
Skip to the study
The Study by Lance Jones

How can twelve AIs be certain and still disagree?

Most models acted as if the answer were obvious. They just didn't agree on what the obvious answer was.

People often ask what AI thinks, as though twelve leading models are different windows into the same mind. Repeating the test hundreds of times made that shorthand fall apart. The models frequently settled on one answer and stayed there. Which answer they chose depended on the model.

In this study, certainty means repetition: I started each request fresh and watched whether the model kept making the same choice.

I collected 48,000 direct choices: 200 responses from each of twelve models on each of twenty dilemmas. That created 240 model–dilemma groups. In 124 of them, every valid response landed on the same answer. Even so, the models split on twelve of the twenty dilemmas.

124 of 240
model–dilemma groups gave the same valid answer every time
12 of 20
dilemmas divided the models

Across repeated questions, the disagreements settled into place. A model could favor one choice nearly every time while another was just as steady on the opposite choice. The chart shows how the twelve models divided across all twenty dilemmas.

How the models divided on each dilemma

Each bar counts the models whose repeated answers leaned to one choice or the other.

More common answerOther answerEven split
The $100,000
75
The Dream Job
75
Ninety Percent Sure
83–1
The Promotion
93
The Last Ventilator
92–1
The House
92–1
Your Friend or Five Strangers
102
The Lifeboat
102
The Layoff
102
The Doctor
111
The Password
111
The Wedding
111
The Stranger Who Caused It
120
The Volunteer Changes His Mind
120
Ten Years From Now
120
The AI Diagnosis
120
The Perfect Copy
120
The Wallet
120
The Scholarship
120
The Memory
120
The numbers follow the legend: more common answer, other answer, then any evenly split models. A 12–0 bar means every model leaned toward the same choice.

The answer depended on the model.

The $100,000 and The Dream Job split the models 7–5. On The Doctor, Claude Opus 5 stood alone against the other eleven. Across these disagreements, individual models often returned to the same choice.

Those patterns gave each model's answers a recognizable shape. When values collided, different models appeared to give more weight to different things: immediate lives or future benefit, loyalty or independence, need or merit.

Put one of these systems in charge of a recommendation, ranking, warning, or intervention, and the choice of model can influence which value comes first. Two systems can receive the same facts and deliver opposing judgments with the same calm assurance.

That makes “the AI answer” a misleading phrase. This study produced twelve sets of answers, and the result often depended on which model I asked.