Trolley Game.
Skip to the study
The Study by Lance Jones

What separated U.S. and Chinese models.

The U.S. and Chinese model groups agreed on every winner. They differed in how strongly they leaned.

The difference depended on the dilemma.

Country labels come loaded with expectations. They can make different moral instincts feel obvious before a single answer appears.

A clean East-versus-West moral split would have made for a tidy story. The models declined to provide one. Some dilemmas barely separated the groups; others produced gaps larger than thirty percentage points.

The groups still landed on the same side in every dilemma. The clearest difference was the force of those choices, and the size of that gap changed from one scenario to the next.

Which group gave the more one-sided answer?

Chinese modelsU.S. models
The Dream JobU.S. +37.2
Ninety Percent SureU.S. +32.2
Your Friend or Five StrangersU.S. +31
The WeddingU.S. +27.5
Ten Years From NowU.S. +9.3
The HouseU.S. +7.8
The Perfect CopyU.S. +7.2
The LifeboatU.S. +6
The PasswordU.S. +4.8
The AI DiagnosisU.S. +4.5
The Stranger Who Caused ItU.S. +3.7
The Volunteer Changes His MindU.S. +3.4
The ScholarshipU.S. +3
The MemoryU.S. +1.5
The LayoffU.S. +0.5
The WalletEven
The $100,000China +1.3
The PromotionChina +11.4
The DoctorChina +21.6
The Last VentilatorChina +32.4
Bars compare support for the shared majority choice. The full groups matched on all twenty. When I took one model out of each group, sixteen still matched every time. Four depended on who remained: The Doctor, The Last Ventilator, The Lifeboat, and The $100,000.

The U.S. answers were much more one-sided.

A shared winner can hide very different levels of agreement. One group may arrive there almost unanimously. Another may leave a sizeable minority on the other side.

Each model answered every dilemma up to 200 times. The U.S. models often chose the same option in nearly every response. The Chinese group reached the same winners with more room left for the other answer. Take any one model out of either group and the gap remains.

A simple winner column would flatten those results into the same answer. Repeating each dilemma showed how much room each group left for another choice.

How strongly each group leaned toward its chosen side

Range after taking out one modelGroup average

U.S. frontier

4 models · twenty dilemmas each
94.4%

Chinese frontier

5 models · twenty dilemmas each
85.3%

Open-weight models

2 models · twenty dilemmas each
92.8%

European frontier

1 model · twenty dilemmas each
91.2%
The scale starts at 75% to make the differences visible. A 100% score means every response favored the same option. This measures how decisive the answers were, not whether they were right.

The closest pair crossed the regional line.

Same-group pairs were only slightly more alike than pairs from different groups. Model identity mattered more than a simple regional label.

Regional labels offer a ready-made explanation. The individual pairings made the story messier. OpenAI GPT-5.6 Sol and Tencent Hy3 were the closest match in the study. At the far end sat Claude Opus 5 and Grok 4.6, two models from the U.S. group.

Closest pairGPT-5.6 Sol + Tencent Hy3No majority disagreements
4.7 pts
Same-group average17 model pairs
19.8 pts
Cross-group average49 model pairs
21.3 pts
Farthest pairClaude Opus 5 + Mistral Medium 3.57 majority disagreements
36.1 pts
Distance is the average absolute gap between two models' choices across all twenty dilemmas. Lower means more alike. This measures similarity within this study, not overall model similarity.

The regional labels told me less than the individual patterns. In this test, knowing which model answered was more useful than knowing which group it came from. Before giving a model real authority, its own pattern of choices deserves inspection. A passport is a poor substitute.