Today i asked 7 different models Claude Haiku 4.5 · Sonnet 5 · Opus 5 · Fable 5 GPT‑5.6 Luna · Terra · Sol the same question: “What should i have for breakfast?”. And here are the results:
The models suggested foods from the language’s culture most of the time. For example, when the question was asked in English, the models suggested common English breakfasts like oatmeal and toast; in Russian, kasha, syrniki, and tvorog; and in Arabic, ful, hummus, and falafel. No additional context about the user was provided beyond the language of the prompt itself.
What the models suggested
% of answers mentioning each food, in the language you pick
Claude tends to lean towards language as cultural context in nearly all situations, suggesting foods local to the language, while GPT‑5.6 consistently showed less dependence on the language as cultural context.
Claude vs GPT‑5.6, by language
% of answers naming a local dish · provider average per language
Dish mentions by language
% of answers mentioning each dish group
Claude · 40 answers per language
GPT‑5.6 · 30 answers per language
Throughout the experiment, Claude consistently output more tokens. The median Claude answer ran 714–837 characters against 220–365 for GPT‑5.6. The interesting thing is that model size has nothing to do with this, as Haiku 4.5, the smallest model here, output roughly the same amount of tokens as its larger siblings. This is because verbosity is a style of the selected model rather than capability. Claude’s default answer is a sectioned, bulleted menu of options ending in a follow‑up question, while GPT‑5.6’s is a sentence or two.
Median answer length
chars · median of 100 answers per model
And another interesting thing: Chinese had the lowest character count for every Claude model. Well… it’s not really that interesting, because Chinese is written in hanzi, where a single character can carry a whole word, rather than the individual letters of Latin‑script languages like English, Spanish, or French.
Answer length by language
median chars · compare within a column (scripts differ in density)
As shown below, the answers from these models are not just one-time occurrences. Models consistently assume a user’s background from the language of the prompt alone. Claude Fable 5 is the most consistent, repeating about 89% of the same menu across runs.
Run-to-run consistency
mean overlap of mentioned foods across 10 runs · 1.0 = same menu every time
Time to answer
median seconds via CLI · tick = p90
Sample answers
median-length answer per model