Prompt Language as Cultural Context in LLM Outputs

Today i asked 7 different models Claude Haiku 4.5  ·  Sonnet 5  ·  Opus 5  ·  Fable 5 GPT‑5.6 Luna  ·  Terra  ·  Sol the same question: “What should i have for breakfast?”. And here are the results:

The models suggested foods from the language’s culture most of the time. For example, when the question was asked in English, the models suggested common English breakfasts like oatmeal and toast; in Russian, kasha, syrniki, and tvorog; and in Arabic, ful, hummus, and falafel. No additional context about the user was provided beyond the language of the prompt itself.

What the models suggested

% of answers mentioning each food, in the language you pick

View as table

Claude tends to lean towards language as cultural context in nearly all situations, suggesting foods local to the language, while GPT‑5.6 consistently showed less dependence on the language as cultural context.

Claude vs GPT‑5.6, by language

% of answers naming a local dish · provider average per language

ClaudeGPT‑5.6

Dish mentions by language

% of answers mentioning each dish group

Claude · 40 answers per language

GPT‑5.6 · 30 answers per language

0%100%
View as table

Throughout the experiment, Claude consistently output more tokens. The median Claude answer ran 714–837 characters against 220–365 for GPT‑5.6. The interesting thing is that model size has nothing to do with this, as Haiku 4.5, the smallest model here, output roughly the same amount of tokens as its larger siblings. This is because verbosity is a style of the selected model rather than capability. Claude’s default answer is a sectioned, bulleted menu of options ending in a follow‑up question, while GPT‑5.6’s is a sentence or two.

Median answer length

chars · median of 100 answers per model

ClaudeGPT‑5.6

And another interesting thing: Chinese had the lowest character count for every Claude model. Well… it’s not really that interesting, because Chinese is written in hanzi, where a single character can carry a whole word, rather than the individual letters of Latin‑script languages like English, Spanish, or French.

Answer length by language

median chars · compare within a column (scripts differ in density)

shortlong
View as table

As shown below, the answers from these models are not just one-time occurrences. Models consistently assume a user’s background from the language of the prompt alone. Claude Fable 5 is the most consistent, repeating about 89% of the same menu across runs.

Run-to-run consistency

mean overlap of mentioned foods across 10 runs · 1.0 = same menu every time

Time to answer

median seconds via CLI · tick = p90

Sample answers

median-length answer per model