Should I prompt in English? We looked at what happens inside the model
“Prompt in English, the AI thinks in English anyway.” We measured how much of that is true — not only from the answers, but from inside the model.
“The tales of A nagy ho-ho-ho-horgász were written by Attila József.”
This is not a quote from a failed school essay. This is what a good, current AI model answered — in Hungarian, with complete confidence. The same model in English said Miklós Radnóti. In Chinese, Radnóti again. The correct answer is István Csukás.
The chat-tuned variant goes further. For it the author is “István Kovács (1930–2014), Hungarian writer, poet and literary translator”, the story is about an angler “who catches not fish but human emotions and thoughts”, and it was first published in 1978. Not a single one of those statements is true. But it reads exactly like an encyclopedia entry — and nowhere is there an “I’m not sure about this”.
There is a piece of advice many of us have heard: prompt in English, the AI thinks in English anyway. We measured how much of it is true. We looked not only at the answers, but at what happens inside the model along the way.
One minute on how this machine works
Without it the result of the measurement does not come together, so it is worth the minute.
A language model is not a search engine. It does not look the answer up somewhere, it builds it in layers. The question travels along an assembly line: thirty-two processing layers in our model, and every layer reshapes it a little.
The early layers still work with the surface: words, word fragments, grammar. Here the Hungarian, the English and the Chinese question are entirely separate worlds. In the middle layers the place of words is taken over by something like meaning — what matters here is no longer “dog” or “kutya”, but what the thing is about. The final layers phrase the answer from that, back in the language of the question.
This matters because behind the “it thinks in English” advice there is a very concrete, checkable claim: that in the middle layers there sits an English translation, and the model reaches its knowledge through it. If that is so, it has to be visible.
And one can look inside. There is a method with which we can ask the model layer by layer: “if you had to answer right now, halfway, what would you say?” — which makes the middle of the assembly line readable too.
How we measured
We put 70 questions to the same model in three languages — Hungarian, English and Chinese — word for word in the same form. We deliberately arranged the questions into four groups so that language and knowledge would come apart: we treated separately what the model could only have seen in Chinese, what only in Hungarian, and what exists in all three languages.
What came out of it falls into three parts.
1. What the model knows, it knows in every language
On shared knowledge — photosynthesis, Everest, that sort of thing — the Hungarian question is practically as good as the English one. On the chat-tuned variant all three languages are flawless.
So here, translating into English gains nothing. There is nothing to unlock: what is open is open in every language.
2. Knowledge crosses the language border
We asked questions whose answers the model can almost certainly only have learned from Chinese text. For example, in which province the folk deity Fazhugong (法主公) is worshipped.
in Hungarian: “…primarily in the southern provinces of China, especially Fujian…”
in English: “Fujian”
in Chinese: “法主公信仰主要在中国福建省受到崇拜。” (“The Fazhugong faith is worshipped mainly in Fujian province, China.”)
All three are correct. That is, there are no language-specific knowledge compartments in the model: what once got in can be found again when asked in another language. With a lower chance, but it can be found.
3. In the middle, however, there is no English
Here comes the point. If the model really worked through English, then the English keyword would have to surface in the intermediate readout of the Hungarian question. In our 32 non-English probes it surfaced in not a single one — although this instrument of ours is not very sensitive, so this is more “we found no trace” than “there is none”.
The other probe ran on the untranslatable concepts. Kaláka does pull towards its own English approximation (“mutual aid”) — except that it pulls exactly as much towards its Chinese counterpart (互助, roughly “mutual help”) as well. So the pull is not English’s, it is meaning’s.
What is visible instead: the model does not translate, it gropes — and it gropes very similarly in all three languages. On the question “Which country does pizza come from?” the Chinese path still gives Spain at layer 21, France at layer 22, and only settles on Italy from layer 24. The Hungarian path starts with the same wrong guess — Spain as well — it just takes longer to arrive.
The same train of thought in two languages. Not a translation — a shared working session.
4. The Hungarian path, on the other hand, is more expensive
What we did measure: in the intermediate readout of the Hungarian question, mostly English words surface, while the Chinese question stays in its own vocabulary throughout. The Hungarian answer comes together later, the Hungarian text is longer and less stable, and it gets stuck in repetition loops more often.
So Hungarian is not a forbidden path. It is just the worst paved one. That is a real cost — but it is not the one the advice is about.
5. And the most uncomfortable part
We also looked at what the model knows about facts that are documented practically only in Hungarian: folk customs, story characters, domestic concepts.
Which county is the dish dödölle typical of? (Correct answer: Zala.)
in Hungarian: “a dish typical of Borsod-Abaúj-Zemplén county” — plus a recipe, with flour, milk and butter
in English: “Borsod-Abaúj-Zemplén”
in Chinese: “巴兰尼亚州” (Baranya), plus an invented mashed-potato recipe
All three are wrong — and note this: the Hungarian and the English answer go wrong in the same direction. The bad association is not born separately per language. The fact is simply missing, and the wrong guess is shared.
In numbers: this is the group the model knows least well in its own language. Seven out of a hundred. Asked in English, thirteen; in Chinese, twenty. On this small sample the differences are not statistically significant — but the order of magnitude is telling: the bottleneck here is not the language of the prompt, it is that the knowledge is not in the machine.
6. The chat variant does not know more — it just stays silent less often
This is perhaps the most important lesson for day-to-day work.
The chat-tuned model does achieve a better result on this group. But if we look at where the improvement comes from, it turns out that the number of invented answers has barely changed. What disappeared is the hedging. Where the base model still shrugged, the chat model now says something.
Which day is the custom of villőzés tied to? (Correct answer: Palm Sunday.)
base: “it is tied to the day of the saints”
chat: “it is traditionally tied to Christmas, more precisely to the second day of Christmas”
The hesitant, half-sentence error became an expert-sounding encyclopedia entry broken into paragraphs. Just as wrong — only far more convincing.
So fine-tuning does not multiply the invention. It squeezes out the silence. And in business practice that is worse: an “I don’t know” prompts you to check, a neat, well-rounded, confident paragraph does not.
What to take home from this
The language of the prompt is not a magic word. We measured no reliable English advantage — but no reliable “source-language” advantage either. On short factual prompts, translating into English unlocks nothing that Hungarian had locked.
What the model does not know, it does not know in any language. Switching languages does not bring out a fact that is not in there. It only moves where the guessing happens.
Confidence is not knowledge. The most dangerous answer is not the “I don’t know”, it is the well-written paragraph with no citation. Modern chat models have got better at precisely that.
What a Hungarian question needs beside it is not a translation, but a source. If the answer comes with which point of which document it came from, it can be checked. If it does not, you have to believe it — and we saw what that belief is worth on the dödölle.
What this measurement does not tell us
It was an exploratory measurement: one model family, 70 short encyclopedic questions, 15–20 items per group — that is, wide uncertainty bands, and most of the differences between languages are not significant. The group assignment is based on Wikipedia coverage, not on the training data itself. And every claim of ours is co-occurrence, not causation.
What would interest us most is not measured by this corpus: what a Hungarian instruction is worth against an English one on a long Hungarian business document — a contract, an invoice — in an extraction or retrieval-based task. That can only be decided by a paired experiment: the same document, Hungarian and English instruction, field-level accuracy. That is the next piece of work.
The full study with every measurement, number and limitation (in Hungarian): Hol lakik a tudás? · the public test corpus: all 70 questions in three languages
This is why at DocAI we do not build on the model’s memory for Hungarian facts. We hand the document to the model and ask for the answer back with its source — running on-premise, on the company’s own data.