Should we fine-tune our own model? We measured it against a cheap alternative
One desktop machine, 485 poems by János Arany, 64 minutes of training — and the question every company asks: is fine-tuning worth it, or is a good prompt enough?
This stanza was written by a machine — an open language model we fine-tuned on the poems of János Arany, the 19th-century Hungarian poet:
Két árva fiú hányni kezd a sárral,
Hozzájuk egy lyány, az is hány, a bárcsal,
Kétszámra gyűlnek, három, de már négy is:
„Ki a pénzt? ki a zsírt? ki a nyers ebgyísz?”
The form is right: a four-line stanza, eleven syllables per line, rhyming couplets. The voice is there too — the pulse of a narrative folk poem. Except that “bárcsal”, “ebgyísz” and “pénzgyísz” are not words. The machine invented them, all at line ends, all to make the rhyme work.
This article is not about poetry. It is about when it is worth fine-tuning a model, and when a cheaper approach is enough — and verse is a good testbed only because “good” here is measurable with a ruler. In Hungarian a syllable count is a vowel count. A rhyme either exists or it doesn’t. You don’t need literary taste, just a well-written program.
The question that comes up at every company
When a model doesn’t do what you want, you have two levers:
- Train it (fine-tuning, LoRA): nudge the model with your own examples. This is the fashionable answer. GPU time, data preparation, maintenance — and all of it again at every model change.
- Filter its output: ask the model several times and let a deterministic program pick the best answer. No training, no GPU hours, and it still works tomorrow with a different model.
The tempting measurement is to compare the fine-tuned model against the raw one. Fine-tuning always wins that comparison spectacularly — and what gets left out is that you could have had the same thing without any training.
That is why you need a cheap opponent. Don’t compare against the raw model, compare against the best result available without training. If your measurement has no such baseline, you did not measure whether it was worth it — only whether something happened.
Our cheap opponent looked like this: the same raw model, a few examples in the prompt, and a few dozen lines of code that pick the most regular output out of eight attempts. Like taking eight photos of your child and framing one.
On form, the cheap option won
On every formal axis: stanza count, line length, syllable count, rhyme rate. And it produced a third as many invented words along the way.
| what we measure | raw + examples | + selection | fine-tuned |
|---|---|---|---|
| syllable count exact | 0.206 | 0.409 | 0.203 |
| stanza size kept | 0.802 | 0.938 | 0.735 |
| rhyme rate | 0.114 | 0.236 | 0.177 |
| invented words (lower is better) | 0.019 | 0.013 | 0.044 |
That is not surprising once you look at the mechanism: the selector maximises exactly what we measure. Fine-tuning only shifts the model’s tendencies; there is no guarantee that any single output will be regular.
So a syllable counter and a rhyme function buy the form more cheaply than 64 minutes of GPU time. If all you want is well-formed output, fine-tuning is a bad deal.
The turn: voice cannot be bought
We measured style with an independent program that never saw the generated texts: does it recognise what the machine wrote as Arany? (On real poems it is 91% accurate; random guessing sits at 59%.)
Look at the second and third bars: the selector does not move style at all (56.4% → 56.0%). That is not an error and not noise. Our program counts syllables and matches line endings; it has no idea whether a text resembles Arany. A selector can only pick what it can measure.
The cleanest comparison is the last two bars: same selector, same 50 tasks, same eight attempts. The only difference is the fine-tuning.
56% → 92%: +36 percentage points, purely from fine-tuning.
This chart is the article in one picture. The horizontal arrow is selection: it improves form, not voice. The vertical one is fine-tuning: it lifts voice, not form. And the two are not alternatives — the best result uses both, and it even runs cheaper, because the fine-tuned model doesn’t need a fifteen-times-longer instruction before every task.
Two traps we nearly fell into
1. The textbook metric picked the wrong model. Training has a standard indicator that tells you when to stop. In our case it recommended the one-epoch checkpoint. We measured the other one too, and the “overfitted” version was better on every axis: rhyme rate, syllable count, repetition, recognisability. The reason is mechanical: the indicator measures how well the model predicts the next word, while what we wanted was the manner. If your indicator doesn’t measure the task, don’t decide on it. A 50-task measurement was enough to get this right, and it cost six minutes.
2. The metric also rewards cheating. “Bárcsal” rhymes perfectly, so the rhyme metric is delighted with it. The fine-tuned model produces three times as many non-existent words as the raw one (0.015 → 0.044). The model did not learn to rhyme; it learned that rhyme matters more than the word. Only a second metric pointing the opposite way caught this — if every number in your measurement points the same way, you are probably not looking at enough numbers.
What we feared, and what didn’t happen
The real corporate risk of fine-tuning isn’t quality, it is that the model memorises the training material and later reproduces it verbatim. You cannot “remove” that from the weights afterwards, only retrain.
We measured it: in no generation did the model return a verbatim stretch longer than eight words from the training poems, and the gap between trained and held-out material is effectively zero — smaller than the raw model’s. Memorisation did not occur in this experiment.
With two caveats. This holds for this corpus size, this training setup and this task type; it is not a general absolution. And that is exactly why the whole thing ran on public-domain material (the works of Arany and Petőfi are free to use) — with copyrighted text you manage the legal risk by choosing the corpus, not by measuring.
What to take away when you decide about a model
- First ask whether the goal is measurable. If you can write a rule for what a good output looks like (field format, mandatory fields, arithmetic consistency), a filter buys it more cheaply and more reliably than training.
- Train for what you cannot put into a formula. Tone, domain jargon, “the way we do it here” — a selector cannot pick that out by definition.
- Put a cheap opponent in your measurement. Raw model + good prompt + a simple post-filter. Without it, fine-tuning always wins on paper.
- Measure the measuring instrument before the model. That was the most boring two days of this work, and without it every number that followed would have been unbelievable.
- The two are not alternatives. The best result pulls both levers.
The full measurement — 10 chapters, from the corpus through the validation of the measuring instrument to the hardware pitfalls, with every number and every limitation: Mit ad hozzá egy LoRA a jó promptoláshoz? in the Research section (in Hungarian).
The measurement is fully reproducible: the corpus is public domain, the model is open-weight, and the measuring tool uses no external library. All 800 generated poems, the training curve, the validation of the instrument and the checksums of the source packages are public.