Back to the blog

The model that doesn’t write, only decides

Jev, Laya, Clef — and what a Hungarian document processor running on its own infrastructure can do with them.

How many times a day do we ask a 35-billion-parameter language model to answer with a single word? Incoming or outgoing. Invoice or delivery note. Do we call the invoice search, or query the receivables?

Even then, the model reads the whole input and produces a JSON object token by token. We parse it, validate it, and in the end use it in the condition of an if. Often, the entire process exists to make one decision.

Since mid-September there has been another way to do this. In barely three weeks, at least three models have appeared that generate no text at all. They receive the current state and a set of predefined questions, and assign a probability to every allowed answer. These are decision models. One of their developers calls them System One models, after Kahneman’s notion of fast, reflexive thinking.

We have already started research in this direction. We have no results of our own to show yet, but we can explain why we took up the topic. In this article we walk through how these models work, what their developers promise, and what limitations the model cards report. Along the way we show where they might fit in a Hungarian document processor running on its own hardware.

One word, 35 billion parameters

A document’s path through DocAI is accompanied by many small decisions. What kind of document has arrived? Did we issue it, or is it addressed to us? Which tool should the chat call to answer the user’s question? Each of these is a choice among options known in advance.

Today, document types are classified by an XGBoost model and a logistic regression model running side by side. They work on a 7,296-dimensional vector concatenated from seven embeddings: the summary, the title, the header, the footer, the topics and a structural fingerprint, plus an image vector.

Classification is backed by selective prediction: if the system is not confident enough in its answer, the document does not get a type automatically. It goes to the “Unidentified” filter, where a person reviews it.

This setup got us to roughly 95% accuracy. The remaining errors, however, are not evenly spread. The most damaging mistakes happen on the boundary between two closely related document types, and the model is often confident when it makes them. The document passes the check, bypasses the “Unidentified” filter, and quietly ends up in the wrong place.

That is why, for us, catching silent misclassifications matters more than raising the share of documents classified automatically. A diagnostic probe showed that the information needed to tell the types apart is already present in the existing features. The current models just don’t make good enough use of it.

We have run into a similar problem with generative models.

Valid JSON, wrong decision

When we hand a decision to a generative model, we tend to worry about the output format. Will it stick to the schema? Will it invent a new field? Will it return the right data type?

In our measurement with Llama-3.3-70B, none of this was a problem. As we showed in “When the eval lies”, all 119 outputs were valid JSON. Not a single format error.

Meanwhile, the model confidently wrote a wrong value into 108 fields. The mistakes clustered precisely on identifying the counterparty and on deciding whose document it was. In other words, on the two groups of fields that all further processing depends on.

This is worth keeping in mind when looking at what decision models promise. For us, the expensive error was the plausible, well-formed, yet wrong answer. An answer with nothing about it to suggest we should doubt it.

Three weeks, three decision models

Two processing paths side by side for the same question: what type of document has arrived. On the left, the generative model processes the input and then writes JSON token by token, from which the program reads the decision after validation; a valid format does not mean a correct answer. On the right, the decision model scores the options it is given (in the example invoice 0.92, delivery note 0.06, none of these 0.02), and then, based on calibrated confidence and a decision threshold, the document goes either to automatic processing or to further, human review. Illustrative figure; the scores are not measurement results.

Jev: decision rules inside the code

TypeSafe AI’s model, Jev, was released on 15 September 2026 in early access. It comes from the workshop of Diogo Almeida, who previously worked at OpenAI on the instruction-following methods behind chat models.

According to the announcement, Jev uses a new architecture, parallel sampling and its own training procedure, called Reinforcement Learning for Calibrated Decisions, or RLCD. The caller specifies the possible outputs in advance, and the model assigns each of them a calibrated probability and a confidence.

It supports three question types: a choice among up to 255 options (choice), a rating on an ordinal scale (score), and a yes/no probability (noul). It costs $0.042 per million input tokens, and output is free. End-to-end response time, according to the published figures, is 70–500 ms.

TypeSafe positions these models as “smart if-statements”: replacements for decision rules where hand-written logic has become too brittle. Jev is available through a closed API; its weights have not been released.

Laya: a small encoder with a separate decision head

Convai Innovations’ Laya model family is available with open weights under the Apache 2.0 licence. It comes in three variants: a 421-million-parameter English model based on ModernBERT-large, a 322-million-parameter multilingual model based on mmBERT-base, and a 421-million-parameter variant fine-tuned for four specific workflows.

The multilingual model’s design is easy to follow. A decision head trained from scratch sits on top of the bidirectional encoder, with two transformer layers and a per-option scorer. Each answer option is scored at its own [MASK] token. A separate head decides whether the system should act or hand the task on.

This means the possible answers can be supplied afresh with every request, without retraining the model. Processing one question takes about 33 ms on a T4 GPU.

Clef: the language model reads, the decision head chooses

Cloudflare’s Clef and Clef-flash models were released on 1 October 2026, also under Apache 2.0, with an interface compatible with Jev’s API.

What interests us most is how they are built. Clef sits on a frozen Qwen3.8-27B, Clef-flash on a frozen Qwen3.5-9B. Cloudflare trained the decision head together with rank-256 LoRA adapters, then refined the calibration with a separate Brier loss.

At inference time, Qwen processes the input in a single prefill pass, just as it would before generating text. The decision head then scores the possible answers in parallel. No text generation is needed.

PropertyJevLaya (multilingual)Clef / Clef-flash
PublisherTypeSafe AIConvai InnovationsCloudflare
Released15 September 2026September 20261 October 2026
WeightsClosed, API onlyOpen, Apache 2.0Open, Apache 2.0
SizeNot disclosed in the announcement322M27B / 9B
BasisOwn architecturemmBERT-base + decision headFrozen Qwen + decision head + LoRA
Context32k1,024 (extendable to 8,192)64k
InputTextTextText and image

So the same task has produced three different engineering answers. TypeSafe built its own model, Convai added a decision head to a small encoder, and Cloudflare built on an existing language model while leaving its base weights untouched. It was this last approach that gave us the idea for our own research.

What the demo promises, and what the model card says

Presentations of decision models often claim that these models don’t hallucinate. TypeSafe spells out what it means by this: because the possible outputs are defined in advance, the model cannot invent a new option or return an answer of the wrong type. The way it works rules that out.

Yet of the 108 wrong fields in our Llama measurement, type safety would have caught not one. Every wrong value matched the expected type.

For us, then, the promise that matters is calibrated confidence: that the number attached to an answer gives a usable signal of uncertainty. That is something a threshold can be built on, below which the software asks for further checks.

How well this works has to be measured. The model cards help a lot here, because they are surprisingly detailed about the limitations too.

Calibration has to be done on your own data

According to its model card, the multilingual Laya model ships uncalibrated. It is systematically over-confident: its mean confidence sits between 0.75 and 0.83, while its accuracy is considerably lower.

Refitting the temperature parameter on separate held-out data, per question type and per number of options, reduces the expected calibration error (ECE) from 0.314 to 0.106. Usable probabilities, in other words, come with calibration on your own data.

Preparing the model for the task matters just as much. Zero-shot, without any fine-tuning, Laya reaches an accuracy of 0.342 on the typed-decisions task. Random guessing scores 0.318, and always picking the most frequent answer scores 0.461. The model card states plainly that the decision-making ability comes from fine-tuning. Laya is best seen as a fast base for specialisation.

Confidence can mask a language failure

As Hungarian developers, we found the section on the English variant particularly instructive. In other languages its performance can collapse while the model stays confident. In Khmer, for example, it shows 0.952 confidence at 0.000 accuracy. And its mean confidence never drops below 0.885 at any accuracy level examined.

A confidence threshold alone cannot filter this out. That is why Convai picks the language route before the model is even called, based on the script of the input. The model’s own certainty does not reveal when it cannot make sense of the text.

It is a familiar problem from the eval article we cited above: we get a number that looks meaningful and would be easy to build the next decision on. Except the number does not signal what we expect it to.

What they measure against matters, too

TypeSafe’s workflow tests have no hand-labelled reference. The expected answer is the average of the answers of two large language models, GPT-6 Astra and Fable 5.1.

TypeSafe points this out itself. It also notes that the 193.6× speed advantage and 444.6× cost advantage shown in the demo come from these workflows, and that by its own account they represent the upper end of the gains to expect. Cloudflare’s comparison table is likewise based on its own runs.

These numbers measure agreement with other models. Whether a Hungarian invoice gets the right decision is something we have to check separately.

Why we started looking into it

We see an opportunity in this model family for three reasons.

The first is selective prediction. Around our document-type classifier we built the dual check and the “Unidentified” filter ourselves. Decision models support this way of working out of the box: alongside the answer they return a probability and an escalation signal. If the calibration holds up, the threshold can also be set in business terms: what error rate do we accept among documents processed automatically?

The second is running locally. Our clients’ Hungarian documents do not leave our own infrastructure. That rules out Jev’s closed API, while the open-weight Laya and Clef remain candidates on this count.

The third is Clef’s design. That is what sparked our own idea: if a frozen Qwen with a decision head and LoRA adapters can handle this kind of task, do we really need a second, separately loaded model for it?

We already run a Qwen3.6-35B-A3B under vLLM on the Spark, processing our documents every day. We want to find out whether this same model can be extended with a decision head, at a small memory overhead, using LoRA adapters that can be swapped per task and per client. Each client’s documents are different; maintaining one adapter per client would take fewer resources than maintaining a whole model.

For now this is a research question. Several important details are still open.

Cloudflare built on dense models. Ours is a MoE, with roughly 3 billion active parameters per token. We don’t yet know whether the approach works as well with routing between experts.

Serving is a separate challenge. vLLM is optimised for generation, whereas the decision head works from the model’s internal states. We need to work out how it can access them at run time.

Nor do we want to guess the speed in advance. The bandwidth formula we presented in the eval article applied to decoding. A decision model that only does prefill has no such loop, so a different limit may dominate on the Spark. Which one, exactly, we will settle by measuring.

The groundwork for the measurement plan is in place. We will use three reference points: our current classifier with its dual check, Clef-flash in zero-shot mode, and the generative Qwen asked the same question.

The main metric will be selective accuracy at a given coverage. That is, if we let a set share of documents through automatically, how many wrong decisions are among them. We will also look at calibration measured on Hungarian documents, response time and memory overhead.

It may also turn out that a ready-made 9-billion-parameter decision model is the better choice, and there is no point in making the system more complex. We will write that up just the same.

Before you set a confidence threshold

Whether you plug Jev, Laya, Clef or a model of your own into a processing pipeline, a few checks are worth doing first. The points below come from the model cards and from our own errors that went unnoticed for a long time.

  1. Measure in your own language. Laya’s English variant can fail confidently in other languages. For Hungarian text, make it explicit which model variant does the work. We have not yet measured where Laya’s router sends accented Hungarian text.
  2. Compare confidence with actual accuracy. On your own labelled, held-out data, group the answers into confidence bands: for example between 0.8 and 0.9, and above 0.9. Check whether the share of correct answers in each band matches the stated confidence. If high confidence comes with markedly lower accuracy, the raw value cannot support a reliable threshold.
  3. Calibrate on your own data. According to Laya’s model card, refitting the temperature per question type and per number of options cut the calibration error to roughly a third. That is a small effort that can make the resulting probabilities far more usable.
  4. Check what the test compares against. If the reference is another model’s answer, the score shows how much the two models agree. Judging actual correctness also takes your own verified data.
  5. Make it possible to hand the decision on. Include a “none of these” option among the answers, and make it clear when a task goes to a person. Selective prediction is only usable if there is also a process for handling the uncertain cases.
  6. Check correctness, not just type safety. A well-formed, allowed answer can still be wrong.

Where we are now

This article reports no measurements of our own on decision models. The scores, speed and cost figures quoted all come from the developers’ own tests and runs.

We don’t yet know how these models perform on Hungarian business documents, so we are not ranking them. On our own development track, too, we have got as far as the question that needs answering. Whether the approach works, the upcoming measurements will show.

The next step

Wherever text has to be written, explained or summarised, the generative model still has a job to do. At other points in the pipeline, though, the task is to pick from a few known options, and the answer is followed straight away by an if. At those steps, decision models offer a more direct solution. The key question is how far we can trust their choice and the confidence that comes with it.

In the next article we measure Clef-flash on Hungarian documents. After that comes our own experiment: a decision head for the model we already have loaded. If the result is convincing, it will become a full research report. If not, we will write about that too — probably more briefly.


This article contains no measurements of our own on decision models. Jev figures come from TypeSafe’s announcement and API documentation, Laya figures from Convai Innovations’ Hugging Face model card, and Clef figures from Cloudflare’s announcement; we have not reproduced them. The Llama-3.3-70B figures (119 of 119 outputs valid JSON, 108 wrong field values) come from the measurement published in “When the eval lies”: 100 real Hungarian documents, 119 evaluation units. The figure is illustrative; the scores in it are not measurement results.