Llama.cpp Decision Models Answer In One Forward Pass

By Indie Kings | October 10, 2026

Updated October 10, 2026: llama.cpp server now supports decision models through the /v1/systemone endpoint, per the Hugging Face blog community article New in llama.cpp: Decision Models by Xuan-Son Nguyen and Victor Mustar of ggml-org dated October 2 2026. Send a state plus typed questions and get a probability for each option in a single forward pass instead of generated text.

llama.cpp decision models header

Image: llama.cpp decision models header. Credit: Hugging Face blog, ggml-org.

What a decision model does

Per HF/ggml-org, a decision model scores options you supply instead of generating text. A chat model needs one forward pass per token plus parsing to turn free text into a choice. A decision model reads the input once and returns an answer that is always one of your options with a probability attached. That design fits routing, moderation, agent step checks, and next action choice.

The Hugging Face blog community article New in llama.cpp: Decision Models by Xuan-Son Nguyen and Victor Mustar explains the contrast in plain terms. Chat models guess the next token over and over until they stop. Decision models score a closed set. The closed set keeps downstream code simple because parsing stays out of the path. This site covers llama.cpp local AI.

The same article ties the server format to the System One format from TypeSafe Jev. TypeSafe is named in prose here as the source of that format idea. Details for the llama.cpp side live in PR 29818. That pull request number is the pointer for readers who want implementation history.

Probability output is central to the whole pitch. Each option gets a number that reads as confidence. Code can route on the top score, hold for review under a cutoff, or log scores for later audit. Scores beat string matching because every run stays inside the allowed set. The model never invents a fourth category when you offered three.

Single forward pass behavior matters for latency. Chat style classification pays per output token and then still needs a parser. Decision style pays once for the state plus questions. That structure is why the article frames decision models as fast judges for pipelines. One read gives a full ballot.

Use cases stay practical in the source. Routing sends a ticket to billing, shipping, or technical support. Moderation marks content as allowed or blocked. Agent step checks ask whether a tool call succeeded or needs a retry. Next action choice picks among a small list of moves. Each case shares the same shape with a state and a question.

For local runners, the appeal is control plus repeatability. The option list lives in your request, not in a prompt that the model might paraphrase. Probabilities let you set policy in code. Small models can judge while a larger chat model keeps talking. That split keeps interactive latency intact.

Attribution stays explicit through this article. Every concrete number, name, speed, and behavior below comes from the Hugging Face blog community article New in llama.cpp: Decision Models by Xuan-Son Nguyen and Victor Mustar of ggml-org dated October 2 2026. No outside benchmark or second source is used.

How the systemone endpoint works

The llama.cpp server supports decision models through the /v1/systemone endpoint. You send a state plus typed questions. The model returns a probability per option in a single forward pass. The article presents this as following the System One format from TypeSafe Jev, with details in PR 29818.

State is the material to judge. The article lists text, JSON, and screenshot as state forms. Text covers tickets, messages, and transcripts. JSON covers structured tool output and form data. Screenshot covers rendered pages and documents when the model supports images. One state can serve several questions at once.

Typed questions define the ballot. Each question carries a type and an option list. The server scores every option and returns probabilities. Your code picks the top option or applies a cutoff. The ballot never leaves your control.

The single forward pass claim is load bearing. Chat models need one forward pass per token plus parsing. Decision models process the state once and score all options together. That is why batching questions on one state can save time. The state encoding is reused instead of repeated.

PR 29818 is cited as the place for endpoint details. Readers who build clients should start from that reference plus the blog walkthrough. The blog is a community post on Hugging Face, not a manual page. Treat the endpoint path and field names as described there.

A simple request shape helps to picture the flow. State holds the ticket text. Question one asks for department with three options. Question two asks for a safety score. Question three asks for a null style check. The response carries probabilities for each option. Usage reports output tokens as 0 because no tokens were generated.

That output tokens detail is explicit in the source. The quick start example shows output_tokens 0 in usage. The zero marks scoring versus generation. Billing and logging code that counts tokens should expect that zero.

Error handling also gets simpler. There is no free text to parse, so there is no malformed label to rescue. Failures look like HTTP errors or low confidence, not creative spelling. Low confidence is a policy decision, not a parse fix.

Model lineup and speed table

The article presents six models in one table. Fields cover size, base, language coverage, image support, license, and median speed. Speeds are median ms per question on one NVIDIA RTX PRO 6000. NVIDIA is named here in prose as the GPU vendor for that measurement rig.

Julia-1 is the smallest entry at 144M on mmBERT-small. It covers 50 plus languages with no images under Apache 2.0 at 3 ms. Laya follows at 421M on ModernBERT-large for English with no images under Apache 2.0 at 5 ms. Both suit high volume triage where latency matters more than nuance.

Kev-4B is 4B on Qwen3.5-4B-Base for English with no images under Apache 2.0 at 12 ms. Lev is 4B on Qwen3.5 for English with no images under Apache 2.0 at 36 ms. The size tie with different speeds shows that base architecture and tuning change cost. Test both before assuming the smaller label means equal speed.

OpenJev is 27B on Qwen3.8-27B with 6 languages plus images under CC BY-NC 4.0 at 43 ms. Clef is 27B on Qwen3.8-27B for English plus images under Apache 2.0 with speed not listed. The table lists Clef without a speed number, so no latency claim is made for it here.

Licenses deserve a careful read. Five entries use Apache 2.0. OpenJev uses CC BY-NC 4.0. That non commercial term changes where you can deploy it. Check the model card before commercial use.

Image support splits the table cleanly. Julia-1, Laya, Kev-4B, and lev do not take images. OpenJev and Clef do take images. Only the 27B pair reads documents and screenshots. Plan hardware and license around that split.

ModelBaseLanguagesImagesLicenseMedian speed per question
Julia-1 144MmmBERT-small50 plus languagesNoApache 2.03 ms
Laya 421MModernBERT-largeEnglishNoApache 2.05 ms
Kev-4B 4BQwen3.5-4B-BaseEnglishNoApache 2.012 ms
lev 4BQwen3.5EnglishNoApache 2.036 ms
OpenJev 27BQwen3.8-27B6 languagesYesCC BY-NC 4.043 ms
Clef 27BQwen3.8-27BEnglishYesApache 2.0Not listed

Per HF/ggml-org, speeds are median ms per question on one NVIDIA RTX PRO 6000. That line is the standing speed note for the table. Medians smooth outliers but local disks, prompts, and builds still shift results. Treat the numbers as ordering guidance, not a guarantee.

Size selection is a cost accuracy trade. Julia-1 answers in 3 ms and fits tight loops. Kev-4B costs 12 ms and brings more judgment. OpenJev costs 43 ms and adds vision plus wider language help. The article advice is to try sizes rather than assume bigger always wins.

Language coverage also shapes choice. Julia-1 lists 50 plus languages, which helps multilingual queues. OpenJev lists 6 languages, which helps mixed document sets with images. Laya, Kev-4B, lev, and Clef are listed as English. Match the queue to the row.

Question types with billing example

The quick start uses three question types named choice, score, and noul. The example routes a support ticket among billing, shipping, and technical options. The same request pattern shows how one state can feed several judgments at once.

Choice picks one label from a list. Score returns a graded judgment. Noul covers the null style check from the example set. The article groups them as the three types to try first. Each type still returns probabilities per option.

The billing, shipping, and technical example makes the idea concrete. A ticket that mentions an invoice maps to billing. A ticket that mentions a parcel maps to shipping. A ticket that mentions an error code maps to technical. Descriptions on the options help the small models tell those apart.

Output accounting stays consistent across types. Usage shows output_tokens 0 because scoring does not emit tokens. That zero is part of the example output. Cost tracking should use the question count and state size, not completion length.

Question typeWhat it doesExample use
choicePicks one option from a closed list with probabilitiesRoute ticket to billing, shipping, or technical
scoreScores the state on a graded scale with probabilitiesRate urgency or safety for the same ticket
noulRuns the null style check from the starter setConfirm whether any action applies to the ticket

Keep option text explicit. The article tip on descriptions applies directly here. Julia-1 gave billing 0.99 with descriptions while shipping without descriptions lost. The same ticket can flip when labels lack context. Write each option as a short definition, not a single word.

Batching is natural with this shape. One ticket state plus three typed questions returns three ballots. The state is processed once on Kev-4B, lev, and OpenJev. That reuse is called out as a speed win. Fewer encodings means lower total latency.

Response handling stays uniform. Read the top probability per question. Compare it to a per model cutoff. Route, queue, or escalate based on that test. Log the full distribution when audit matters.

Quick start with Kev 4B

The article quick start serves Kev-4B from GGUF with a single serve command. The command is llama serve with -hf ggml-org/Kev-4B-GGUF. That line pulls the packaged weights and starts the local server. No separate chat template setup is described.

Point requests at /v1/systemone once the server is up. Build a state from ticket text or JSON. Add choice, score, and noul questions in one payload. The billing, shipping, and technical example is the suggested first test.

Check /v1/models to confirm the loaded id. Router mode can load models on demand, so the id you request must match a known entry. The article says the endpoint lists ids. Use that list instead of guessing names.

Inspect usage after each call. Expect output_tokens 0 in usage for decision calls. That zero confirms scoring path rather than generation path. If completions appear, the request missed the decision endpoint.

Start with clear option descriptions. Copy the article pattern where billing carries a full description. Confirm that probabilities look sane before wiring automation. A 0.99 on billing for an invoice ticket is the kind of crisp result to expect when text matches.

Then test a vague ticket. The article contrast is sharp with 0.25 on Julia-1 versus 0.80 on Kev-4B for the same vague input. That gap shows why cutoffs must be per model. One global threshold will either spam review or miss true cases.

Local runners should note the GGUF angle. The -hf flag points at ggml-org/Kev-4B-GGUF. That packaging is built for llama.cpp serving. Keep the server build fresh enough to include PR 29818 support.

Images and vision support

OpenJev reads documents and screenshots. That vision ability is limited to the image capable rows. Clef shares the image capable flag. The other four models take text or JSON state only.

The projector detail matters for setup. The vision projector auto downloads on first image use. Expect a one time fetch before the first screenshot scores. Later calls reuse the cached piece.

Screenshot state suits rendered content. Think receipts, error dialogs, scanned forms, and web captures. Text state suits tickets and logs. JSON state suits tool output with fields intact. Pick the state form that preserves the signal.

License and size travel with vision. OpenJev pairs vision with 27B size, 6 languages, CC BY-NC 4.0, and 43 ms median. Clef pairs vision with 27B size, English, Apache 2.0, and no listed speed. There is no small vision option in this table.

Latency planning should reflect that. Vision calls cost more than text triage on Julia-1 at 3 ms. Reserve OpenJev for cases where the image carries the answer. Keep text only models in front for cheap first pass routing.

Hugging Face hosts the blog and the model cards behind this release. Hugging Face is named here in prose as the publisher of the community article. Model pages carry the license and projector notes.

Router mode and model listing

Router mode loads models on demand. You can pick a different decision model per request. That choice lets one server front several judges without holding all weights in memory at once.

The /v1/models endpoint lists ids. Use those ids in requests to select Julia-1, Laya, Kev-4B, lev, OpenJev, or Clef as needed. The article frames this as pick per request. Small tickets can use Julia-1 while tricky tickets escalate to Kev-4B.

On demand loading adds a first hit cost. The first request for an unloaded model waits for weights. Later requests reuse the loaded model. Design warmup calls for latency sensitive paths.

Per request selection also helps testing. Run the same state against Julia-1 and Kev-4B and compare distributions. The vague ticket contrast of 0.25 versus 0.80 came from that kind of comparison. Keep the state fixed and swap only the model id.

Operationally, router mode fits mixed queues. Billing triage at high volume can stay on Julia-1. Edge cases can spill to Kev-4B. Document checks can go to OpenJev. One endpoint fronts the whole policy.

Tuning tips for accuracy and speed

Try sizes before locking in. Julia-1 at 144M, Laya at 421M, Kev-4B at 4B, lev at 4B, and OpenJev at 27B span three orders of magnitude. Larger does not always win enough to justify latency. Measure on your own tickets.

Describe options in full sentences. The article example shows Julia-1 giving billing 0.99 with descriptions while shipping without descriptions lost. Single word labels starve small encoders. One line of context per option can move the top score.

Set per model confidence cutoffs. A vague ticket scored 0.25 on Julia-1 but 0.80 on Kev-4B in the article example. That spread means a 0.50 global cutoff would treat the two models very differently. Calibrate each model on held out tickets and store the threshold with the model id.

Batch questions on one state. The state is processed once on Kev-4B, lev, and OpenJev. Three questions in one call cost less than three separate calls. Group department, urgency, and safety checks together.

Try quants, including Q8_0. The article names Q8_0 as a quantization to test. Higher precision quants can recover accuracy at some speed and memory cost. Lower quants can save memory for router fleets. Test the quant on your cutoff, not just top one accuracy.

Log scores during rollout. Probabilities make drift visible before users complain. Track top score histograms per model and per question. Recalibrate cutoffs when the mix shifts.

Keep the request set minimal. Extra options add scoring work and confuse close calls. Remove dead labels from the ballot. Split overloaded questions into two typed checks.

When to use a decision model

Routing is the headline use. Billing, shipping, and technical triage maps cleanly to choice questions. High volume queues benefit most because 3 ms and 5 ms judges keep up. Escalate low confidence to a larger model or a human.

Moderation is a second fit. Allowed versus blocked is a closed set with audit needs. Scores give a natural hold queue. Policy text lives in option descriptions for consistency.

Agent step checks are a third fit. After each tool call, ask whether the step succeeded, needs retry, or needs a new plan. The closed set keeps loops tight. Probabilities gate retries without string parsing.

Next action choice closes the loop. Offer the agent a short menu of moves. Score once and act on top choice above cutoff. The single forward pass keeps the control loop fast.

Chat models still own open ended writing. Decision models own the ballot. That split is the core mental model from the article. Generate with one, judge with the other.

TypeSafe Jev is the named origin for the System One request format. TypeSafe is named here in prose without a link, per house sourcing. The llama.cpp server adopts that shape at /v1/systemone with details in PR 29818.

FAQ

What is the new llama.cpp decision model endpoint?
Per the Hugging Face blog community article by Xuan-Son Nguyen and Victor Mustar of ggml-org, it is /v1/systemone, which takes a state plus typed questions and returns a probability per option in a single forward pass.

Which models are available and how fast are they?
Per HF/ggml-org, Julia-1 runs 3 ms, Laya runs 5 ms, Kev-4B runs 12 ms, lev runs 36 ms, and OpenJev runs 43 ms as median ms per question on one NVIDIA RTX PRO 6000, while Clef speed is not listed.

What are the three starter question types?
Per the article quick start, they are choice, score, and noul, shown with a billing, shipping, and technical example where usage reports output_tokens 0.

How do I start a local server?
Per the article, run llama serve with -hf ggml-org/Kev-4B-GGUF, confirm ids at /v1/models, then send state plus questions to /v1/systemone.

Which models take images?
Per HF/ggml-org, OpenJev and Clef take images and read documents and screenshots, with the vision projector auto downloading on first image use, while the other four rows are text or JSON only.

How should I tune cutoffs and cost?
Per the article, describe options in full, set per model cutoffs because a vague ticket scored 0.25 on Julia-1 versus 0.80 on Kev-4B, batch questions because state is processed once on Kev-4B, lev, and OpenJev, and try quants including Q8_0.

Bottom Line

llama.cpp decision models replace chat text parsing with closed set scoring at /v1/systemone in a single forward pass, per the Hugging Face blog community article by Xuan-Son Nguyen and Victor Mustar of ggml-org. Start with Kev-4B for balanced judgment, keep Julia-1 for 3 ms triage, use OpenJev for screenshot and document checks, describe options, calibrate per model cutoffs, batch questions, and test Q8_0 before scaling.

Sources

Per the Hugging Face blog community article New in llama.cpp: Decision Models by Xuan-Son Nguyen and Victor Mustar of ggml-org dated October 2 2026, with implementation details in PR 29818: New in llama.cpp: Decision Models on the Hugging Face blog.

Share