New in llama.cpp: Decision Models

Community Article
Published October 2, 2026

llama.cpp server now supports decision models through the /v1/systemone endpoint. You send a state (text, JSON, a screenshot) and typed questions. The model returns a probability for each option in a single forward pass.

The API follows the System One format introduced with TypeSafe's Jev model, so existing clients only need a new base URL. Implementation details are in PR #29818.

What is a decision model? A model that answers by scoring the options you give it, instead of generating text. A chat model spends one forward pass per output token, and its output still has to be parsed. A decision model reads the input once, and its answer is always one of your options, with a probability attached. Typical uses are routing a request, moderating content, checking that an agent's step worked, or choosing its next action.

Supported models

Model Size Based on Languages Images License Speed*
Julia-1 144M mmBERT-small 50+ no Apache 2.0 3 ms
Laya 421M ModernBERT-large English no Apache 2.0 5 ms
Kev-4B 4B Qwen3.5-4B-Base English no Apache 2.0 12 ms
lev 4B Qwen3.5-4B English no Apache 2.0 36 ms
OpenJev 27B Qwen3.8-27B en, de, fr, hi, zh, ja yes CC BY-NC 4.0 43 ms
Clef 27B Qwen3.8-27B English yes Apache 2.0 —

*Median time to answer one question, on one NVIDIA RTX PRO 6000.

Find these models in the Decision models collection, with more coming. The community Decision Index shows how they compare.

Quick start

Get the latest llama.cpp from llama.app (or run llama update), then start a model:

llama serve -hf ggml-org/Kev-4B-GGUF

A request contains a state and one or more questions. There are three question types:

Type You send You get
choice options, with optional descriptions the top option, plus a probability per option
score 2 to 10 levels, lowest first the expected level (can fall between two)
noul a yes/no question the probability of yes

Send a request with your state and questions:

curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "state": "Customer message: I was charged twice for my order last week and nobody has replied.",
    "questions": {
      "route": {
        "type": "choice",
        "instructions": "Which team should handle this?",
        "criteria": {
          "billing": "payments, charges, refunds, invoices",
          "shipping": "delivery, tracking, lost or late parcels",
          "technical": "bugs, errors, login problems"
        }
      },
      "angry": {
        "type": "noul",
        "instructions": "Is the customer angry?"
      },
      "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["can wait", "this week", "today", "right now"]
      }
    }
  }'

Response (values rounded):

{
  "model": "ggml-org/Kev-4B-GGUF",
  "answers": {
    "route": {
      "type": "choice",
      "choice": "billing",
      "probabilities": {"billing": 0.9049, "shipping": 0.0275, "technical": 0.0676},
      "confidence": 0.8574
    },
    "angry": {
      "type": "noul",
      "noul": 0.8208
    },
    "urgency": {
      "type": "score",
      "score": 2.2821,
      "legend": {"0": "can wait", "1": "this week", "2": "today", "3": "right now"},
      "probabilities": {"0": 0.036, "1": 0.1937, "2": 0.2225, "3": 0.5478},
      "confidence": 0.2821
    }
  },
  "usage": {"input_tokens": 130, "output_tokens": 0}
}

The full reference is in the server docs.

Images

Some models (OpenJev at the time of writing) can also read images, such as documents or screenshots. The vision projector downloads automatically:

llama serve -hf ggml-org/OpenJev-GGUF

For example, to classify an uploaded document:

import base64
import requests

with open("document.png", "rb") as f:
    image = "data:image/png;base64," + base64.b64encode(f.read()).decode()

response = requests.post("http://localhost:8080/v1/systemone", json={
    "state": "A file uploaded by a customer.",
    "images": [image],
    "questions": {
        "kind": {
            "type": "choice",
            "instructions": "What kind of document is this?",
            "criteria": {"invoice": None, "receipt": None, "contract": None, "other": None},
        },
    },
})

print(response.json()["answers"]["kind"]["choice"])  # invoice

The state can also be a list of chat messages. Any image_url part (data URL) is read as an image, same as chat completions.

Several models, one server

In router mode, models load on demand and you pick one per request:

llama serve
curl http://localhost:8080/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{"model": "ggml-org/Julia-1-GGUF:Q8_0", "state": "...", "questions": {...}}'

/v1/models lists the ids. With a single model loaded, the model field is ignored.

Tips

  • Try several models, of different sizes. Small models are faster, large ones know more. The Decision Index compares them.
  • Describe your options. Julia-1 routed "I was charged twice" to shipping with bare labels, and to billing (0.99) once each option had a description.
  • Pick your confidence cutoff per model. A common pattern is to act on confident answers and send the rest to a human. The right cutoff depends on the model: a vague ticket ("Hi, quick question about my account") scored 0.25 with Julia-1 but 0.80 with Kev-4B. Test on your own examples before choosing it.
  • Batch your questions. They are answered independently, and Kev-4B, lev and OpenJev process the state only once.
  • Try different quantizations. Like any GGUF, these models come in several precisions, for example llama serve -hf ggml-org/Kev-4B-GGUF:Q8_0.

What's next

New open decision models come out every week, and we'll keep adding the best ones. Is there one you particularly want? Tell us in the comments.

Community

fastino/GLiNER2.5-Decide compatibility ? 👀

Crafted a Space Playground to help understand these new kind of models: https://huggingface.co/spaces/fffiloni/decision-models-playground

Sign up or log in to comment