Instructions to use PelaAI/KnowLine-4B-Gen4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PelaAI/KnowLine-4B-Gen4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="PelaAI/KnowLine-4B-Gen4")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("PelaAI/KnowLine-4B-Gen4") model = AutoModelForMultimodalLM.from_pretrained("PelaAI/KnowLine-4B-Gen4", device_map="auto") - Notebooks
- Google Colab
- Kaggle
KnowLine-4B-Gen4
The fourth release of PelaAI's 4B System One decision model, trained on a single all-in-one machine through an agent-driven data loop.
Inference guide ยท ไธญๆ่ฏดๆ ยท weights Apache-2.0 ยท previous version: KnowLine-4B-Gen3
Highlights:
Decision Index 0.3 public suite: 64.90 (self-run), 1.79 points above Gen3 (63.11). Most of the gain comes from a few benchmarks whose tasks appear in our training data, led by BPoMP; see What changed in Gen4. For comparison, on the board Perplexity Decider v1.1 (27B) scores 62.25 and Jev 1.13 scores 57.96, and the best model of 5B or smaller has a public score of 50.82; see Comparison.
Harder to steer with planted instructions: an instruction planted in the state changes the answer 1.6% of the time, down from 3.4% for Gen3. With injection templates not seen in training, the rate is 6.2%, down from 10.2%.
Calibrated probabilities: a temperature is folded into the weights. Computed with the board's method, ECE is about 0.04-0.05 (board median: 0.084) and the Brier score is about 0.29-0.31, roughly 2nd-3rd among the 114 models on the board; see Calibration.
On par with Gen3 on our other evaluations: JevBench / JevBench-hard 86.1 / 72.1 (Gen3: 86.6 / 72.1), Open-Jev 1.1 test / OOD 88.4 / 87.1 (88.4 / 87.2), C-Eval 78.7 (79.2).
Images without any image training: Gen4 was trained on text only, yet its serving path accepts images and the decision skills carry over. On our reconstruction of the Decision Index 0.3.1 vision public suite it scores 66.4, against 59.4 for the stock Qwen3.5-4B on the same harness (self-run); see Images.
KOF '98 harness: 14-3-1 against Jev 1.13 and 10-4-4 against StartLux-Decision-4B in single-bout mirror matches; see Game harness.
An AI agent ran the whole loop end to end. In each round it:
- identified where the model was weak;
- proposed a targeted data group;
- built and decontaminated the data;
- trained and evaluated the model on it;
- kept or rejected the change based on the evidence.
For this round, see What changed in Gen4.
KnowLine is an independent model. It is not affiliated with, endorsed by, or derived from TypeSafe or Jev.
What it does
You send a state (text or a chat) and up to 64 typed questions: yes/no, choose one of k, or score on a rubric. The model answers every question with a probability distribution over its options:
- one forward pass per question, with no text generation;
- the output is the probability of each option's label token;
- existing Jev clients only need to change the base URL, and users of earlier versions only need to change the model name, because the interface and serving settings are unchanged.
| Base model | Qwen/Qwen3.5-4B (Apache-2.0) |
| Training | LoRA SFT (rank 32, alpha 64, language model only), merged into bf16 weights |
| Release format | bf16 weights, with a readout temperature T = 1.32 folded into the final norm weight (see Calibration). The vision tower and MTP head are unchanged from the base model; the config, tokenizer and chat template are identical to those of Gen1-Gen3. |
| Languages | English, Simplified Chinese, Traditional Chinese |
Quickstart
pip install "sglang==0.5.21" "transformers==5.12.1" requests
bash serve_knowline.sh PelaAI/KnowLine-4B-Gen4 0 8080 # SGLang (FP8 at load) on :9080 + /v1/systemone on :8080
curl -s http://127.0.0.1:8080/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "m",
"state": "Customer: my order arrived broken, I want my money back.",
"questions": {"refund": {"type": "noul", "instructions": "Should the agent offer a refund?"},
"tone": {"type": "choice", "instructions": "Customer tone?",
"criteria": {"angry": "Angry", "neutral": "Neutral", "happy": "Happy"}}}}'
- Front end:
knowline_server.pyis a single file that depends only on transformers and requests. It is the same file as in Gen3. - Full settings: the exact settings used for our Decision Index run are in INFERENCE.md.
What changed in Gen4
Finding the gaps. Knowledge and reasoning (43.9) remained the weakest area. Beyond that, the agent looked for abilities that did not carry over to new settings: the same question asked in a different format, unseen injection templates, tool selection among near-identical tools, and tasks from domains absent from training.
New data. Data groups aimed at those gaps: task instructions written inside the state, format variants of existing items, more tool-retrieval and hard tool-selection items, product-search relevance, a dozen or so new task families, injection-invariance pairs with new templates (the same state with and without a planted instruction), and news-topic classification. Arithmetic, knowledge and Jev-style items from earlier rounds were replayed. All of the data went through a revised decontamination check; see Disclosures.
Training. Gen3 was trained further on this data, and the result was averaged 50/50 with Gen3. The extra training made the model more overconfident, so a single temperature was fitted and folded into the weights. It changes the size of the probabilities, but not which option ranks first.
Selection. This round compared 37 candidates (new runs, Gen3 trained further, and their averages with Gen3), with Decision Index 0.3 public scores between 61.18 and 65.39. Every candidate had to pass the same checks against Gen3: no significant drop on our held-out sets, Open-Jev 1.1 OOD or JevBench-hard; an injection rate of at most 4%; and no regression beyond set margins on new-domain tasks, sibling skills, format consistency, unseen injection templates or calibration (measured after applying the temperature). Gen4 has the highest public score among the candidates that passed. Several averages with higher public scores were rejected because they lost JevBench-hard items or format consistency.
Where the gain comes from. About 1.58 of the 1.79 points (88%) come from benchmarks whose tasks were added to the training data in this round:
- BPoMP, 54.5 โ 90.7 (0.60 index points): original-versus-perturbed limerick items in the same style. Earlier versions had no such data.
- Amazon ESCI, 38.9 โ 45.6: query-product pairs from the ESCI train split.
- cfcolor, 25.4 โ 35.2: palettes from the same source dataset that the benchmark uses (O'Donovan et al., 2014).
- Humicroedit, 23.1 โ 32.0: the SemEval-2020 Task 7 train and dev splits and FunLines; the benchmark uses the test split.
- Smaller gains: ToolRet (61.2 โ 64.9), BRIGHT (39.2 โ 42.0), SATA-Bench (29.2 โ 33.6), FinEntity (86.8 โ 90.2) and ACOS (43.3 โ 46.1).
On same-skill benchmarks that we did not train on (BLiMP for BPoMP, HaHackathon for Humicroedit, WANDS for ESCI, Toucan for ToolRet), the changes are small or mixed, so most of this gain is specific to those tasks. The largest drops are all small; the biggest is WinoGrande (78.8 โ 77.7).
Comparison
Scores on the Decision Index 0.3 public suite. The full 0.3 score also includes private tests that only the maintainers run (weighted 0.5 same-skill and 0.3 new-domain); our full score is not available yet.
| model | size | DI 0.3 public | DI 0.3 full | source |
|---|---|---|---|---|
| KnowLine-4B-Gen4 | 4B | 64.90 | not yet scored | self-run, official kit |
| KnowLine-4B-Gen3 | 4B | 63.11 | not yet scored | self-run, official kit |
| KnowLine-4B-Gen2 | 4B | 62.54 | not yet scored | self-run, official kit |
| KnowLine-4B-Gen1 | 4B | 60.47 | not yet scored | self-run, official kit |
| Perplexity Decider v1.1 | 27B | 62.25 | 62.75 | board |
| Clef | 27B | 61.71 | 53.08 | board |
| Jev 1.13 | (API) | 57.96 | 60.11 | board |
| RSI-Jev v6.1-VL | 4B | 50.98 | not listed | self-reported |
| ezjev 4B s2 | 4B | 50.82 | 46.95 | board |
| jiwo 4B | 4B | 45.76 | 42.86 | board |
| Nox 4B | 4B | 44.21 | 44.95 | board |
Public-suite scores by area (DI 0.3):
| area | KnowLine-4B-Gen4 | KnowLine-4B-Gen3 | Perplexity Decider v1.1 (27B) | Jev 1.13 |
|---|---|---|---|---|
| Knowledge & reasoning | 44.6 | 43.9 | 52.0 | 53.9 |
| Language | 65.5 | 64.8 | 67.5 | 59.2 |
| Retrieval & routing | 72.5 | 70.9 | 61.3 | 55.4 |
| Tools & agents | 87.6 | 86.6 | 78.9 | 75.1 |
| Arts & taste | 59.2 | 49.8 | 47.0 | 39.1 |
Releases
The model is trained in a self-evolving loop, and new versions will follow.
| model | date | DI 0.3 public | DI 0.2.1 | golden held-out (en / zh-Hans / zh-Hant) | notes |
|---|---|---|---|---|---|
| KnowLine-4B-Gen4 (this model) | 2026-10-10 | 64.90 | โ | 69.7 / 70.1 / 65.1 | fourth release |
| KnowLine-4B-Gen3 | 2026-10-09 | 63.11 | โ | 69.4 / 70.8 / 65.0 | third release |
| KnowLine-4B-Gen2 | 2026-10-08 | 62.54 | โ | 69.7 / 71.2 / 64.8 | second release |
| KnowLine-4B-Gen1 | 2026-10-07 | 60.47 | 60.92 | 68.7 / 69.9 / 64.1 | first release |
Since Gen2, we run only Decision Index 0.3, not 0.2.1.
Evaluation
All results are self-run and have not been verified by a third party.
Held-out and Chinese evaluations
| suite | Gen4 | Gen3 |
|---|---|---|
| In-house evaluation set, English | 69.7 | 69.4 |
| In-house evaluation set, Simplified Chinese | 70.1 | 70.8 |
| In-house evaluation set, Traditional Chinese | 65.1 | 65.0 |
| C-Eval (4 categories, macro) | 78.7 | 79.2 |
| Open-Jev 1.1 test / OOD | 88.4 / 87.1 | 88.4 / 87.2 |
| Prompt injection: answers changed (lower is better) | 1.6% | 3.4% |
| Prompt injection with templates not seen in training (lower is better) | 6.2% | 10.2% |
None of the differences on the in-house sets or on Open-Jev 1.1 OOD is significant in a paired sign test.
Web operation (Mind2Web official test splits, evaluation only)
| split | Gen4 element selection | Gen4 operation (balanced) | Gen3 element selection | Gen3 operation (balanced) |
|---|---|---|---|---|
| test_task (websites seen in training, new tasks) | 92.9 | 97.0 | 92.7 | 97.0 |
| test_website (new websites) | 91.9 | 95.7 | 91.3 | 97.3 |
| test_domain (new domains) | 91.7 | 97.2 | 91.8 | 97.2 |
- About the same as Gen3. The lower operation score on test_website comes from 3 more misses among its 72 SELECT steps; the differences are within noise.
- So far, we use this only to explore the model for RPA-style automation and to check that it generalises to some degree.
- The task is to pick the target element from a candidate set: the target plus up to 5 other elements sampled from the page. This is easier than the original Mind2Web protocol, so do not compare these results with the Mind2Web leaderboard.
Images (zero-shot)
- No image training: Gen4 keeps the base model's vision tower unchanged and saw no images in training.
knowline_server.pyaccepts OpenAI-styleimage/image_urlcontent parts in a chat state and passes them to SGLang; each question is still one prefill. - How we measured it: we rebuilt the Decision Index 0.3.1 vision public suite as closely as the board's descriptions allow: 11 benchmarks, the board's row counts where the sources allow, images capped at 1.6 MP, chance-corrected skill and the board's weights. We ran it through the release server: 12,192 requests, 0 failed.
| benchmark | Gen4 | stock Qwen3.5-4B (same front end) |
|---|---|---|
| CV-Bench | 75.7 | 72.1 |
| BLINK | 58.1 | 59.9 |
| RealWorldQA | 64.7 | 59.7 |
| CharXiv | 66.3 | 52.4 |
| InfographicVQA | 86.3 | 73.2 |
| Mind2Web | 77.4 | 65.3 |
| Winoground | 80.5 | 73.5 |
| KIE (CORD+FUNSD) | 95.5 | 95.7 |
| Moderation (Hateful Memes) | 27.1 | 18.4 |
| R-Bench-M | 14.6 | 6.3 |
| MMMU-Pro vision | 31.9 | 21.4 |
| weighted public index | 66.4 | 59.4 |
- The decision skills generalise to images: Gen4 is 7.0 points above the stock model, with the largest gains on charts, infographics, web screenshots and exam pages, although no image data was used in training.
- Indicative only: these are self-run numbers on our reconstruction, not on the board's rows. Our CharXiv distractors lean towards numeric questions, and our Mind2Web marks and KIE questions are our own construction.
- Weakest areas: Hateful Memes, the expert questions of R-Bench-M and MMMU-Pro, and BLINK (multi-image perception).
Game harness (KOF '98)
- Setup: single-bout character-mirror matches, 18 games per pair, argmax actions; the same settings as for Gen1-Gen3.
- Result: 14-3-1 against Jev 1.13, 10-4-4 against StartLux-Decision-4B and 3-11-4 against Gen3; 27-18-9 overall, a score of 0.583 [0.45, 0.71]. Against Jev and StartLux this is slightly below Gen3 (16-1-1 and 12-4-2), but within the wide intervals.
- Play style: it chooses moves by distance: mostly the 623C anti-air uppercut at close range (79%), and special_2 at mid range (69%) and long range (92%). Overall, it uses special_2 52% of the time and the uppercut 28%.
- Caveats:
- 18 games per pair give wide intervals.
- 39 of the 54 games ended at time-out, so most wins were decided on remaining health. Gen4 won 7 games by K.O. (4 against Jev, 2 against StartLux, 1 against Gen3); 8 of its 11 losses to Gen3 were by K.O.
- This is a measured result in this harness, not general fighting-game skill.
Calibration
Computed on the 0.3 public suite using the method described by the Decision Index board: each field is scored right or wrong, confidence is the probability of the chosen option, benchmarks are weighted equally, and 10 equal-width bins are used.
| metric | Gen4 | Gen3 | board median (114 models) | Jev 1.13 |
|---|---|---|---|---|
| ECE (lower is better) | 0.04-0.05 | 0.07-0.08 | 0.084 | 0.074 |
| Brier score (lower is better) | 0.29-0.31 | 0.31-0.33 | 0.49 | 0.36 |
| Confidence โฅ95% but wrong | 1.7% | 2.7% | 1.3% | 2.1% |
| Mean confidence / accuracy | 0.82 / 0.77 | 0.84 / 0.76 | 0.81 / 0.74 |
- Temperature in the weights: the final norm weight includes a readout temperature T = 1.32, so every logit is divided by 1.32 before the softmax over the option labels. The top-ranked option does not change, so neither do the Decision Index scores; the probabilities are just less extreme. T was fitted on a set of our own evaluation items whose distribution is covered by the training data, not on the held-out sets or on Decision Index rows.
- Why ranges: the board does not publish which 32 benchmarks it uses. We checked our computation against three board models with public results; our ECE was within about 0.015 of the board's values, so we report ranges over the plausible benchmark sets.
- Still slightly overconfident: mean confidence is about 3-5 points above accuracy. Most of the gap comes from hard reasoning (HLE, CRUXEval, MuSR) and New Yorker caption matching.
- Recommendation: the temperature was fitted on our data. If your use depends on probability thresholds, check calibration on your own data.
Disclosures
- Decision Index format training data: Gen4 is the product of four training rounds. In this round, about 34% of the rows were built in the request format of Decision Index benchmarks (some with reworded instructions), and another 11% are copies of the same items rendered in other formats. These include train splits of public datasets rewritten in that format (for example GSM8K, WinoGrande XL, ACOS and Amazon ESCI) and synthetic items written in the same format.
- Decontamination:
- All training data of this round, new and replayed, was checked with a revised tool against every Decision Index 0.3 row, all of our evaluation sets and the other public benchmarks we report (Intern-Decision, S1MB, JevBench, MMDM).
- Rows that nearly duplicate an evaluation item, repeat its question or claim, or contain its whole state were removed: 1,909 of 539,647 rows. Rows that only share a source passage with an evaluation item, such as the same Wikipedia paragraph or tool description, were kept (23,684 rows).
- BPoMP and cfcolor training items come from the same source datasets as these benchmarks (OEDILF limericks; the O'Donovan et al. colour ratings), with every benchmark item excluded. An item-level check against all 5,000 Decision Index rows of each benchmark found no shared poems and no shared rating items.
- Selection on evaluations: of the 37 candidates, we picked the one with the highest Decision Index 0.3 public score among those that passed all of our release checks.
- Game data: labels come from simulator rollouts or engines. Seeds and start states are disjoint from the evaluation states.
- No model outputs as labels: no output of Jev or any other decision model was used as a training label.
- Teacher-labelled synthetic data: LLM teachers wrote and labelled our synthetic tasks. The labels were filtered, but not all of them were checked by a human.
Limitations
- Knowledge-heavy reasoning: weaker than larger models. The knowledge area score is 44.6 (DI 0.3); MMLU-Pro barely moved (51.6), and HLE remains at chance level.
- Math: answers are given without step-by-step reasoning. The rebuilt 0.3 GSM8K scores 61.6, about the same as Gen3; this score comes from training on the GSM8K train split rewritten in the same format.
- New domains: on our own tasks from domains not in training, Gen4 scores about the same as Gen3 (49.3 vs 49.5). The gains of this round are in robustness and calibration, not in new-domain ability.
- Prompt injection: an instruction planted in the state changes the answer about 2% of the time on our injection set, and about 6% of the time with templates not seen in training. Keep untrusted text clearly delimited.
- Images: there was no image training. The image scores above are zero-shot and self-run on our own reconstruction of the suite, so check them on your own data before relying on them.
- Private tests: most of the public-suite gain over Gen3 comes from benchmarks whose tasks are in our training data, so the 0.3 private tests may show a much smaller gain, or even a lower score.
Citation
@misc{knowline4bgen4,
title = {KnowLine-4B-Gen4: a 4B decision model},
author = {PelaAI},
year = {2026},
url = {https://huggingface.co/PelaAI/KnowLine-4B-Gen4}
}
License
- Weights: Apache-2.0, the same as the base model.
- Code:
knowline_server.pyis MIT (see the file header).
- Downloads last month
- 8