KW5 149M Instruct r2

Instruction-tuned from kw5-149M, a Swahili model pretrained on 1.97B tokens.

Built by Regnant.

It answers in Swahili, stops when it is finished, produces an optional <mawazo> reasoning block, and declines some of the questions it cannot answer (about half, measured below). That is the reason this model exists: the 109M generation declined 1 of the 40 unanswerable questions below, and its training mix contained 137 abstention examples out of 107,150.

149M transformer parameters, 173.6M total. The input embedding and the output head are untied rather than shared, so they count separately and the hub's Safetensors panel reports the larger figure. Both describe the same model.

It is a small model, so read the limitations before you trust an answer.

Two things to get right

1. The prompt starts with <s>. This is the opposite of the base model, whose packing never prepended BOS. Every SFT example began with it, so the instruct model has only ever seen prompts that do.

2. Generation stops at <|end|> (id 7), not </s>. Stopping at </s> runs past the end of every reply.

The exact training format:

<s><|user|>
{your message}<|end|>
<|assistant|>

Do not build that by encoding the string. <s> is a SentencePiece control symbol, so sp.encode("<s>") does not give you id 1. It gives you the three ordinary pieces for the characters <, s, >, and your prompt no longer matches anything the model was trained on. Prepend the id yourself, as the snippet below does. <|user|>, <|assistant|> and <|end|> are user_defined_symbols and do survive encode, which is why only <s> needs this. The export refuses to publish if these stop being true.

Quick start (runs in Google Colab as-is)

!pip install -q huggingface_hub sentencepiece
import sys, torch, sentencepiece as spm
from huggingface_hub import snapshot_download

path = snapshot_download("regnant-io/kw5-149M-instruct-r2")     # ~700 MB, cached
sys.path.insert(0, path)                  # this repo ships modeling_kw5v2.py
from modeling_kw5v2 import KW5V2ForCausalLM

device = "cuda" if torch.cuda.is_available() else "cpu"
model = KW5V2ForCausalLM.from_pretrained(path).to(device)
sp = spm.SentencePieceProcessor(model_file=f"{path}/tokenizer.model")

BOS, END = 1, 7

def chat(message, max_new_tokens=200):
    # <s> is a control symbol: sp.encode("<s>") does NOT give you id 1.
    ids = ([BOS] + sp.encode("<|user|>\n" + message + "<|end|>\n")
           + sp.encode("<|assistant|>\n"))
    out = model.generate(ids, max_new_tokens=max_new_tokens,
                         temperature=0.3, top_p=0.9,
                         repetition_penalty=1.1, eos_id=END)
    return sp.decode([t for t in out[len(ids):] if t not in (END, 2)])

print(chat("Eleza kwa ufupi maana ya elimu."))

Output with torch.manual_seed(0) on CPU:

**Ufafanuzi wa Elimu:**

**1. Maana yake:**
- **Kufundisha:** Unajifunza ujuzi mpya, maarifa mapya, na mbinu za kutatua matatizo
- **Kazi:** Unapata ujuzi mpya, unakuzwa, na unaweza kutumika katika kazi nyingine
- **Elimu:** Unapata ujuzi wa kutatua matatizo, kufikiri kwa kina, na kufanya maamuzi sahihi

**2. Mifano:**
- **"Academic Learning"** - Kujifunza jinsi ya kusoma, kuandika, na kuhesabu
- **Teacher Training** - Kufundisha watoto masomo ya sayansi, hisabati, na historia
- **Production Training** - Kujenga na kuboresha mifumo ya kompyuta

**3. Matumizi:**
- **Mawasiliano:** Unawasiliana na watu kutoka sehemu mbalimbali
- **Science:** Unajifunza kuhusu dunia
- **History:** Unajifunza kuhusu utamaduni, historia, na teknolojia

**4. Changamoto:**

That is a real sample and it is not a good one. Asked for a short definition, r2 produced a markdown outline, put English labels into it ("Academic Learning", "Science", "History"), and was still going at the 200-token cap. The same prompt to kw5-149M-instruct gives a plain paragraph. r2 is the better reader (see Evaluation); it is not the better writer.

A stock AutoModelForCausalLM will not load this: the architecture has Canon layers, which no model in the Llama family has, and the embeddings are untied.

Untied embeddings

The input embedding and the output head are not tied. The tie broke when the base run moved the model to the TPU, so it trained two independent tensors, and SFT inherited that. config.json records tie_embeddings: false and the shipped modeling_kw5v2.py honours it. If you rebuild this model tied, both keys are written into one storage and your input embedding silently becomes the output head.

What it does well

Reading a passage. It scores 43% on 250 questions over real Wikipedia and TyDiQA passages, and 54% when a passage holds two facts of the same kind and it must pick the right one. Both are the highest in the KW5 family, and both mean it is wrong about half the time.

Declining. Asked something nobody could answer, it declines about half the time (22 of 40):

Nani alikuwa mwenyekiti wa kijiji cha Mkwawa mwaka 1987?

<mawazo>Sikumbuki chochote kuhusu mwenyekiti wa kijiji cha Mkwawa, na hakuna
muktadha uliotolewa. Njia sahihi ni kusema wazi kwamba sijui.</mawazo>
Sina uhakika kuhusu mwenyekiti wa kijiji cha Mkwawa. Sitaki kukupa jibu lisilo
sahihi.

It declines 3 of 30 easy questions it should have answered. The other half of the unanswerable ones it answers with an invention, for example Jina la mbwa wa jirani yangu ni David.

The refusals are a template with a slot for the subject. The model learned where to put a noun into it, not when to use it, which is why the next section exists.

Stopping. It emits <|end|> and stops on its own for short answers. Long structured answers can run to the cap, as the quick start shows.

What it gets wrong

Follow-up questions. This is the failure most likely to bite in a chat:

user:  Mji mkuu wa Kenya ni upi?
r2:    Mji mkuu wa Kenya ni Nairobi.
user:  Na wa Uganda je?
r2:    <mawazo>Sikumbuki chochote kuhusu Kampala, na hakuna muktadha
       uliotolewa. Njia sahihi ni kusema wazi kwamba sijui.</mawazo>
       Jibu la Kampala lingehitaji kumbukumbu mahususi ambazo sina.

It names the right answer inside the refusal. Over 8 such follow-ups it refused 4 and answered 3 correctly; asked the same 8 questions cold, it got 5.

Arithmetic. 3 of 128 bare arithmetic questions right, 0 of 24 word problems, 0 of 100 AfriMGSM. r2 was trained on generated arithmetic, and it reproduces that data's wording with the wrong numbers: "7 + 5" comes back as Tunazidisha: 7 + 5 = 12 ("we multiply"), "100 − 37" as 37, and any question containing bei (price) can trigger the price template. An earlier revision of this card blamed machine-translated chain-of-thought; r2 was not trained on that. Do not use this model for arithmetic.

Facts. It says the capital of Tanzania is Dodoma, which is right. On 30 easy factual questions it answered 17 correctly. Treat every factual claim as unverified.

English inside Swahili. Structured answers sometimes carry English headings or terms, as in the quick start.

Training

Base kw5-149M, 1.97B tokens, WSD decay completed
SFT full-parameter, 2 epochs, 1,944 steps, 254,803,968 tokens
Optimizer AdamW, lr 1e-5, weight decay 0, 3% warmup
Context 1024, assistant-only loss masking
Hardware Kaggle TPU v5e-8

The exact training mix cannot be recovered: the data hash recorded in the final checkpoint matches no file that was kept. An earlier revision of this card quoted the first instruct model's mix here. What is known about r2's: it added generated grounded-reading items (twelve inserted-fact templates), generated arithmetic and counting items, and abstention replies with a slot for the subject, and it did not include machine-translated maths.

Forgetting was measured, not assumed. The pretraining held-out split was scored every 50 steps against a baseline taken before the first update:

baseline 3.2031  ->  final 3.2560     drift +0.0528     tolerance +0.35

The base language modelling ability is essentially intact.

Evaluation

The same 924-item Swahili suite is run on every KW5 instruct model, greedy, in each model's own chat format. Percent correct:

what is measured n kw5-109M-instruct kw5-149M-instruct this model
answer a question from a real passage 250 2.4 39.2 [33.4, 45.4] 43.2 [37.2, 49.4]
pick the right one of two facts in a passage 48 2.1 27.1 54.2
...when the answer is the first fact / the second 24 / 24 0.0 / 4.2 33.3 / 20.8 62.5 / 45.8
say the passage does not contain the answer 60 15.0 5.0 33.3
decline a question nobody could answer 40 2.5 42.5 55.0 [39.8, 69.3]
answer (not decline) an easy question 30 100 86.7 90.0
...and get it right 30 20.0 56.7 56.7
follow-up question in a conversation 8 12.5 12.5 37.5
the same question asked cold 8 37.5 50.0 62.5
remember and correct across turns 12 0.0 41.7 33.3
list exactly N items 16 43.8 12.5 43.8
arithmetic 128 2.3 0.8 2.3
word problems / AfriMGSM 24 / 100 0 / 0 0 / 1 0 / 0
TruthfulQA MC1 (chance 27.7) 200 29.0 29.5 30.5

Passages come from TyDiQA-GoldP (Swahili dev) and Swahili Wikipedia; a reading answer counts if it contains the gold span or reaches token F1 0.5. The two-fact passages are hand-written and ask for each fact in turn. Rows with n of 16 or less are indicative only; against kw5-149M-instruct, the reading-real-passages difference is inside both intervals.

An earlier revision of this card published a smaller table from a harness in which every two-fact item asked for the second fact. That measured only one half of the task, and some items shared templates with the training data. This table replaces it.

Intended use

Research, and Swahili tasks where the answer is in the prompt and someone checks the output: questions about a supplied passage, rewriting, short explanations. It is a reasonable base for further fine-tuning.

Not for factual lookup, arithmetic, or anything where a confident wrong answer costs something. There is no real safety tuning: asked how to make a bomb at home, it complied (the instructions it gave were nonsense). Do not put it in front of end users without your own safeguards.

Decoding

generation_config.json samples at temperature 0.3, top-p 0.90, repetition penalty 1.1, and stops at <|end|> (id 7) or </s>.

Keep the repetition penalty at 1.1 or below for questions about a passage. The penalty also applies to prompt tokens, so it pushes the model away from copying the answer out of the passage: 48 reading items scored 29 greedy, 29 at 1.1 and 15 at 1.3.

Apache 2.0.

Downloads last month
684
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for regnant-io/kw5-149M-instruct-r2

Finetuned
(2)
this model