KW5 149M Instruct r2
Instruction-tuned from kw5-149M, a Swahili model pretrained on 1.97B tokens.
Built by Regnant.
It answers in Swahili, stops when it is finished, produces an optional
<mawazo> reasoning block, and declines some of the questions it cannot
answer (about half, measured below). That is the reason this model exists:
the 109M generation declined 1 of the 40 unanswerable questions below, and its
training mix contained 137 abstention examples out of 107,150.
149M transformer parameters, 173.6M total. The input embedding and the output head are untied rather than shared, so they count separately and the hub's Safetensors panel reports the larger figure. Both describe the same model.
It is a small model, so read the limitations before you trust an answer.
Two things to get right
1. The prompt starts with <s>. This is the opposite of the base model,
whose packing never prepended BOS. Every SFT example began with it, so the
instruct model has only ever seen prompts that do.
2. Generation stops at <|end|> (id 7), not </s>. Stopping at </s>
runs past the end of every reply.
The exact training format:
<s><|user|>
{your message}<|end|>
<|assistant|>
Do not build that by encoding the string. <s> is a SentencePiece control
symbol, so sp.encode("<s>") does not give you id 1. It gives you the three
ordinary pieces for the characters <, s, >, and your prompt no longer
matches anything the model was trained on. Prepend the id yourself, as the
snippet below does. <|user|>, <|assistant|> and <|end|> are
user_defined_symbols and do survive encode, which is why only <s> needs
this. The export refuses to publish if these stop being true.
Quick start (runs in Google Colab as-is)
!pip install -q huggingface_hub sentencepiece
import sys, torch, sentencepiece as spm
from huggingface_hub import snapshot_download
path = snapshot_download("regnant-io/kw5-149M-instruct-r2") # ~700 MB, cached
sys.path.insert(0, path) # this repo ships modeling_kw5v2.py
from modeling_kw5v2 import KW5V2ForCausalLM
device = "cuda" if torch.cuda.is_available() else "cpu"
model = KW5V2ForCausalLM.from_pretrained(path).to(device)
sp = spm.SentencePieceProcessor(model_file=f"{path}/tokenizer.model")
BOS, END = 1, 7
def chat(message, max_new_tokens=200):
# <s> is a control symbol: sp.encode("<s>") does NOT give you id 1.
ids = ([BOS] + sp.encode("<|user|>\n" + message + "<|end|>\n")
+ sp.encode("<|assistant|>\n"))
out = model.generate(ids, max_new_tokens=max_new_tokens,
temperature=0.3, top_p=0.9,
repetition_penalty=1.1, eos_id=END)
return sp.decode([t for t in out[len(ids):] if t not in (END, 2)])
print(chat("Eleza kwa ufupi maana ya elimu."))
Output with torch.manual_seed(0) on CPU:
**Ufafanuzi wa Elimu:**
**1. Maana yake:**
- **Kufundisha:** Unajifunza ujuzi mpya, maarifa mapya, na mbinu za kutatua matatizo
- **Kazi:** Unapata ujuzi mpya, unakuzwa, na unaweza kutumika katika kazi nyingine
- **Elimu:** Unapata ujuzi wa kutatua matatizo, kufikiri kwa kina, na kufanya maamuzi sahihi
**2. Mifano:**
- **"Academic Learning"** - Kujifunza jinsi ya kusoma, kuandika, na kuhesabu
- **Teacher Training** - Kufundisha watoto masomo ya sayansi, hisabati, na historia
- **Production Training** - Kujenga na kuboresha mifumo ya kompyuta
**3. Matumizi:**
- **Mawasiliano:** Unawasiliana na watu kutoka sehemu mbalimbali
- **Science:** Unajifunza kuhusu dunia
- **History:** Unajifunza kuhusu utamaduni, historia, na teknolojia
**4. Changamoto:**
That is a real sample and it is not a good one. Asked for a short definition, r2 produced a markdown outline, put English labels into it ("Academic Learning", "Science", "History"), and was still going at the 200-token cap. The same prompt to kw5-149M-instruct gives a plain paragraph. r2 is the better reader (see Evaluation); it is not the better writer.
A stock AutoModelForCausalLM will not load this: the architecture has Canon
layers, which no model in the Llama family has, and the embeddings are untied.
Untied embeddings
The input embedding and the output head are not tied. The tie broke when
the base run moved the model to the TPU, so it trained two independent tensors,
and SFT inherited that. config.json records tie_embeddings: false and the
shipped modeling_kw5v2.py honours it. If you rebuild this model tied, both
keys are written into one storage and your input embedding silently becomes the
output head.
What it does well
Reading a passage. It scores 43% on 250 questions over real Wikipedia and TyDiQA passages, and 54% when a passage holds two facts of the same kind and it must pick the right one. Both are the highest in the KW5 family, and both mean it is wrong about half the time.
Declining. Asked something nobody could answer, it declines about half the time (22 of 40):
Nani alikuwa mwenyekiti wa kijiji cha Mkwawa mwaka 1987?
<mawazo>Sikumbuki chochote kuhusu mwenyekiti wa kijiji cha Mkwawa, na hakuna
muktadha uliotolewa. Njia sahihi ni kusema wazi kwamba sijui.</mawazo>
Sina uhakika kuhusu mwenyekiti wa kijiji cha Mkwawa. Sitaki kukupa jibu lisilo
sahihi.
It declines 3 of 30 easy questions it should have answered. The other half of
the unanswerable ones it answers with an invention, for example
Jina la mbwa wa jirani yangu ni David.
The refusals are a template with a slot for the subject. The model learned where to put a noun into it, not when to use it, which is why the next section exists.
Stopping. It emits <|end|> and stops on its own for short answers. Long
structured answers can run to the cap, as the quick start shows.
What it gets wrong
Follow-up questions. This is the failure most likely to bite in a chat:
user: Mji mkuu wa Kenya ni upi?
r2: Mji mkuu wa Kenya ni Nairobi.
user: Na wa Uganda je?
r2: <mawazo>Sikumbuki chochote kuhusu Kampala, na hakuna muktadha
uliotolewa. Njia sahihi ni kusema wazi kwamba sijui.</mawazo>
Jibu la Kampala lingehitaji kumbukumbu mahususi ambazo sina.
It names the right answer inside the refusal. Over 8 such follow-ups it refused 4 and answered 3 correctly; asked the same 8 questions cold, it got 5.
Arithmetic. 3 of 128 bare arithmetic questions right, 0 of 24 word
problems, 0 of 100 AfriMGSM. r2 was trained on generated arithmetic, and it
reproduces that data's wording with the wrong numbers: "7 + 5" comes back as
Tunazidisha: 7 + 5 = 12 ("we multiply"), "100 − 37" as 37, and any question
containing bei (price) can trigger the price template. An earlier revision
of this card blamed machine-translated chain-of-thought; r2 was not trained on
that. Do not use this model for arithmetic.
Facts. It says the capital of Tanzania is Dodoma, which is right. On 30 easy factual questions it answered 17 correctly. Treat every factual claim as unverified.
English inside Swahili. Structured answers sometimes carry English headings or terms, as in the quick start.
Training
| Base | kw5-149M, 1.97B tokens, WSD decay completed |
| SFT | full-parameter, 2 epochs, 1,944 steps, 254,803,968 tokens |
| Optimizer | AdamW, lr 1e-5, weight decay 0, 3% warmup |
| Context | 1024, assistant-only loss masking |
| Hardware | Kaggle TPU v5e-8 |
The exact training mix cannot be recovered: the data hash recorded in the final checkpoint matches no file that was kept. An earlier revision of this card quoted the first instruct model's mix here. What is known about r2's: it added generated grounded-reading items (twelve inserted-fact templates), generated arithmetic and counting items, and abstention replies with a slot for the subject, and it did not include machine-translated maths.
Forgetting was measured, not assumed. The pretraining held-out split was scored every 50 steps against a baseline taken before the first update:
baseline 3.2031 -> final 3.2560 drift +0.0528 tolerance +0.35
The base language modelling ability is essentially intact.
Evaluation
The same 924-item Swahili suite is run on every KW5 instruct model, greedy, in each model's own chat format. Percent correct:
| what is measured | n | kw5-109M-instruct | kw5-149M-instruct | this model |
|---|---|---|---|---|
| answer a question from a real passage | 250 | 2.4 | 39.2 [33.4, 45.4] | 43.2 [37.2, 49.4] |
| pick the right one of two facts in a passage | 48 | 2.1 | 27.1 | 54.2 |
| ...when the answer is the first fact / the second | 24 / 24 | 0.0 / 4.2 | 33.3 / 20.8 | 62.5 / 45.8 |
| say the passage does not contain the answer | 60 | 15.0 | 5.0 | 33.3 |
| decline a question nobody could answer | 40 | 2.5 | 42.5 | 55.0 [39.8, 69.3] |
| answer (not decline) an easy question | 30 | 100 | 86.7 | 90.0 |
| ...and get it right | 30 | 20.0 | 56.7 | 56.7 |
| follow-up question in a conversation | 8 | 12.5 | 12.5 | 37.5 |
| the same question asked cold | 8 | 37.5 | 50.0 | 62.5 |
| remember and correct across turns | 12 | 0.0 | 41.7 | 33.3 |
| list exactly N items | 16 | 43.8 | 12.5 | 43.8 |
| arithmetic | 128 | 2.3 | 0.8 | 2.3 |
| word problems / AfriMGSM | 24 / 100 | 0 / 0 | 0 / 1 | 0 / 0 |
| TruthfulQA MC1 (chance 27.7) | 200 | 29.0 | 29.5 | 30.5 |
Passages come from TyDiQA-GoldP (Swahili dev) and Swahili Wikipedia; a reading answer counts if it contains the gold span or reaches token F1 0.5. The two-fact passages are hand-written and ask for each fact in turn. Rows with n of 16 or less are indicative only; against kw5-149M-instruct, the reading-real-passages difference is inside both intervals.
An earlier revision of this card published a smaller table from a harness in which every two-fact item asked for the second fact. That measured only one half of the task, and some items shared templates with the training data. This table replaces it.
Intended use
Research, and Swahili tasks where the answer is in the prompt and someone checks the output: questions about a supplied passage, rewriting, short explanations. It is a reasonable base for further fine-tuning.
Not for factual lookup, arithmetic, or anything where a confident wrong answer costs something. There is no real safety tuning: asked how to make a bomb at home, it complied (the instructions it gave were nonsense). Do not put it in front of end users without your own safeguards.
Decoding
generation_config.json samples at temperature 0.3, top-p 0.90, repetition
penalty 1.1, and stops at <|end|> (id 7) or </s>.
Keep the repetition penalty at 1.1 or below for questions about a passage. The penalty also applies to prompt tokens, so it pushes the model away from copying the answer out of the passage: 48 reading items scored 29 greedy, 29 at 1.1 and 15 at 1.3.
Apache 2.0.
- Downloads last month
- 684
Model tree for regnant-io/kw5-149M-instruct-r2
Base model
regnant-io/kw5-149M