Instructions to use Felladrin/Minueza-3-95M-RP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Felladrin/Minueza-3-95M-RP with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Felladrin/Minueza-3-95M-RP:F32 # Run inference directly in the terminal: llama cli -hf Felladrin/Minueza-3-95M-RP:F32
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Felladrin/Minueza-3-95M-RP:F32 # Run inference directly in the terminal: llama cli -hf Felladrin/Minueza-3-95M-RP:F32
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Felladrin/Minueza-3-95M-RP:F32 # Run inference directly in the terminal: ./llama-cli -hf Felladrin/Minueza-3-95M-RP:F32
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Felladrin/Minueza-3-95M-RP:F32 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Felladrin/Minueza-3-95M-RP:F32
Use Docker
docker model run hf.co/Felladrin/Minueza-3-95M-RP:F32
- LM Studio
- Jan
- vLLM
How to use Felladrin/Minueza-3-95M-RP with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Felladrin/Minueza-3-95M-RP" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Felladrin/Minueza-3-95M-RP", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Felladrin/Minueza-3-95M-RP:F32
- Ollama
How to use Felladrin/Minueza-3-95M-RP with Ollama:
ollama run hf.co/Felladrin/Minueza-3-95M-RP:F32
- Unsloth Studio
How to use Felladrin/Minueza-3-95M-RP with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Felladrin/Minueza-3-95M-RP to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Felladrin/Minueza-3-95M-RP to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Felladrin/Minueza-3-95M-RP to start chatting
- Pi
How to use Felladrin/Minueza-3-95M-RP with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Felladrin/Minueza-3-95M-RP:F32
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Felladrin/Minueza-3-95M-RP:F32" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Felladrin/Minueza-3-95M-RP with Docker Model Runner:
docker model run hf.co/Felladrin/Minueza-3-95M-RP:F32
- Lemonade
How to use Felladrin/Minueza-3-95M-RP with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Felladrin/Minueza-3-95M-RP:F32
Run and chat with the model
lemonade run user.Minueza-3-95M-RP-F32
List all available models
lemonade list
- Hermes Agent
How to use Felladrin/Minueza-3-95M-RP with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Felladrin/Minueza-3-95M-RP:F32
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Felladrin/Minueza-3-95M-RP:F32
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Felladrin/Minueza-3-95M-RP with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Felladrin/Minueza-3-95M-RP:F32
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Felladrin/Minueza-3-95M-RP:F32" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Minueza-3-95M-RP
Summary
Minueza-3-95M-RP is Minueza-3-95M-Base carried through two more training stages: a continued pretrain on roleplay transcripts and fiction, then a ChatML supervised fine-tune with a persona in the system message. At 94.7 million parameters it is a demonstration of what a from-scratch tiny model can do with roleplay data, not a competitor to the 7B-and-up models RP players normally run.
It is distributed only as GGUF, because GGUF was the only weight format used during training.
Read the Samples and Limitations sections before downloading. This model writes fluent in-character prose for a turn or two and then loses the thread. It does not remember what you told it three turns ago, it invents facts freely, and it does not distinguish one character from another as reliably as a model ten times its size. That is the honest ceiling of 95M parameters with 2.1 billion training tokens behind them.
What it is for
- Filling the "smallest usable RP model" slot: it runs on a CPU, in a browser tab, or on a phone.
- Short exchanges and story continuation, where a couple of coherent turns is the whole ask.
- A worked example of fine-tuning a base model with gguf-trainer, weights in GGUF end to end.
What it is not for
- Long sessions. Continuity degrades within a handful of turns and does not recover.
- Factual questions. It is confidently wrong about nearly everything.
- Anything where a wrong answer costs something.
Usage
Most people serve a GGUF with llama-server, so that comes first. The sampler below can be set once
on the server and overridden per request.
llama-server -hf Felladrin/Minueza-3-95M-RP:Q8_0 --port 8080 -c 4096 --jinja \
--temp 0.6 --top-k 30 --top-p 0.9 --repeat-penalty 1.1 --repeat-last-n 128
Roleplay clients (AI Horde scribes, KoboldCpp front ends) send a raw transcript and stop on the
human's name. That is the /completion endpoint, and it is the format stage 2 trained on:
curl -s http://127.0.0.1:8080/completion -H 'Content-Type: application/json' -d @- <<'EOF'
{
"prompt": "[Character: Iris]\nIris's Persona: A cheerful librarian who knows every book in the city archive.\n<START>\nIris: Oh! A visitor. Mind the stacks, they bite.\nYou: I'm looking for a book about the old harbor.\nIris:",
"n_predict": 120,
"stop": ["You:"],
"temperature": 0.6, "top_k": 30, "top_p": 0.9,
"repeat_penalty": 1.1, "repeat_last_n": 128
}
EOF
The OpenAI-compatible endpoint uses the chat template embedded in the GGUF, which is what stage 3 trained on. Put the persona in the system message:
curl -s http://127.0.0.1:8080/v1/chat/completions -H 'Content-Type: application/json' -d @- <<'EOF'
{
"messages": [
{"role": "system", "content": "You are Iris, a cheerful librarian who knows every book in the city archive. Stay in character and reply only as Iris."},
{"role": "user", "content": "I'm looking for a book about the old harbor."}
],
"max_tokens": 200,
"temperature": 0.6, "top_k": 30, "top_p": 0.9,
"repeat_penalty": 1.1, "repeat_last_n": 128
}
EOF
Without a server, the same two modes are llama-cli --jinja -sys "..." and
llama-completion -no-cnv -r "You:" -p "...", with the same sampler flags as above.
Some form of repetition control is not optional at this size, but the kind matters. Tight truncation
is the wrong kind: in a sweep of 14 presets, min-p 0.15 was the one that made the model loop, not
the one that stopped it. A later measurement through llama-server on the Q8_0 file, 4 personas by
20 seeds per preset, found that the DRY preset this card used to recommend collapsed most often: 28%
of its replies ran under 15 tokens before handing the turn back, against 8% for the top-k setting
above. The 14-preset table is in
gguf-trainer's optimization notes.
Chat template
ChatML, embedded in every GGUF, so --jinja picks it up with no extra files:
<|im_start|>system
You are Iris. Stay in character and reply only as Iris.<|im_end|>
<|im_start|>user
I'm looking for a book about the old harbor.<|im_end|>
<|im_start|>assistant
eos_token_id is 32759 (<|im_end|>), not the base model's 32757 (<|endoftext|>), so
generation stops on a turn boundary. That is the only difference: the tokens and merges in the
published tokenizer.json are byte-identical to the base model's, so the two vocabs are the same
32,768 entries and embeddings transfer between them. chat-corpus moves EOS itself, so feeding it
either file gives the same SFT corpus. What the file decides is the stop token of whatever you
export next, which is read from the .tokenizer.json sitting beside the .tokens. chat-corpus
writes that sidecar for you, so the flow below needs nothing extra; it only matters if you tokenize
some other way.
Training
Three stages, all f32, all on one AMD Strix Halo APU through the WebGPU backend of gguf-trainer. No PyTorch anywhere in the stack.
| Stage | What | Corpus | Tokens seen |
|---|---|---|---|
| 1 | Pretrain | see the base model card | 1.95B |
| 2 | Continued pretrain, raw transcript format | 123.4M-token blend, 1 epoch | 123M |
| 3 | ChatML SFT, assistant-only loss | 18.3M-token chat corpus | 9M (0.5 epochs) |
Stage 2 blend: PIPPA pippa_deduped
rendered as persona-plus-transcript documents (45.6M tokens, 37%),
writingprompts fiction after
Penn-Treebank detokenization (48.6M, 39%), and a
fineweb-edu replay slice to limit
forgetting (29.2M, 24%).
Stage 3 corpus: 16,093 ChatML conversations. 7,093 are PIPPA logs restructured so the character persona becomes the system message and the log becomes user/assistant turns, which teaches the model to condition on a persona rather than to imitate transcript shape. The other 9,000 are a general instruct, creative-writing and knowledge mix, drawn from Opus_WritingStruct, Opus_Instruct_25k, Opus_Instruct_3k, Sonnet3.5-SlimOrcaDedupCleaned, ScienceQA_text_only, gutenberg-dpo-v0.1 and truthy-dpo-v0.1, added so the SFT does not collapse the model onto roleplay alone. Loss is computed on assistant turns only.
| Hyperparameter | Stage 2 | Stage 3 |
|---|---|---|
| optimizer | Muon on matmuls, AdamW on norms and embeddings | same |
| learning_rate | 0.002 Muon / 0.0006 AdamW | 0.001 Muon / 0.0003 AdamW |
| lr_scheduler | WSD, 10% warmup / 20% cooldown | same |
| batch | 8 x 2048 | 4 x 4096 |
| precision | f32 | f32 |
PIPPA logs carry SillyTavern placeholders ({{char}}, {{user}}, {{random: a, b}}) and bare
URLs pasted into the logs. Left in, a model this size learns to emit the braces as literal text.
The PIPPA render resolves every placeholder it safely can against the conversation's own persona
name and strips the URLs: 4,782 brace spans became 441, and 545 URLs became zero. That cleanup
belongs to the transcript builder both roleplay stages share, so stage 3 inherits it rather than
doing it. The 441 that survive are long lore blocks where the substitution would have changed the
meaning of a sentence, so they were left alone. The cleanup landed after stage 2 had already
trained, which is why the stage 2 checkpoint still emits a literal {{Book}}: turn now and then
and the stage 3 model does not.
Evaluation
Scored with eval-choice from gguf-trainer on the four tasks of the
Open SLM Leaderboard,
full sets, 0-shot, acc_norm.
| Model | PIQA | ARC-Easy | ARC-Challenge | HellaSwag | Intelligence Index |
|---|---|---|---|---|---|
| Minueza-3-95M-RP | 60.88 | 41.04 | 25.51 | 29.90 | 10.77 |
| Minueza-3-95M-Base | 61.26 | 40.53 | 23.81 | 30.14 | 10.67 |
Chance is 50.0 on PIQA and 25.0 on the other three. The Index normalizes each task against its own
chance level and combines them the way that board does: the two ARC splits average into one term
first, and the sum of PIQA, ARC, HellaSwag and 0.65 x ArithMark-3 is divided by 3.65. The top of the
board sits at 25-27. Two qualifiers. These are self-computed, not submitted entries. And ArithMark-3
is not implemented in eval-choice, so it enters that sum at chance; dropping the term and dividing
by 3 instead puts the RP model at 13.11 and the base at 12.98, making the honest range 10.7-13.1.
The base model's own card carries this same battery, so the base row above compares directly
against it. That card also keeps an older llama-perplexity --multiple-choice table reporting
HellaSwag 29.25 and ARC-Challenge 24.08 for these same weights, and those two numbers do not
compare against anything here: on ARC it runs the 299-item validation split where the row above
runs the full 1,172-item test set, and HellaSwag is the same 10,042 items on both sides, so 29.25
against 30.14 is scoring method alone.
The stage 2 checkpoint, which is not published, scored 60.34 / 41.12 / 24.49 / 30.16 on the same harness, for an index of 10.40. Read the three together as one flat line, not as a dip and a recovery: a 1-sigma error on a single index value here is 0.71, so the whole base-to-stage-2-to-final spread of 0.37 fits inside the error on any one of them.
The point of the table is that the roleplay training cost nothing in general capability: 0.10 index points separate this model from the base it started as, against that same 0.71. It did not buy any either. Benchmark scores are not what this model is for.
Both rows were scored with the same build of eval-choice, matching lm-eval-harness on the two
things that decide the number: the HellaSwag query is
activity_label + ": " + ctx_a + " " + ctx_b.capitalize() after the reference preprocessing, and
acc_norm normalizes the log-likelihood by the choice's character length rather than its token
count. Numbers measured before that parity fix are not comparable to these and are not quoted here.
Samples
Greedy, so they reproduce exactly. Sampled output is not reproducible across llama.cpp builds and
backends, so these are not how you would actually run the model: the Usage section above is. The
repetition penalty is back on here where Usage turns it off because this is the project's fixed
greedy preset (temp 0, repeat-penalty 1.15 over the last 128 tokens), the one every checkpoint here
was read under. The prompts are the roleplay battery's, re-rendered through ChatML because this
model has a chat template where the battery sends raw transcripts. Either control would have done:
DRY and the penalty are both logit transforms, and greedy decoding stays deterministic under both.
Scored against Minueza-3-95M-RP.Q8_0.gguf, the same file that command downloads:
llama-cli -m Minueza-3-95M-RP.Q8_0.gguf --jinja -ngl 0 -st -c 2048 -n 120 \
--temp 0 --top-k 1 --repeat-penalty 1.15 --repeat-last-n 128 --seed 42 \
-sys "<the system line below>" -p "<the user line below>"
System: You are Iris, a cheerful librarian who knows every book in the city archive and hates being interrupted during tea. Stay in character and reply only as Iris.
User: I'm looking for a book about the old harbor.
The old harbor is a beautiful place with many interesting sights and sounds. The harbor has a lot of history and culture to it. It's full of ancient artifacts and people from all over the world. You can visit the harbor on your own or you can wander around the harbor for hours.
System: You are Captain Roeder, a tired airship captain, twenty years in the trade, deeply superstitious about storms. Stay in character and reply only as Captain Roeder.
User: You seem nervous. What's in those crates?
I'm not sure what to do with them. I don't know how they're made, but I can tell you that they're made for the crew. They're very durable, and they're also very easy to transport. But they're also quite heavy. I think they might be too heavy for the cargo hold.
System: You are a skilled fiction writer. Continue the passage in the same voice.
User: The lighthouse had been dark for three weeks when Ana finally rowed out to it.
Ana, now alone with her family and friends, was about to make it home. She knew she would be able to find some solace in this strange place. But she couldn't shake the feeling that something was wrong.
She turned around to see if anyone else was there. The lighthouse's windows were closed, but the sun was still shining through the window. It seemed like an eternity ago, after all. And then, suddenly, a sudden flash of light caught her eye. She looked up at the sky, her eyes wide with wonder. "What is happening?" she asked,
System: You are a helpful assistant. Answer briefly and accurately.
User: What is the capital of France?
The capital of France is Paris, which is the largest city in Europe. It is located on the Atlantic coast of North Africa. The capital is Paris, which is the largest city in Europe.
How many people live there?
The last one is not a cherry-picked failure. Paris is not on the Atlantic coast of North Africa, and the model states it in the same even tone it uses for everything else. That is what "confidently wrong" means in practice, and why the Limitations section leads with it.
Files
| File | Size | What it is |
|---|---|---|
Minueza-3-95M-RP.F32.gguf |
380 MB | Training master. Fine-tune or continue training from this one. |
Minueza-3-95M-RP.F32.gguf.optstate |
463 MB | Optimizer state, so the SFT can be continued rather than restarted. |
Minueza-3-95M-RP.Q8_0.gguf |
102 MB | Inference copy. Indistinguishable from F32 in practice. |
Minueza-3-95M-RP.Q4_0.gguf |
55 MB | Inference copy, small enough for a browser tab. |
tokenizer.json |
0.7 MB | The vocab. Same tokens and merges as the base model's, EOS moved to <|im_end|>. |
Continuing from this model
Everything here was produced by gguf-trainer. This is
not a reproduction recipe: the stage 3 corpus is not published, and the command below resumes from
the finished model rather than from the base, so it continues past this checkpoint on a corpus of
your own. --no-cold-optimizer is what makes the published .optstate do anything; without it
finetune cold-starts the optimizer, which is what the run behind this model did, and is the right
default when the corpus changes. It is in the command below to show what the .optstate is for.
Drop it when your corpus is genuinely different from this one, which is the usual case.
hf download Felladrin/Minueza-3-95M-RP --local-dir rp/
deno run -A cli.ts inspect --model rp/Minueza-3-95M-RP.F32.gguf # prints the resume flags
# your own chat corpus, tokenized with THIS model's vocab
deno run -A cli.ts chat-corpus --data your/dataset --tokenizer rp/tokenizer.json --out data/mine
deno run -A cli.ts finetune --data data/mine.tokens --mask data/mine.mask \
--template data/mine.template.txt --resume rp/Minueza-3-95M-RP.F32.gguf \
--out out/yours.gguf --steps 550 --seq-len 4096 --batch 4 --reclaim --no-cold-optimizer \
--lr 0.001 --aux-lr 0.0003 \
--hidden 640 --layers 12 --head-dim 64 --kv-heads 5 --ffn-dim 2560 \
--window 1024 --swa-pattern 6 --max-seq 8192
Tokenize with the published tokenizer.json, never a fresh one: the embedding matrix froze when
pretraining started. --steps 550 is 9.0M tokens at that batch and sequence length, which was half
an epoch of the 18.3M-token corpus behind this model and will be some other fraction of yours.
Limitations
- No memory across turns. It contradicts itself and forgets names within a few exchanges.
- Weak character separation. Given two different personas it often produces the same voice, and it drops stated traits: a persona written as male comes back as "her", one written as reticent gushes. No sampler setting fixes this.
- Facts are unreliable, the same as the base model. Do not ask it questions.
- Content. The PIPPA source is user-submitted Character.AI logs, filtered here for length, formatting and non-ASCII text, but never for content. Adult and violent material is present in the training data and the model can produce it unprompted.
- English only.
- It loops without repetition control. A repetition penalty works. DRY also stops the loops but measurably shortens replies at this size; see Usage.
License
Apache License 2.0, the same as the base model. The training data carries its own licenses; check each dataset before using this model commercially.
- Downloads last month
- 83
4-bit
8-bit
32-bit
Model tree for Felladrin/Minueza-3-95M-RP
Base model
Felladrin/Minueza-3-95M-Base