Instructions to use suzuvoice/LMT-60-0.6B-mlx-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use suzuvoice/LMT-60-0.6B-mlx-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("suzuvoice/LMT-60-0.6B-mlx-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use suzuvoice/LMT-60-0.6B-mlx-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "suzuvoice/LMT-60-0.6B-mlx-8bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "suzuvoice/LMT-60-0.6B-mlx-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use suzuvoice/LMT-60-0.6B-mlx-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "suzuvoice/LMT-60-0.6B-mlx-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "suzuvoice/LMT-60-0.6B-mlx-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "suzuvoice/LMT-60-0.6B-mlx-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use suzuvoice/LMT-60-0.6B-mlx-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "suzuvoice/LMT-60-0.6B-mlx-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default suzuvoice/LMT-60-0.6B-mlx-8bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use suzuvoice/LMT-60-0.6B-mlx-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "suzuvoice/LMT-60-0.6B-mlx-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "suzuvoice/LMT-60-0.6B-mlx-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LMT-60-0.6B, MLX 8-bit
This is NiuTrans/LMT-60-0.6B converted to the MLX format and quantized to 8 bits. Nothing else was changed: same weights, same tokenizer, same chat template, same 60 languages.
It exists because Suzu, a macOS voice assistant, translates on device and needed a build that loads directly into MLX on Apple Silicon. No public MLX conversion of this checkpoint existed.
Attribution
The original model is the work of NiuTrans, released under the Apache License 2.0. All credit for the model itself goes to them. Please cite their paper rather than this repository:
- Model card: https://huggingface.co/NiuTrans/LMT-60-0.6B
- Paper: arXiv:2511.07003
How it was converted
Reproducible in one command, with mlx-lm 0.31.3:
mlx_lm.convert \
--hf-path NiuTrans/LMT-60-0.6B \
--mlx-path LMT-60-0.6B-mlx-8bit \
-q --q-bits 8 --q-group-size 64
Result: 8.501 bits per weight, 615 MiB on disk against 1.5 GB for the bf16 original.
LICENSE and the four tokenizer files that mlx_lm.convert does not copy
(added_tokens.json, special_tokens_map.json, merges.txt, vocab.json) were added by hand
from the upstream repository.
Why 8-bit and not 4-bit
Measured on an M3 Pro, same sentence, Chinese into French:
| Build | On disk | Per sentence | Control sentence |
|---|---|---|---|
| bf16 | 1500 MB | 0.40 s | "régions nordiques" ✅ |
| 8-bit | 645 MB | 0.29 s | "régions nordiques" ✅ |
| 4-bit | 347 MB | 0.21 s | "en Norvécie" ❌ invented word |
4-bit saves 300 MB and makes up words. 8-bit does not.
Prompt format
Use the chat template, role user, greedy sampling (temperature 0). Language names are written
out in English, not ISO codes:
Translate the following text from {SourceLanguage} into {TargetLanguage}:
{SourceLanguage}: {text}
{TargetLanguage}:
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("suzuvoice/LMT-60-0.6B-mlx-8bit")
prompt = tokenizer.apply_chat_template(
[{"role": "user", "content":
"Translate the following text from English into French:\n"
"English: Sales fell in the third quarter.\nFrench:"}],
add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256,
sampler=make_sampler(temp=0.0)))
Known limitation, inherited from the base model
LMT-60 is trained in a star around English and Chinese: 234 directions, not all 3540. A pair where neither side is English or Chinese can come out fluent and wrong. Measured example, Korean into French direct: "ont diminué de manière prévisible" where the source says unexpectedly, a reversed meaning no reader could catch.
Pivot through English whenever neither the source nor the target is English or Chinese. That fixes German, Russian, Korean and Japanese into French in our tests, at a cost of about 0.2 s per sentence. Arabic at this size stays unreliable even with the pivot: that is a limit of the 0.6B checkpoint, use a larger one if you need it.
- Downloads last month
- 55
8-bit