Text Generation
MLX
Safetensors
qwen3
conversational
8-bit precision

LMT-60-0.6B, MLX 8-bit

This is NiuTrans/LMT-60-0.6B converted to the MLX format and quantized to 8 bits. Nothing else was changed: same weights, same tokenizer, same chat template, same 60 languages.

It exists because Suzu, a macOS voice assistant, translates on device and needed a build that loads directly into MLX on Apple Silicon. No public MLX conversion of this checkpoint existed.

Attribution

The original model is the work of NiuTrans, released under the Apache License 2.0. All credit for the model itself goes to them. Please cite their paper rather than this repository:

How it was converted

Reproducible in one command, with mlx-lm 0.31.3:

mlx_lm.convert \
  --hf-path NiuTrans/LMT-60-0.6B \
  --mlx-path LMT-60-0.6B-mlx-8bit \
  -q --q-bits 8 --q-group-size 64

Result: 8.501 bits per weight, 615 MiB on disk against 1.5 GB for the bf16 original.

LICENSE and the four tokenizer files that mlx_lm.convert does not copy (added_tokens.json, special_tokens_map.json, merges.txt, vocab.json) were added by hand from the upstream repository.

Why 8-bit and not 4-bit

Measured on an M3 Pro, same sentence, Chinese into French:

Build On disk Per sentence Control sentence
bf16 1500 MB 0.40 s "régions nordiques" ✅
8-bit 645 MB 0.29 s "régions nordiques" ✅
4-bit 347 MB 0.21 s "en Norvécie" ❌ invented word

4-bit saves 300 MB and makes up words. 8-bit does not.

Prompt format

Use the chat template, role user, greedy sampling (temperature 0). Language names are written out in English, not ISO codes:

Translate the following text from {SourceLanguage} into {TargetLanguage}:
{SourceLanguage}: {text}
{TargetLanguage}:
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("suzuvoice/LMT-60-0.6B-mlx-8bit")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content":
      "Translate the following text from English into French:\n"
      "English: Sales fell in the third quarter.\nFrench:"}],
    add_generation_prompt=True,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=256,
               sampler=make_sampler(temp=0.0)))

Known limitation, inherited from the base model

LMT-60 is trained in a star around English and Chinese: 234 directions, not all 3540. A pair where neither side is English or Chinese can come out fluent and wrong. Measured example, Korean into French direct: "ont diminué de manière prévisible" where the source says unexpectedly, a reversed meaning no reader could catch.

Pivot through English whenever neither the source nor the target is English or Chinese. That fixes German, Russian, Korean and Japanese into French in our tests, at a cost of about 0.2 s per sentence. Arabic at this size stays unreliable even with the pivot: that is a limit of the 0.6B checkpoint, use a larger one if you need it.

Downloads last month
55
Safetensors
Model size
0.6B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for suzuvoice/LMT-60-0.6B-mlx-8bit

Quantized
(3)
this model

Dataset used to train suzuvoice/LMT-60-0.6B-mlx-8bit

Paper for suzuvoice/LMT-60-0.6B-mlx-8bit