gemma-4-26B-A4B-it EAGLE-3 Draft (Korean-optimized)

An EAGLE-3 draft (speculator) for accelerating Korean generation from gemma-4-26B-A4B-it (Gemma 4 MoE; 26B total, ~4B active) via speculative decoding. The public RedHat draft is English-only (Magpie + UltraChat) and accepts Korean tokens poorly, so this draft was retrained on Korean prompts with on-policy responses regenerated by the verifier.

  • Method: EAGLE-3 (vllm-project/speculators)
  • Verifier (target): RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic (FP8 MoE; serving + hidden-state extraction)
  • Warm start: RedHatAI/gemma-4-26B-A4B-it-speculator.eagle3
  • Training data: ~150k prompts sampled from sh2orc/bccard-maywell-jojo0217-markai-lcw99-kendamarron-microsoft (1.71M-row Korean QA; instruction column only). Answers regenerated on-policy by the verifier (text-only).

Serving (vLLM)

VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic -tp 1 \
  --max-model-len 4096 \
  --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
  --speculative-config '{
    "model": "<this-repo-id>",
    "num_speculative_tokens": 3,
    "method": "eagle3",
    "draft_tensor_parallel_size": 1
  }'

Tune num_speculative_tokens (3–6) on measured acceptance / TPS. The draft uses the verifier's tokenizer and chat template. The 26B-A4B is multimodal; for text workloads pass --limit-mm-per-prompt to disable image/audio/video.

Notes for the MoE target

  • The verifier is a Mixture-of-Experts model (128 experts, top-8). On some Blackwell parts (sm_120 / sm_121) the FP8 fused-MoE Triton kernel can exceed shared memory; if serving fails with an out of resource: shared memory error, use a vLLM MoE fallback config or the NVFP4 path.
  • Acceptance is measured against RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic. Pairing with a different target (or a base, non--it model) will change results.

Limitations

  • Trained on general Korean QA. Domain-specific traffic (e.g. finance) may benefit from another training cycle on domain-matched data.

License

Apache 2.0. The base Gemma 4 (Apache 2.0), the RedHat FP8 verifier (Apache 2.0), and the RedHat EAGLE-3 warm-start checkpoint (Apache 2.0) are all Apache 2.0, so this draft is released under Apache 2.0 as well. (Informational, not legal advice.)

Downloads last month
35
Safetensors
Model size
0.9B params
Tensor type
I64
·
BF16
·
BOOL
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BCCard/MoAI-gemma-4-26B-A4B-it-speculator.eagle3

Finetuned
(1)
this model

Dataset used to train BCCard/MoAI-gemma-4-26B-A4B-it-speculator.eagle3

Collection including BCCard/MoAI-gemma-4-26B-A4B-it-speculator.eagle3