gemma-4-26B-A4B-it EAGLE-3 Draft (Korean-optimized)
An EAGLE-3 draft (speculator) for accelerating Korean generation from
gemma-4-26B-A4B-it (Gemma 4 MoE; 26B total, ~4B active) via speculative decoding.
The public RedHat draft is English-only (Magpie + UltraChat) and accepts Korean
tokens poorly, so this draft was retrained on Korean prompts with on-policy
responses regenerated by the verifier.
- Method: EAGLE-3 (
vllm-project/speculators) - Verifier (target):
RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic(FP8 MoE; serving + hidden-state extraction) - Warm start:
RedHatAI/gemma-4-26B-A4B-it-speculator.eagle3 - Training data: ~150k prompts sampled from
sh2orc/bccard-maywell-jojo0217-markai-lcw99-kendamarron-microsoft
(1.71M-row Korean QA;
instructioncolumn only). Answers regenerated on-policy by the verifier (text-only).
Serving (vLLM)
VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic -tp 1 \
--max-model-len 4096 \
--limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
--speculative-config '{
"model": "<this-repo-id>",
"num_speculative_tokens": 3,
"method": "eagle3",
"draft_tensor_parallel_size": 1
}'
Tune num_speculative_tokens (3–6) on measured acceptance / TPS. The draft uses the
verifier's tokenizer and chat template. The 26B-A4B is multimodal; for text workloads
pass --limit-mm-per-prompt to disable image/audio/video.
Notes for the MoE target
- The verifier is a Mixture-of-Experts model (128 experts, top-8). On some Blackwell
parts (sm_120 / sm_121) the FP8 fused-MoE Triton kernel can exceed shared memory;
if serving fails with an
out of resource: shared memoryerror, use a vLLM MoE fallback config or the NVFP4 path. - Acceptance is measured against
RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic. Pairing with a different target (or a base, non--itmodel) will change results.
Limitations
- Trained on general Korean QA. Domain-specific traffic (e.g. finance) may benefit from another training cycle on domain-matched data.
License
Apache 2.0. The base Gemma 4 (Apache 2.0), the RedHat FP8 verifier (Apache 2.0), and the RedHat EAGLE-3 warm-start checkpoint (Apache 2.0) are all Apache 2.0, so this draft is released under Apache 2.0 as well. (Informational, not legal advice.)
- Downloads last month
- 35
Model tree for BCCard/MoAI-gemma-4-26B-A4B-it-speculator.eagle3
Base model
google/gemma-4-26B-A4B