Instructions to use google/gemma-4-E2B-it with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use google/gemma-4-E2B-it with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("google/gemma-4-E2B-it") model = AutoModelForMultimodalLM.from_pretrained("google/gemma-4-E2B-it", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- AMD Developer Cloud
fix: Remediate 233 glitch tokens in tokenizer.json
Glitch Token Remediation — Tokenizer Patch
Summary
This PR patches tokenizer.json to remediate 233 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.
No model weights are modified — this is a tokenizer-only fix.
Technique: Hybrid (merge pruning + placeholder rename)
The patch uses two complementary strategies:
1. Merge pruning (232 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.
Example —
mSwisTrackCore(glitch token ID 192495):Before (original tokenizer):
Input (full rendered): <bos><|turn>user Please repeat the following string: " mSwisTrackCore"<turn|> <|turn>model Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the', '▁following', '▁string', ':', '▁"', '▁mSwisTrackCore', '"', ...] ↑ single collapsed-embedding token (ID 192495) Model output: "matchCondition" ← WRONG (hallucinated, unrelated string)After (patched tokenizer):
Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the', '▁following', '▁string', ':', '▁"', '▁m', 'Sw', 'isTrack', 'Core', '"', ...] ↑ decomposed into well-trained subwords Model output: "mSwisTrackCore" ← CORRECT
267 merges were removed (rather than 232) because some glitch tokens have multiple
BPE merge paths that must all be eliminated (e.g., adiganik can be produced by both["adig", "anik"] and ["adigan", "ik"]).
2. Placeholder rename (1 token): The single-character glitch token𡉺 (U+2127A, ID 249788) — a rare ideograph from the CJK Unified Ideographs Extension B
block (U+20000–U+2A6DF, containing ~42K archaic/historical characters not used in
modern writing) — has no producing merge rule. Its vocab string is renamed to<glitch_pruned_249788>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.
Example —
𡉺(U+2127A, glitch token ID 249788):Before (original tokenizer):
Input (full rendered): <bos><|turn>user Please repeat the following string: "𡉺"<turn|> <|turn>model Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the', '▁following', '▁string', ':', '▁"', '𡉺', '"', ...] ↑ single collapsed-embedding token (ID 249788) Model output: "🖨️" ← WRONG (hallucinated emoji, unrelated)After (patched tokenizer):
Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the', '▁following', '▁string', ':', '▁"', '<0xF0>', '<0xA1>', '<0x89>', '<0xBA>', '"', ...] ↑ UTF-8 byte fallback tokens Model output: "𡉺" ← CORRECT
Safety: The placeholder string <glitch_pruned_249788> cannot be triggered by user
input. BPE builds tokens bottom-up from characters via merge rules only — since no merge
rule produces this string, it is permanently unreachable. Additionally, Gemma's
split-digits tokenization policy ensures the digit sequence 249788 is always
tokenized as individual digit tokens (2,4,9,7,8,8), preventing any
character-level path from assembling the full placeholder string. This was verified
against adversarial inputs including quoting ("<glitch_pruned_249788>"), concatenation,
and Unicode adjacency — the placeholder ID is never produced.
⚠️ Note for Agentic Systems
Many of the remediated glitch tokens (e.g., mSwisTrackCore, AWSJavaScript,CYCLONEDB) are shared across Gemma model families (Gemma 1, 2, 3, 4) since they
inherit the same base vocabulary. In agentic pipelines where text is passed between
models or where prompts are programmatically constructed, care should be taken that
these token strings are not inadvertently injected into prompts for unpatched models,
as they may trigger unpredictable behavior. We recommend applying this patch
consistently across all Gemma models used in a pipeline.
What is preserved
- ✅
len(tokenizer)— unchanged (262,144) - ✅ All token IDs — stable, no re-indexing
- ✅ Chat template — identical to original (18,567 chars)
- ✅
tokenizer_config.json— identical to original - ✅ All 24 special tokens and
added_tokens— unchanged - ✅ Normal text tokenization — verified via benchmarks
- ✅ Zero model weight changes
Validation results
| Check | Result |
|---|---|
| Vocab size | 262,144 → 262,144 ✅ |
| Chat template | Identical ✅ |
| Benchmark corpus | All sentences tokenize identically ✅ |
| Special tokens (24) | All preserved ✅ |
| Glitch tokens fixed | 233/233 (100%) ✅ |
| Repeat-the-glitch test (before) | 1.7% pass |
| Repeat-the-glitch test (after) | 82.0% pass |
| Regressions | 0 ✅ |
The remaining ~18% non-passing tokens in the repeat test are model behavior failures
(the token tokenizes correctly after patching, but the model still can't repeat rare
character sequences — expected for a 2B parameter model).
Diff summary
| Original | Patched | Delta | |
|---|---|---|---|
| Vocab size | 262,144 | 262,144 | 0 |
| Merges | 514,906 | 514,639 | −267 |
| Vocab renamed | — | 1 | +1 |
Vocab rename (1 entry)
| ID | Original | Patched | Unicode Block |
|---|---|---|---|
| 249788 | 𡉺 (U+2127A) |
<glitch_pruned_249788> |
CJK Unified Ideographs Extension B |
Removed merges (all 267)
Click to expand full list of 267 removed merges
["."+"|", ""+"] → ."+"|"+"
["AWS", "JavaScript"] → AWSJavaScript
["Ad", "xRtList"] → AdxRtList
["AllFilesIn", "MkvDir"] → AllFilesInMkvDir
["CYCL", "ONEDB"] → CYCLONEDB
["Case", "Missense"] → CaseMissense
["Case", "PTV"] → CasePTV
["Collin", "ToDO"] → CollinToDO
["Consec", "HtIdx"] → ConsecHtIdx
["Control", "PTV"] → ControlPTV
["Custom", "Glare"] → CustomGlare
["CustomGlare", "Def"] → CustomGlareDef
["DT", "MakeRectInCM"] → DTMakeRectInCM
["Days", "GE"] → DaysGE
["Denovo", "Mis"] → DenovoMis
["ELEASE", "STR"] → ELEASESTR
["ENEMY", "PLACE"] → ENEMYPLACE
["Fcm", "Php"] → FcmPhp
["Foldout", "GC"] → FoldoutGC
["GTBase", "Alert"] → GTBaseAlert
["Get", "ParaPkg"] → GetParaPkg
["Go", "PrintError"] → GoPrintError
["Go", "RawContext"] → GoRawContext
["Go", "SrvGroupIndex"] → GoSrvGroupIndex
["GoObject", "AllRef"] → GoObjectAllRef
["H", "IMQTTRVM"] → HIMQTTRVM
["HAO", "AVOA"] → HAOAVOA
["IMQTTR", "VM"] → IMQTTRVM
["KAKKIA", "INEN"] → KAKKIAINEN
["Lua", "ToJavaResult"] → LuaToJavaResult
["MakeRect", "InCM"] → MakeRectInCM
["Mesh", "VD"] → MeshVD
["ND", "IndexArray"] → NDIndexArray
["ObjectDefer", "red"] → ObjectDeferred
["Paf", "Handle"] → PafHandle
["ParaPkg", "Print"] → ParaPkgPrint
["Pattern", "INED"] → PatternINED
["Print", "Error"] → PrintError
["Process", "Avg"] → ProcessAvg
["RAW", "CONTEXT"] → RAWCONTEXT
["SRP", "Go"] → SRPGo
["SRPGo", "Get"] → SRPGoGet
["SRPGo", "SetStr"] → SRPGoSetStr
["Sample", "Height"] → SampleHeight
["Sample", "Width"] → SampleWidth
["Spring", "ObjectID"] → SpringObjectID
["SrvGroup", "Class"] → SrvGroupClass
["SrvGroup", "Index"] → SrvGroupIndex
["Star", "SXml"] → StarSXml
["StarSXml", "Class"] → StarSXmlClass
["Student", "No"] → StudentNo
["Student", "Vector"] → StudentVector
["SwisTrack", "Core"] → SwisTrackCore
["TP", "ASDW"] → TPASDW
["Term", "ObjectDefer"] → TermObjectDefer
["TestAvg", "Callback"] → TestAvgCallback
["To", "GoObject"] → ToGoObject
["To", "JavaResult"] → ToJavaResult
["YYYY", "yyy"] → YYYYyyy
["YYYYyyy", "y"] → YYYYyyyy
["add", "ConfigureArg"] → addConfigureArg
["add", "SBOM"] → addSBOM
["addKill", "Penalty"] → addKillPenalty
["adig", "anik"] → adiganik
["adigan", "ik"] → adiganik
["ak", "arantadhatu"] → akarantadhatu
["angolo", "Rad"] → angoloRad
["angolo", "Tocco"] → angoloTocco
["arant", "adhatu"] → arantadhatu
["arantad", "hatu"] → arantadhatu
["atthavid", "u"] → atthavidu
["attup", "adani"] → attupadani
["attu", "vasena"] → attuvasena
["avac", "ako"] → avacako
["avacak", "o"] → avacako
["ban", "ipi"] → banipi
["bani", "pi"] → banipi
["block", "idcoin"] → blockidcoin
["blusas", "Fem"] → blusasFem
["capture", "cpu"] → capturecpu
["ch", "ccgi"] → chccgi
["check", "katore"] → checkkatore
["co", "OrdinateTuple"] → coOrdinateTuple
["colour", "CodeDict"] → colourCodeDict
["country", "geocode"] → countrygeocode
["custom", "Glare"] → customGlare
["df", "sonic"] → dfsonic
["dfs", "onic"] → dfsonic
["done", "ProcessAvg"] → doneProcessAvg
["drawingCode", "hint"] → drawingCodehint
["dw", "RetJpegLen"] → dwRetJpegLen
["eco", "expr"] → ecoexpr
["edLeft", "Shape"] → edLeftShape
["edRight", "Shape"] → edRightShape
["faulse", "Ans"] → faulseAns
["foe", "Place"] → foePlace
["get", "HDRProcessor"] → getHDRProcessor
["get", "starcore"] → getstarcore
["getstarcore", "data"] → getstarcoredata
["grafo", "Existe"] → grafoExiste
["him", "qttrvm"] → himqttrvm
["icoter", "zi"] → icoterzi
["ineed", "follower"] → ineedfollower
["inertia", "Seq"] → inertiaSeq
["int", "Fragmentation"] → intFragmentation
["isTrack", "Core"] → isTrackCore
["jols", "endev"] → jolsendev
["js", "bpmOb"] → jsbpmOb
["kit", "opssynth"] → kitopssynth
["last", "DamageTook"] → lastDamageTook
["m", "BlitzID"] → mBlitzID
["mark", "UpdateChoice"] → markUpdateChoice
["match", "StudentNo"] → matchStudentNo
["og", "Choice"] → ogChoice
["opencamer", "astudio"] → opencamerastudio
["opencamera", "studio"] → opencamerastudio
["ops", "synth"] → opssynth
["opss", "ynth"] → opssynth
["pJ", "PEGBuf"] → pJPEGBuf
["paren", "macro"] → parenmacro
["partial", "owner"] → partialowner
["pmm", "Imp"] → pmmImp
["pos", "Tocco"] → posTocco
["qttr", "vm"] → qttrvm
["respArray", "All"] → respArrayAll
["right", "squig"] → rightsquig
["sad", "urdu"] → sadurdu
["sadurdu", "poetry"] → sadurdupoetry
["sal", "expr"] → salexpr
["selectTable", "X"] → selectTableX
["selectTable", "Y"] → selectTableY
["smo", "io"] → smoio
["sor", "finaly"] → sorfinaly
["squarePos", "Vecchio"] → squarePosVecchio
["start", "ZielPanel"] → startZielPanel
["tcp", "UniqueID"] → tcpUniqueID
["testGet", "Popup"] → testGetPopup
["time", "PlusEvents"] → timePlusEvents
["tochy", "odikwa"] → tochyodikwa
["total", "BlockFit"] → totalBlockFit
["trad", "uitEnCPP"] → traduitEnCPP
["uit", "EnCPP"] → uitEnCPP
["wired", "Elems"] → wiredElems
["ய்ய", "மணி"] → ய்யமணி
["వెట్", "స్కీ"] → వెట్స్కీ
["▁", "::::::::"] → ▁::::::::
["▁::", "::::::"] → ▁::::::::
["▁", "ControlPTV"] → ▁ControlPTV
["▁", "FuncParamNum"] → ▁FuncParamNum
["▁", "GoRawContext"] → ▁GoRawContext
["▁", "GoSrvGroupIndex"] → ▁GoSrvGroupIndex
["▁", "HIMQTTRVM"] → ▁HIMQTTRVM
["▁", "LuaToJavaResult"] → ▁LuaToJavaResult
["▁", "NewParaPkg"] → ▁NewParaPkg
["▁", "PafHandle"] → ▁PafHandle
["▁", "StarSXml"] → ▁StarSXml
["▁", "TPASDW"] → ▁TPASDW
["▁", "TermObjectDefer"] → ▁TermObjectDefer
["▁", "YYYY"] → ▁YYYY
["▁", "angoloRad"] → ▁angoloRad
["▁", "cytyle"] → ▁cytyle
["▁", "doneProcessAvg"] → ▁doneProcessAvg
["▁", "ecoexpr"] → ▁ecoexpr
["▁", "matchStudentNo"] → ▁matchStudentNo
["▁", "nohVP"] → ▁nohVP
["▁", "yyyy"] → ▁yyyy
["▁A", "fdPar"] → ▁AfdPar
["▁Afd", "Par"] → ▁AfdPar
["▁Archers", "Unit"] → ▁ArchersUnit
["▁C", "AdxRtList"] → ▁CAdxRtList
["▁CC", "BUNDLE"] → ▁CCBUNDLE
["▁Control", "Missense"] → ▁ControlMissense
["▁Control", "PTV"] → ▁ControlPTV
["▁DT", "MakeRect"] → ▁DTMakeRect
["▁FROM", "VS"] → ▁FROMVS
["▁Func", "ParamNum"] → ▁FuncParamNum
["▁Go", "RawContext"] → ▁GoRawContext
["▁Go", "SRP"] → ▁GoSRP
["▁Go", "SrvGroupIndex"] → ▁GoSrvGroupIndex
["▁GoObject", "To"] → ▁GoObjectTo
["▁H", "IMQTTRVM"] → ▁HIMQTTRVM
["▁LG", "AGEmoji"] → ▁LGAGEmoji
["▁Lua", "ToGoObject"] → ▁LuaToGoObject
["▁Lua", "ToJavaResult"] → ▁LuaToJavaResult
["▁ND", "IndexArray"] → ▁NDIndexArray
["▁New", "ParaPkg"] → ▁NewParaPkg
["▁Paf", "Handle"] → ▁PafHandle
["▁RUTARE", "AL"] → ▁RUTAREAL
["▁RUTARE", "L"] → ▁RUTAREL
["▁RUTAREL", "ATIV"] → ▁RUTARELATIV
["▁Ref", "ToGoObject"] → ▁RefToGoObject
["▁SIINFE", "KLC"] → ▁SIINFEKLC
["▁SIINFEKL", "C"] → ▁SIINFEKLC
["▁SRP", "Go"] → ▁SRPGo
["▁SRPGo", "Get"] → ▁SRPGoGet
["▁SRPGo", "SetStr"] → ▁SRPGoSetStr
["▁Spring", "ObjectID"] → ▁SpringObjectID
["▁SrvGroup", "Class"] → ▁SrvGroupClass
["▁Star", "SXml"] → ▁StarSXml
["▁StarSXml", "Class"] → ▁StarSXmlClass
["▁TP", "ASDW"] → ▁TPASDW
["▁Term", "ObjectDefer"] → ▁TermObjectDefer
["▁TestAvg", "Callback"] → ▁TestAvgCallback
["▁YY", "YY"] → ▁YYYY
["▁add", "ConfigureArg"] → ▁addConfigureArg
["▁add", "SBOM"] → ▁addSBOM
["▁ak", "ammak"] → ▁akammak
["▁angolo", "Rad"] → ▁angoloRad
["▁app", "asidd"] → ▁appasidd
["▁atth", "udd"] → ▁atthudd
["▁bhuv", "adigane"] → ▁bhuvadigane
["▁browsing", "Stamp"] → ▁browsingStamp
["▁check", "HDROffsets"] → ▁checkHDROffsets
["▁coi", "Alarm"] → ▁coiAlarm
["▁cy", "tyle"] → ▁cytyle
["▁cyt", "yle"] → ▁cytyle
["▁dSample", "Height"] → ▁dSampleHeight
["▁dSample", "Width"] → ▁dSampleWidth
["▁diff", "formul"] → ▁diffformul
["▁ditt", "iyam"] → ▁dittiyam
["▁done", "ProcessAvg"] → ▁doneProcessAvg
["▁eco", "expr"] → ▁ecoexpr
["▁evam", "adisu"] → ▁evamadisu
["▁ic", "capi"] → ▁iccapi
["▁icc", "adini"] → ▁iccadini
["▁icc", "api"] → ▁iccapi
["▁iccad", "ini"] → ▁iccadini
["▁inner", "WallArray"] → ▁innerWallArray
["▁jaû", "nes"] → ▁jaûnes
["▁jobSearch", "Repo"] → ▁jobSearchRepo
["▁m", "SwisTrackCore"] → ▁mSwisTrackCore
["▁magick", "woods"] → ▁magickwoods
["▁match", "StudentNo"] → ▁matchStudentNo
["▁mdl", "MeshVD"] → ▁mdlMeshVD
["▁min", "Goto"] → ▁minGoto
["▁neighbor", "Indexs"] → ▁neighborIndexs
["▁nibb", "acan"] → ▁nibbacan
["▁noh", "VP"] → ▁nohVP
["▁o", "LetterLocation"] → ▁oLetterLocation
["▁process", "PerRow"] → ▁processPerRow
["▁sc", "StudentVector"] → ▁scStudentVector
["▁sdx", "Concept"] → ▁sdxConcept
["▁student", "LVector"] → ▁studentLVector
["▁subTest", "Avg"] → ▁subTestAvg
["▁subTest", "HDR"] → ▁subTestHDR
["▁subTest", "Panorama"] → ▁subTestPanorama
["▁suddhak", "att"] → ▁suddhakatt
["▁total", "BlockUsed"] → ▁totalBlockUsed
["▁totalBlockUsed", "A"] → ▁totalBlockUsedA
["▁yy", "yy"] → ▁yyyy
["▁চিদা", "ভ"] → ▁চিদাভ
["▁শরনার্থ", "িদের"] → ▁শরনার্থিদের
["▁শরনার্থি", "দের"] → ▁শরনার্থিদের
["▁బ్లా", "వెట్స్కీ"] → ▁బ్లావెట్స్కీ
["▁⏮", "\"," ] → ▁⏮",
["加入", "参数向量中"] → 加入参数向量中
Files changed
tokenizer.json— BPE merge rules pruned, 1 vocab entry renamed
How to verify
from transformers import AutoTokenizer
# Load original and patched
orig = AutoTokenizer.from_pretrained("google/gemma-4-E2B-it")
patched = AutoTokenizer.from_pretrained("gkielian/gemma-glitch-staging",
subfolder="google_gemma-4-E2B-it")
# Verify vocab size unchanged
assert len(orig) == len(patched) == 262144
# Verify normal text is identical
text = "Hello world! 你好世界 🎉 def foo(): return 42"
assert orig(text)["input_ids"] == patched(text)["input_ids"]
# Verify a glitch token is no longer produced
# "mSwisTrackCore" was glitch token ID 192495
glitch_ids_orig = orig("mSwisTrackCore")["input_ids"]
glitch_ids_patched = patched("mSwisTrackCore")["input_ids"]
assert 192495 in glitch_ids_orig # original produces glitch ID
assert 192495 not in glitch_ids_patched # patched decomposes to subwords
Algorithm Details
For full details on the glitch token collection algorithm and the remediation techniques
(including candidate merge pruning, dual-pronged vocab deletion, and the hybrid approach
used here), please reach out to Gregory Kielian.
Related
This is the first in a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).