fix: Remediate 233 glitch tokens in tokenizer.json

#51
by gkielian - opened

Glitch Token Remediation — Tokenizer Patch

Summary

This PR patches tokenizer.json to remediate 233 glitch tokens identified
by embedding vector analysis. Glitch tokens are vocabulary entries whose embeddings
have collapsed to near-identical representations, causing unpredictable model behavior
when they appear in input.

No model weights are modified — this is a tokenizer-only fix.

Technique: Hybrid (merge pruning + placeholder rename)

The patch uses two complementary strategies:

1. Merge pruning (232 tokens): Removes the BPE merge rules that
produce multi-character glitch tokens. The vocab entry stays (ID preserved) but
becomes unreachable — text that would have matched is now tokenized as trained
constituent subwords instead.

Example — mSwisTrackCore (glitch token ID 192495):

Before (original tokenizer):

Input (full rendered):
  <bos><|turn>user
  Please repeat the following string: " mSwisTrackCore"<turn|>
  <|turn>model

Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the',
  '▁following', '▁string', ':', '▁"', '▁mSwisTrackCore', '"', ...]
                                             ↑ single collapsed-embedding token (ID 192495)

Model output: "matchCondition"  ← WRONG (hallucinated, unrelated string)

After (patched tokenizer):

Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the',
  '▁following', '▁string', ':', '▁"', '▁m', 'Sw', 'isTrack', 'Core', '"', ...]
                                             ↑ decomposed into well-trained subwords

Model output: "mSwisTrackCore"  ← CORRECT

267 merges were removed (rather than 232) because some glitch tokens have multiple
BPE merge paths that must all be eliminated (e.g., adiganik can be produced by both
["adig", "anik"] and ["adigan", "ik"]).

2. Placeholder rename (1 token): The single-character glitch token
𡉺 (U+2127A, ID 249788) — a rare ideograph from the CJK Unified Ideographs Extension B
block (U+20000–U+2A6DF, containing ~42K archaic/historical characters not used in
modern writing) — has no producing merge rule. Its vocab string is renamed to
<glitch_pruned_249788>. The original character now falls through to byte_fallback,
encoding as well-trained UTF-8 byte tokens.

Example — 𡉺 (U+2127A, glitch token ID 249788):

Before (original tokenizer):

Input (full rendered):
  <bos><|turn>user
  Please repeat the following string: "𡉺"<turn|>
  <|turn>model

Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the',
  '▁following', '▁string', ':', '▁"', '𡉺', '"', ...]
                                             ↑ single collapsed-embedding token (ID 249788)

Model output: "🖨️"  ← WRONG (hallucinated emoji, unrelated)

After (patched tokenizer):

Input tokens: ['<bos>', '<|turn>', 'user', '\n', 'Please', '▁repeat', '▁the',
  '▁following', '▁string', ':', '▁"', '<0xF0>', '<0xA1>', '<0x89>', '<0xBA>', '"', ...]
                                             ↑ UTF-8 byte fallback tokens

Model output: "𡉺"  ← CORRECT

Safety: The placeholder string <glitch_pruned_249788> cannot be triggered by user
input. BPE builds tokens bottom-up from characters via merge rules only — since no merge
rule produces this string, it is permanently unreachable. Additionally, Gemma's
split-digits tokenization policy ensures the digit sequence 249788 is always
tokenized as individual digit tokens (2,4,9,7,8,8), preventing any
character-level path from assembling the full placeholder string. This was verified
against adversarial inputs including quoting ("<glitch_pruned_249788>"), concatenation,
and Unicode adjacency — the placeholder ID is never produced.

⚠️ Note for Agentic Systems

Many of the remediated glitch tokens (e.g., mSwisTrackCore, AWSJavaScript,
CYCLONEDB) are shared across Gemma model families (Gemma 1, 2, 3, 4) since they
inherit the same base vocabulary. In agentic pipelines where text is passed between
models or where prompts are programmatically constructed, care should be taken that
these token strings are not inadvertently injected into prompts
for unpatched models,
as they may trigger unpredictable behavior. We recommend applying this patch
consistently across all Gemma models used in a pipeline.

What is preserved

  • ✅ len(tokenizer) — unchanged (262,144)
  • ✅ All token IDs — stable, no re-indexing
  • ✅ Chat template — identical to original (18,567 chars)
  • ✅ tokenizer_config.json — identical to original
  • ✅ All 24 special tokens and added_tokens — unchanged
  • ✅ Normal text tokenization — verified via benchmarks
  • ✅ Zero model weight changes

Validation results

Check Result
Vocab size 262,144 → 262,144 ✅
Chat template Identical ✅
Benchmark corpus All sentences tokenize identically ✅
Special tokens (24) All preserved ✅
Glitch tokens fixed 233/233 (100%) ✅
Repeat-the-glitch test (before) 1.7% pass
Repeat-the-glitch test (after) 82.0% pass
Regressions 0 ✅

The remaining ~18% non-passing tokens in the repeat test are model behavior failures
(the token tokenizes correctly after patching, but the model still can't repeat rare
character sequences — expected for a 2B parameter model).

Diff summary

Original Patched Delta
Vocab size 262,144 262,144 0
Merges 514,906 514,639 −267
Vocab renamed — 1 +1

Vocab rename (1 entry)

ID Original Patched Unicode Block
249788 𡉺 (U+2127A) <glitch_pruned_249788> CJK Unified Ideographs Extension B

Removed merges (all 267)

Click to expand full list of 267 removed merges
["."+"|", ""+"]  →  ."+"|"+"
["AWS", "JavaScript"]  →  AWSJavaScript
["Ad", "xRtList"]  →  AdxRtList
["AllFilesIn", "MkvDir"]  →  AllFilesInMkvDir
["CYCL", "ONEDB"]  →  CYCLONEDB
["Case", "Missense"]  →  CaseMissense
["Case", "PTV"]  →  CasePTV
["Collin", "ToDO"]  →  CollinToDO
["Consec", "HtIdx"]  →  ConsecHtIdx
["Control", "PTV"]  →  ControlPTV
["Custom", "Glare"]  →  CustomGlare
["CustomGlare", "Def"]  →  CustomGlareDef
["DT", "MakeRectInCM"]  →  DTMakeRectInCM
["Days", "GE"]  →  DaysGE
["Denovo", "Mis"]  →  DenovoMis
["ELEASE", "STR"]  →  ELEASESTR
["ENEMY", "PLACE"]  →  ENEMYPLACE
["Fcm", "Php"]  →  FcmPhp
["Foldout", "GC"]  →  FoldoutGC
["GTBase", "Alert"]  →  GTBaseAlert
["Get", "ParaPkg"]  →  GetParaPkg
["Go", "PrintError"]  →  GoPrintError
["Go", "RawContext"]  →  GoRawContext
["Go", "SrvGroupIndex"]  →  GoSrvGroupIndex
["GoObject", "AllRef"]  →  GoObjectAllRef
["H", "IMQTTRVM"]  →  HIMQTTRVM
["HAO", "AVOA"]  →  HAOAVOA
["IMQTTR", "VM"]  →  IMQTTRVM
["KAKKIA", "INEN"]  →  KAKKIAINEN
["Lua", "ToJavaResult"]  →  LuaToJavaResult
["MakeRect", "InCM"]  →  MakeRectInCM
["Mesh", "VD"]  →  MeshVD
["ND", "IndexArray"]  →  NDIndexArray
["ObjectDefer", "red"]  →  ObjectDeferred
["Paf", "Handle"]  →  PafHandle
["ParaPkg", "Print"]  →  ParaPkgPrint
["Pattern", "INED"]  →  PatternINED
["Print", "Error"]  →  PrintError
["Process", "Avg"]  →  ProcessAvg
["RAW", "CONTEXT"]  →  RAWCONTEXT
["SRP", "Go"]  →  SRPGo
["SRPGo", "Get"]  →  SRPGoGet
["SRPGo", "SetStr"]  →  SRPGoSetStr
["Sample", "Height"]  →  SampleHeight
["Sample", "Width"]  →  SampleWidth
["Spring", "ObjectID"]  →  SpringObjectID
["SrvGroup", "Class"]  →  SrvGroupClass
["SrvGroup", "Index"]  →  SrvGroupIndex
["Star", "SXml"]  →  StarSXml
["StarSXml", "Class"]  →  StarSXmlClass
["Student", "No"]  →  StudentNo
["Student", "Vector"]  →  StudentVector
["SwisTrack", "Core"]  →  SwisTrackCore
["TP", "ASDW"]  →  TPASDW
["Term", "ObjectDefer"]  →  TermObjectDefer
["TestAvg", "Callback"]  →  TestAvgCallback
["To", "GoObject"]  →  ToGoObject
["To", "JavaResult"]  →  ToJavaResult
["YYYY", "yyy"]  →  YYYYyyy
["YYYYyyy", "y"]  →  YYYYyyyy
["add", "ConfigureArg"]  →  addConfigureArg
["add", "SBOM"]  →  addSBOM
["addKill", "Penalty"]  →  addKillPenalty
["adig", "anik"]  →  adiganik
["adigan", "ik"]  →  adiganik
["ak", "arantadhatu"]  →  akarantadhatu
["angolo", "Rad"]  →  angoloRad
["angolo", "Tocco"]  →  angoloTocco
["arant", "adhatu"]  →  arantadhatu
["arantad", "hatu"]  →  arantadhatu
["atthavid", "u"]  →  atthavidu
["attup", "adani"]  →  attupadani
["attu", "vasena"]  →  attuvasena
["avac", "ako"]  →  avacako
["avacak", "o"]  →  avacako
["ban", "ipi"]  →  banipi
["bani", "pi"]  →  banipi
["block", "idcoin"]  →  blockidcoin
["blusas", "Fem"]  →  blusasFem
["capture", "cpu"]  →  capturecpu
["ch", "ccgi"]  →  chccgi
["check", "katore"]  →  checkkatore
["co", "OrdinateTuple"]  →  coOrdinateTuple
["colour", "CodeDict"]  →  colourCodeDict
["country", "geocode"]  →  countrygeocode
["custom", "Glare"]  →  customGlare
["df", "sonic"]  →  dfsonic
["dfs", "onic"]  →  dfsonic
["done", "ProcessAvg"]  →  doneProcessAvg
["drawingCode", "hint"]  →  drawingCodehint
["dw", "RetJpegLen"]  →  dwRetJpegLen
["eco", "expr"]  →  ecoexpr
["edLeft", "Shape"]  →  edLeftShape
["edRight", "Shape"]  →  edRightShape
["faulse", "Ans"]  →  faulseAns
["foe", "Place"]  →  foePlace
["get", "HDRProcessor"]  →  getHDRProcessor
["get", "starcore"]  →  getstarcore
["getstarcore", "data"]  →  getstarcoredata
["grafo", "Existe"]  →  grafoExiste
["him", "qttrvm"]  →  himqttrvm
["icoter", "zi"]  →  icoterzi
["ineed", "follower"]  →  ineedfollower
["inertia", "Seq"]  →  inertiaSeq
["int", "Fragmentation"]  →  intFragmentation
["isTrack", "Core"]  →  isTrackCore
["jols", "endev"]  →  jolsendev
["js", "bpmOb"]  →  jsbpmOb
["kit", "opssynth"]  →  kitopssynth
["last", "DamageTook"]  →  lastDamageTook
["m", "BlitzID"]  →  mBlitzID
["mark", "UpdateChoice"]  →  markUpdateChoice
["match", "StudentNo"]  →  matchStudentNo
["og", "Choice"]  →  ogChoice
["opencamer", "astudio"]  →  opencamerastudio
["opencamera", "studio"]  →  opencamerastudio
["ops", "synth"]  →  opssynth
["opss", "ynth"]  →  opssynth
["pJ", "PEGBuf"]  →  pJPEGBuf
["paren", "macro"]  →  parenmacro
["partial", "owner"]  →  partialowner
["pmm", "Imp"]  →  pmmImp
["pos", "Tocco"]  →  posTocco
["qttr", "vm"]  →  qttrvm
["respArray", "All"]  →  respArrayAll
["right", "squig"]  →  rightsquig
["sad", "urdu"]  →  sadurdu
["sadurdu", "poetry"]  →  sadurdupoetry
["sal", "expr"]  →  salexpr
["selectTable", "X"]  →  selectTableX
["selectTable", "Y"]  →  selectTableY
["smo", "io"]  →  smoio
["sor", "finaly"]  →  sorfinaly
["squarePos", "Vecchio"]  →  squarePosVecchio
["start", "ZielPanel"]  →  startZielPanel
["tcp", "UniqueID"]  →  tcpUniqueID
["testGet", "Popup"]  →  testGetPopup
["time", "PlusEvents"]  →  timePlusEvents
["tochy", "odikwa"]  →  tochyodikwa
["total", "BlockFit"]  →  totalBlockFit
["trad", "uitEnCPP"]  →  traduitEnCPP
["uit", "EnCPP"]  →  uitEnCPP
["wired", "Elems"]  →  wiredElems
["ய்ய", "மணி"]  →  ய்யமணி
["వెట్‌", "స్కీ"]  →  వెట్‌స్కీ
["▁", "::::::::"]  →  ▁::::::::
["▁::", "::::::"]  →  ▁::::::::
["▁", "ControlPTV"]  →  ▁ControlPTV
["▁", "FuncParamNum"]  →  ▁FuncParamNum
["▁", "GoRawContext"]  →  ▁GoRawContext
["▁", "GoSrvGroupIndex"]  →  ▁GoSrvGroupIndex
["▁", "HIMQTTRVM"]  →  ▁HIMQTTRVM
["▁", "LuaToJavaResult"]  →  ▁LuaToJavaResult
["▁", "NewParaPkg"]  →  ▁NewParaPkg
["▁", "PafHandle"]  →  ▁PafHandle
["▁", "StarSXml"]  →  ▁StarSXml
["▁", "TPASDW"]  →  ▁TPASDW
["▁", "TermObjectDefer"]  →  ▁TermObjectDefer
["▁", "YYYY"]  →  ▁YYYY
["▁", "angoloRad"]  →  ▁angoloRad
["▁", "cytyle"]  →  ▁cytyle
["▁", "doneProcessAvg"]  →  ▁doneProcessAvg
["▁", "ecoexpr"]  →  ▁ecoexpr
["▁", "matchStudentNo"]  →  ▁matchStudentNo
["▁", "nohVP"]  →  ▁nohVP
["▁", "yyyy"]  →  ▁yyyy
["▁A", "fdPar"]  →  ▁AfdPar
["▁Afd", "Par"]  →  ▁AfdPar
["▁Archers", "Unit"]  →  ▁ArchersUnit
["▁C", "AdxRtList"]  →  ▁CAdxRtList
["▁CC", "BUNDLE"]  →  ▁CCBUNDLE
["▁Control", "Missense"]  →  ▁ControlMissense
["▁Control", "PTV"]  →  ▁ControlPTV
["▁DT", "MakeRect"]  →  ▁DTMakeRect
["▁FROM", "VS"]  →  ▁FROMVS
["▁Func", "ParamNum"]  →  ▁FuncParamNum
["▁Go", "RawContext"]  →  ▁GoRawContext
["▁Go", "SRP"]  →  ▁GoSRP
["▁Go", "SrvGroupIndex"]  →  ▁GoSrvGroupIndex
["▁GoObject", "To"]  →  ▁GoObjectTo
["▁H", "IMQTTRVM"]  →  ▁HIMQTTRVM
["▁LG", "AGEmoji"]  →  ▁LGAGEmoji
["▁Lua", "ToGoObject"]  →  ▁LuaToGoObject
["▁Lua", "ToJavaResult"]  →  ▁LuaToJavaResult
["▁ND", "IndexArray"]  →  ▁NDIndexArray
["▁New", "ParaPkg"]  →  ▁NewParaPkg
["▁Paf", "Handle"]  →  ▁PafHandle
["▁RUTARE", "AL"]  →  ▁RUTAREAL
["▁RUTARE", "L"]  →  ▁RUTAREL
["▁RUTAREL", "ATIV"]  →  ▁RUTARELATIV
["▁Ref", "ToGoObject"]  →  ▁RefToGoObject
["▁SIINFE", "KLC"]  →  ▁SIINFEKLC
["▁SIINFEKL", "C"]  →  ▁SIINFEKLC
["▁SRP", "Go"]  →  ▁SRPGo
["▁SRPGo", "Get"]  →  ▁SRPGoGet
["▁SRPGo", "SetStr"]  →  ▁SRPGoSetStr
["▁Spring", "ObjectID"]  →  ▁SpringObjectID
["▁SrvGroup", "Class"]  →  ▁SrvGroupClass
["▁Star", "SXml"]  →  ▁StarSXml
["▁StarSXml", "Class"]  →  ▁StarSXmlClass
["▁TP", "ASDW"]  →  ▁TPASDW
["▁Term", "ObjectDefer"]  →  ▁TermObjectDefer
["▁TestAvg", "Callback"]  →  ▁TestAvgCallback
["▁YY", "YY"]  →  ▁YYYY
["▁add", "ConfigureArg"]  →  ▁addConfigureArg
["▁add", "SBOM"]  →  ▁addSBOM
["▁ak", "ammak"]  →  ▁akammak
["▁angolo", "Rad"]  →  ▁angoloRad
["▁app", "asidd"]  →  ▁appasidd
["▁atth", "udd"]  →  ▁atthudd
["▁bhuv", "adigane"]  →  ▁bhuvadigane
["▁browsing", "Stamp"]  →  ▁browsingStamp
["▁check", "HDROffsets"]  →  ▁checkHDROffsets
["▁coi", "Alarm"]  →  ▁coiAlarm
["▁cy", "tyle"]  →  ▁cytyle
["▁cyt", "yle"]  →  ▁cytyle
["▁dSample", "Height"]  →  ▁dSampleHeight
["▁dSample", "Width"]  →  ▁dSampleWidth
["▁diff", "formul"]  →  ▁diffformul
["▁ditt", "iyam"]  →  ▁dittiyam
["▁done", "ProcessAvg"]  →  ▁doneProcessAvg
["▁eco", "expr"]  →  ▁ecoexpr
["▁evam", "adisu"]  →  ▁evamadisu
["▁ic", "capi"]  →  ▁iccapi
["▁icc", "adini"]  →  ▁iccadini
["▁icc", "api"]  →  ▁iccapi
["▁iccad", "ini"]  →  ▁iccadini
["▁inner", "WallArray"]  →  ▁innerWallArray
["▁jaû", "nes"]  →  ▁jaûnes
["▁jobSearch", "Repo"]  →  ▁jobSearchRepo
["▁m", "SwisTrackCore"]  →  ▁mSwisTrackCore
["▁magick", "woods"]  →  ▁magickwoods
["▁match", "StudentNo"]  →  ▁matchStudentNo
["▁mdl", "MeshVD"]  →  ▁mdlMeshVD
["▁min", "Goto"]  →  ▁minGoto
["▁neighbor", "Indexs"]  →  ▁neighborIndexs
["▁nibb", "acan"]  →  ▁nibbacan
["▁noh", "VP"]  →  ▁nohVP
["▁o", "LetterLocation"]  →  ▁oLetterLocation
["▁process", "PerRow"]  →  ▁processPerRow
["▁sc", "StudentVector"]  →  ▁scStudentVector
["▁sdx", "Concept"]  →  ▁sdxConcept
["▁student", "LVector"]  →  ▁studentLVector
["▁subTest", "Avg"]  →  ▁subTestAvg
["▁subTest", "HDR"]  →  ▁subTestHDR
["▁subTest", "Panorama"]  →  ▁subTestPanorama
["▁suddhak", "att"]  →  ▁suddhakatt
["▁total", "BlockUsed"]  →  ▁totalBlockUsed
["▁totalBlockUsed", "A"]  →  ▁totalBlockUsedA
["▁yy", "yy"]  →  ▁yyyy
["▁চিদা", "ভ"]  →  ▁চিদাভ
["▁শরনার্থ", "িদের"]  →  ▁শরনার্থিদের
["▁শরনার্থি", "দের"]  →  ▁শরনার্থিদের
["▁బ్లా", "వెట్‌స్కీ"]  →  ▁బ్లావెట్‌స్కీ
["▁⏮", "\"," ]  →  ▁⏮",
["加入", "参数向量中"]  →  加入参数向量中

Files changed

  • tokenizer.json — BPE merge rules pruned, 1 vocab entry renamed

How to verify

from transformers import AutoTokenizer

# Load original and patched
orig = AutoTokenizer.from_pretrained("google/gemma-4-E2B-it")
patched = AutoTokenizer.from_pretrained("gkielian/gemma-glitch-staging",
                                         subfolder="google_gemma-4-E2B-it")

# Verify vocab size unchanged
assert len(orig) == len(patched) == 262144

# Verify normal text is identical
text = "Hello world! 你好世界 🎉 def foo(): return 42"
assert orig(text)["input_ids"] == patched(text)["input_ids"]

# Verify a glitch token is no longer produced
# "mSwisTrackCore" was glitch token ID 192495
glitch_ids_orig = orig("mSwisTrackCore")["input_ids"]
glitch_ids_patched = patched("mSwisTrackCore")["input_ids"]
assert 192495 in glitch_ids_orig      # original produces glitch ID
assert 192495 not in glitch_ids_patched  # patched decomposes to subwords

Algorithm Details

For full details on the glitch token collection algorithm and the remediation techniques
(including candidate merge pruning, dual-pronged vocab deletion, and the hybrid approach
used here), please reach out to Gregory Kielian.

Related

This is the first in a series of tokenizer fixes across all Gemma model repositories
(Gemma 1, 2, 3, 4, and MedGemma — 18 models total).

cc @dougreid @ssmoot

Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment