Kashmiri ByT5-small Diacritizer — checkpoint 58897

This repository is a safety release of the Kashmiri diacritizer training run completed so far. It contains the directly loadable model weights from the latest safe checkpoint plus training checkpoints, code, and run artifacts needed to resume or inspect the run.

Current status

  • Model: ByT5-small sequence-to-sequence diacritizer
  • Base / resume source: previous checkpoint /outputs/byt5small-ksdiac/checkpoints/checkpoint-6650
  • Latest safe checkpoint: checkpoint-58897
  • Training progress: 58,897 / 160,620 steps
  • Epoch: 11.0 / 30
  • Run: byt5small-ksdiac-l4-bf16-full
  • GPU: NVIDIA L4
  • Precision: bf16=True, fp16=False, tf32=True
  • Status: partial fine-tune, stopped during post-checkpoint test/eval phase; checkpoint itself is complete and resumable.

Repository layout

  • Root files: latest model files copied from checkpoint-58897 for direct from_pretrained() loading.
  • training-checkpoints/checkpoint-58897/: full latest trainer checkpoint, including optimizer/scheduler/RNG state.
  • training-checkpoints/checkpoint-37479/: earlier retained checkpoint from the same run.
  • artifacts/history.jsonl: train/eval log history.
  • artifacts/run_config.json: exact run configuration.
  • artifacts/dataset_stats.json: dataset filtering/split statistics.
  • source/: training/inference code used for this run.

Dataset summary

Input files were four XLSX sources mounted in Modal:

  • kashmiri_parallel_dataset_dia-non_dia_text_1.xlsx
  • kashmiri_parallel_dataset_dia-non_dia_text_2.xlsx
  • kashmiri_parallel_dataset_dia-non_dia_text_3.xlsx
  • kashmiri_parallel_dataset_dia-non_dia_text_4.xlsx

Filtering and split stats:

  • Rows in: 207,047
  • Dropped by length: 1,321
  • Dropped duplicates: 15,457
  • Rows kept: 190,269
  • Train rows: 171,344
  • Validation rows: 9,536 (1,500 max per epoch eval)
  • Test rows: 9,389

Training configuration

  • Epoch target: 30
  • Completed: 11 epochs
  • Train batch size: 4
  • Gradient accumulation: 8
  • Effective batch size: 32
  • Eval batch size: 1
  • Optimizer: adamw_torch
  • Learning rate: 5e-4
  • Scheduler: cosine
  • Warmup ratio: 0.05
  • Weight decay: 0.01
  • Gradient checkpointing: enabled
  • Group by length: enabled
  • Max source length: 256
  • Max target length: 256
  • Save total limit: 2

Loss track

Training loss reduced strongly and stably:

  • Step 50: 0.2546
  • Step 1,000: 0.0941
  • Step 5,000: 0.0624
  • Step 10,000: 0.0539
  • Step 16,000: 0.0492
  • Step 21,400: 0.0425
  • Step 26,700: 0.0382
  • Step 32,100: 0.0356
  • Step 37,400: 0.0308
  • Step 42,800: 0.0284
  • Step 48,100: 0.0250
  • Step 53,500: 0.0214
  • Step 58,850: 0.0178
  • Best observed train loss: 0.0154 at step 53,850

Approximate epoch train averages:

  • Epoch 1: 0.0575
  • Epoch 2: 0.0492
  • Epoch 3: 0.0434
  • Epoch 4: 0.0387
  • Epoch 5: 0.0346
  • Epoch 6: 0.0309
  • Epoch 7: 0.0273
  • Epoch 8: 0.0239
  • Epoch 9: 0.0205
  • Epoch 10: ~`0.017–0.019`
  • Epoch 11 boundary latest: 0.0178

Validation metrics by epoch

Lower DER/WER is better; higher exact match is better.

  • Epoch 1 / step 5354: eval_loss 0.0676, DER_marked 0.2378, DER_all 0.0702, WER 0.2176, exact 0.1300
  • Epoch 2 / step 10709: eval_loss 0.0655, DER_marked 0.2123, DER_all 0.0665, WER 0.2106, exact 0.1407
  • Epoch 3 / step 16062: eval_loss 0.0608, DER_marked 0.2284, DER_all 0.0638, WER 0.2033, exact 0.1460
  • Epoch 4 / step 21417: eval_loss 0.0585, DER_marked 0.1984, DER_all 0.0606, WER 0.1915, exact 0.1660
  • Epoch 5 / step 26770: eval_loss 0.0583, DER_marked 0.1961, DER_all 0.0591, WER 0.1868, exact 0.1713
  • Epoch 6 / step 32125: eval_loss 0.0584, DER_marked 0.1923, DER_all 0.0590, WER 0.1868, exact 0.1707
  • Epoch 7 / step 37479: eval_loss 0.0587, DER_marked 0.1844, DER_all 0.0586, WER 0.1856, exact 0.1747
  • Epoch 8 / step 42834: eval_loss 0.0599, DER_marked 0.1946, DER_all 0.0573, WER 0.1828, exact 0.1753
  • Epoch 9 / step 48188: eval_loss 0.0643, DER_marked 0.1927, DER_all 0.0572, WER 0.1842, exact 0.1860
  • Epoch 10 / step 53543: eval_loss 0.0663, DER_marked 0.2016, DER_all 0.0573, WER 0.1827, exact 0.1940
  • Epoch 11 / step 58897: eval_loss 0.0712, DER_marked 0.1870, DER_all 0.0570, WER 0.1818, exact 0.1753

How to use

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
import torch

repo_id = "Omarrran/koshur-diacritizer-byt5-small-v2"

tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSeq2SeqLM.from_pretrained(repo_id)
model.eval()

text = "یہاں اپنا غیر اعرابی کشمیری متن لکھیں"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=256)

with torch.no_grad():
    ids = model.generate(
        **inputs,
        max_new_tokens=256,
        num_beams=1,
        do_sample=False,
    )

print(tokenizer.decode(ids[0], skip_special_tokens=True))

How to resume training

The latest full Trainer checkpoint is:

training-checkpoints/checkpoint-58897

For a Hugging Face Seq2SeqTrainer continuation, restore/copy that folder as the training output checkpoint and call:

trainer.train(resume_from_checkpoint="training-checkpoints/checkpoint-58897")

If continuing on Modal, keep the same data orientation, all 4 XLSX sources, bf16/L4 settings, and output/run format used in artifacts/run_config.json.

Important caveats

  • This is not the final 30-epoch model. It is the best safe snapshot completed so far at epoch 11.
  • The last run stopped after the checkpoint was saved, during the longer test/eval phase. The checkpoint is structurally complete.
  • The source XLSX files themselves are not bundled here unless explicitly uploaded separately; the repo contains model/checkpoint/code/artifacts needed to use or resume when the dataset files are available.
Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Omarrran/koshur-diacritizer-byt5-small-v2

Finetuned
(334)
this model

Space using Omarrran/koshur-diacritizer-byt5-small-v2 1