Kashmiri ByT5-small Diacritizer — checkpoint 58897
This repository is a safety release of the Kashmiri diacritizer training run completed so far. It contains the directly loadable model weights from the latest safe checkpoint plus training checkpoints, code, and run artifacts needed to resume or inspect the run.
Current status
- Model: ByT5-small sequence-to-sequence diacritizer
- Base / resume source: previous checkpoint
/outputs/byt5small-ksdiac/checkpoints/checkpoint-6650 - Latest safe checkpoint:
checkpoint-58897 - Training progress:
58,897 / 160,620steps - Epoch:
11.0 / 30 - Run:
byt5small-ksdiac-l4-bf16-full - GPU: NVIDIA L4
- Precision:
bf16=True,fp16=False,tf32=True - Status: partial fine-tune, stopped during post-checkpoint test/eval phase; checkpoint itself is complete and resumable.
Repository layout
- Root files: latest model files copied from
checkpoint-58897for directfrom_pretrained()loading. training-checkpoints/checkpoint-58897/: full latest trainer checkpoint, including optimizer/scheduler/RNG state.training-checkpoints/checkpoint-37479/: earlier retained checkpoint from the same run.artifacts/history.jsonl: train/eval log history.artifacts/run_config.json: exact run configuration.artifacts/dataset_stats.json: dataset filtering/split statistics.source/: training/inference code used for this run.
Dataset summary
Input files were four XLSX sources mounted in Modal:
kashmiri_parallel_dataset_dia-non_dia_text_1.xlsxkashmiri_parallel_dataset_dia-non_dia_text_2.xlsxkashmiri_parallel_dataset_dia-non_dia_text_3.xlsxkashmiri_parallel_dataset_dia-non_dia_text_4.xlsx
Filtering and split stats:
- Rows in:
207,047 - Dropped by length:
1,321 - Dropped duplicates:
15,457 - Rows kept:
190,269 - Train rows:
171,344 - Validation rows:
9,536(1,500max per epoch eval) - Test rows:
9,389
Training configuration
- Epoch target:
30 - Completed:
11epochs - Train batch size:
4 - Gradient accumulation:
8 - Effective batch size:
32 - Eval batch size:
1 - Optimizer:
adamw_torch - Learning rate:
5e-4 - Scheduler: cosine
- Warmup ratio:
0.05 - Weight decay:
0.01 - Gradient checkpointing: enabled
- Group by length: enabled
- Max source length:
256 - Max target length:
256 - Save total limit:
2
Loss track
Training loss reduced strongly and stably:
- Step
50:0.2546 - Step
1,000:0.0941 - Step
5,000:0.0624 - Step
10,000:0.0539 - Step
16,000:0.0492 - Step
21,400:0.0425 - Step
26,700:0.0382 - Step
32,100:0.0356 - Step
37,400:0.0308 - Step
42,800:0.0284 - Step
48,100:0.0250 - Step
53,500:0.0214 - Step
58,850:0.0178 - Best observed train loss:
0.0154at step53,850
Approximate epoch train averages:
- Epoch 1:
0.0575 - Epoch 2:
0.0492 - Epoch 3:
0.0434 - Epoch 4:
0.0387 - Epoch 5:
0.0346 - Epoch 6:
0.0309 - Epoch 7:
0.0273 - Epoch 8:
0.0239 - Epoch 9:
0.0205 - Epoch 10: ~`0.017–0.019`
- Epoch 11 boundary latest:
0.0178
Validation metrics by epoch
Lower DER/WER is better; higher exact match is better.
- Epoch 1 / step 5354: eval_loss
0.0676, DER_marked0.2378, DER_all0.0702, WER0.2176, exact0.1300 - Epoch 2 / step 10709: eval_loss
0.0655, DER_marked0.2123, DER_all0.0665, WER0.2106, exact0.1407 - Epoch 3 / step 16062: eval_loss
0.0608, DER_marked0.2284, DER_all0.0638, WER0.2033, exact0.1460 - Epoch 4 / step 21417: eval_loss
0.0585, DER_marked0.1984, DER_all0.0606, WER0.1915, exact0.1660 - Epoch 5 / step 26770: eval_loss
0.0583, DER_marked0.1961, DER_all0.0591, WER0.1868, exact0.1713 - Epoch 6 / step 32125: eval_loss
0.0584, DER_marked0.1923, DER_all0.0590, WER0.1868, exact0.1707 - Epoch 7 / step 37479: eval_loss
0.0587, DER_marked0.1844, DER_all0.0586, WER0.1856, exact0.1747 - Epoch 8 / step 42834: eval_loss
0.0599, DER_marked0.1946, DER_all0.0573, WER0.1828, exact0.1753 - Epoch 9 / step 48188: eval_loss
0.0643, DER_marked0.1927, DER_all0.0572, WER0.1842, exact0.1860 - Epoch 10 / step 53543: eval_loss
0.0663, DER_marked0.2016, DER_all0.0573, WER0.1827, exact0.1940 - Epoch 11 / step 58897: eval_loss
0.0712, DER_marked0.1870, DER_all0.0570, WER0.1818, exact0.1753
How to use
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
import torch
repo_id = "Omarrran/koshur-diacritizer-byt5-small-v2"
tokenizer = AutoTokenizer.from_pretrained(repo_id)
model = AutoModelForSeq2SeqLM.from_pretrained(repo_id)
model.eval()
text = "یہاں اپنا غیر اعرابی کشمیری متن لکھیں"
inputs = tokenizer(text, return_tensors="pt", truncation=True, max_length=256)
with torch.no_grad():
ids = model.generate(
**inputs,
max_new_tokens=256,
num_beams=1,
do_sample=False,
)
print(tokenizer.decode(ids[0], skip_special_tokens=True))
How to resume training
The latest full Trainer checkpoint is:
training-checkpoints/checkpoint-58897
For a Hugging Face Seq2SeqTrainer continuation, restore/copy that folder as the training output checkpoint and call:
trainer.train(resume_from_checkpoint="training-checkpoints/checkpoint-58897")
If continuing on Modal, keep the same data orientation, all 4 XLSX sources, bf16/L4 settings, and output/run format used in artifacts/run_config.json.
Important caveats
- This is not the final 30-epoch model. It is the best safe snapshot completed so far at epoch 11.
- The last run stopped after the checkpoint was saved, during the longer test/eval phase. The checkpoint is structurally complete.
- The source XLSX files themselves are not bundled here unless explicitly uploaded separately; the repo contains model/checkpoint/code/artifacts needed to use or resume when the dataset files are available.
- Downloads last month
- 21
Model tree for Omarrran/koshur-diacritizer-byt5-small-v2
Base model
google/byt5-small