Koshur OCRv6 β Kashmiri OCR Recognition Model
Koshur OCRv6 is a Kashmiri optical character recognition (OCR) text-recognition model trained with the PaddleOCR PP-OCRv6 small recognition stack. It is intended for recognizing Kashmiri text from cropped text-line images, including Perso-Arabic/Nastaliq-style Kashmiri text.
This repository contains the trained model checkpoints, exact source code/configuration used for training, training logs, metrics, and reproducibility artifacts. The training dataset itself is intentionally not included in this model repository.
Repository status
- Training status: completed successfully
- PBS/HPC exit status: 0
- Training runtime: 16h 34m 27s
- Final epoch: 200 / 200
- Final global step: 532,800
- Recommended checkpoint:
checkpoints/best_accuracy.pdparams - Best epoch: 191
What is included
checkpoints/best_accuracy.pdparamsβ best validation checkpoint bynorm_edit_dis; recommended for inference/evaluation.latest.pdparams,latest.pdopt,latest.statesβ latest/final training state.iter_epoch_200.*β final epoch checkpoint.iter_epoch_191.*β best epoch checkpoint, when available.all-epoch-checkpoints/β full epoch checkpoint archive is being uploaded separately when available.
artifacts/train.logβ full PaddleOCR training log.metrics_summary.jsonβ parsed training/evaluation summary.eval_history.jsonlβ validation history extracted from logs.train_loss_tail.jsonlβ tail of parsed step-level training loss records.assets/β pretrained/base artifacts used for the run where applicable.
source/koshur-ocrv6/β project code, scripts, configs, tests, and documentation used for this run.PaddleOCR/β PaddleOCR source snapshot used by the training command.
What is not included
The image dataset is not uploaded here. No train/validation/test images or label-list files are included in this model repo.
Dataset provenance and split counts are documented below for reproducibility.
Dataset
Source dataset: Omarrran/Koshur_Pixel
The dataset was prepared for OCR recognition training as cropped text-image samples with labels. It was exported into PaddleOCR recognition-list format on Pragya HPC, then split into train/validation/test partitions.
Dataset split counts used in this training run:
- Training samples: 330,078
- Validation samples: 10,388
- Test samples: 6,995
- Total samples: 347,461
- Character dictionary size: 124 entries
Dataset exclusion policy:
- The dataset is excluded from this repository to keep the model release separate from data distribution.
- This repo publishes the trained OCR model, code, configs, logs, metrics, and model artifacts only.
Architecture
This run uses PaddleOCR's PP-OCRv6 small recognition architecture configured for Kashmiri OCR recognition.
High-level structure:
- OCR task: cropped text-line recognition / image-to-text
- Framework: PaddleOCR / PaddlePaddle
- Recognition stack: PP-OCRv6 small recognition model
- Losses observed in training logs:
- CTC loss (
CTCLoss) - NRTR loss (
NRTRLoss) - Combined training loss (
loss)
- CTC loss (
- Main validation indicator:
norm_edit_dis - Secondary metric: exact-match
acc - Max text length: 160
- Character dictionary:
source/koshur-ocrv6/artifacts/charset/koshur_dict.txt
The training was initialized from:
artifacts/assets/pretrained/PP-OCRv6_small_rec_pretrained.pdparams
Methodology
Training pipeline:
- Build/verify Kashmiri character dictionary.
- Export OCR recognition examples into PaddleOCR-compatible recognition data.
- Train PP-OCRv6 small recognition model for 200 epochs.
- Evaluate periodically on the validation split using normalized edit distance and exact-match accuracy.
- Save epoch checkpoints and a best validation checkpoint.
- Preserve source code, configs, training logs, and checkpoints for reproducibility.
Training configuration highlights:
- Epochs: 200
- Global steps: 532,800
- GPUs: 4 Γ NVIDIA A100-SXM4-80GB
- Batch size per GPU/card: 64
- Eval batch step:
[0,1000] - Save epoch step: 1
- Max text length: 160
- Mixed precision flag in best metric:
is_float16=False
The exact config used for training is included under:
source/koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml
Results
Primary validation metric: norm_edit_dis (normalized edit-distance similarity; higher is better).
Best validation result:
- Best epoch: 191
- Best
norm_edit_dis: 0.42845766519166395 - Best exact-match
acc: 0.007893723519547821 - Validation FPS at best checkpoint: 1565.0712993756804
Final recorded validation result:
- Final eval index: 532
- Final
norm_edit_dis: 0.42539299199315106 - Final exact-match
acc: 0.006738544467906677 - Final validation FPS: 1567.270772256869
Final training step:
- Epoch: 200 / 200
- Global step: 532,800
- Learning rate: 0.000001
- Batch exact-match acc: 0.015625
- Batch
norm_edit_dis: 0.385448 - CTC loss: 1.182781
- NRTR loss: 1.165309
- Combined loss: 2.351904
- Throughput: 278.53996 samples/sec
Full log and parsed histories:
- Full training log:
artifacts/train.log - Parsed summary:
artifacts/metrics_summary.json - Eval history:
artifacts/eval_history.jsonl - Training-loss tail:
artifacts/train_loss_tail.jsonl
Loading and using the model
This is a PaddleOCR/PaddlePaddle checkpoint release, not a standard Hugging Face Transformers model. Use the included PaddleOCR source snapshot and config.
Example setup:
git clone https://huggingface.co/Omarrran/koshur-ocrv6
cd koshur-ocrv6
python -m venv .venv
source .venv/bin/activate
pip install -r source/koshur-ocrv6/requirements.txt
Example inference/evaluation pattern:
cd source/PaddleOCR
python tools/infer_rec.py \
-c ../koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml \
-o Global.pretrained_model=../../checkpoints/best_accuracy \
Global.character_dict_path=../koshur-ocrv6/artifacts/charset/koshur_dict.txt \
Global.infer_img=/path/to/cropped_text_line_image.png
For batch evaluation, adapt:
source/koshur-ocrv6/scripts/06_eval_rec.py
Recommended checkpoint for inference/evaluation:
checkpoints/best_accuracy.pdparams
If PaddleOCR expects the checkpoint prefix rather than the full .pdparams filename, pass:
checkpoints/best_accuracy
Reproducing training
The original training command used Paddle distributed launch on Pragya HPC with 4 GPUs. The essential training arguments were:
python -m paddle.distributed.launch --devices <4 GPUs> tools/train.py \
-c source/koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml \
-o Global.epoch_num=200 \
Global.save_model_dir=output/koshur_ppocrv6_small_rec_full \
Global.use_gpu=true \
Global.distributed=true \
Global.character_dict_path=source/koshur-ocrv6/artifacts/charset/koshur_dict.txt \
Global.pretrained_model=artifacts/assets/pretrained/PP-OCRv6_small_rec_pretrained.pdparams \
Global.eval_batch_step=[0,1000] \
Global.save_epoch_step=1 \
Global.max_text_length=160 \
Metric.main_indicator=norm_edit_dis \
Metric.ignore_space=false \
Train.loader.batch_size_per_card=64 \
Train.loader.num_workers=8 \
Eval.loader.batch_size_per_card=64 \
Eval.loader.num_workers=4
Dataset paths are intentionally omitted here because the dataset is not included in this model repo. Reproduction requires preparing the Omarrran/Koshur_Pixel dataset into PaddleOCR recognition-list format.
Limitations
- Exact-match accuracy is low because full-string OCR exact match is strict, especially for long Kashmiri text strings and orthographic variation.
norm_edit_disis the primary metric used for checkpoint selection.- This checkpoint should be evaluated on downstream page/line crops before deployment.
- The model card reports validation metrics extracted from the training log; separate held-out test-set evaluation should be added before making benchmark or SOTA claims.
- The release is PaddleOCR-format, not Transformers-format; use the included PaddleOCR code/configs.
Citation / attribution
If you use this model, please cite or acknowledge:
- The Koshur OCRv6 training run and model release:
Omarrran/koshur-ocrv6 - Dataset provenance:
Omarrran/Koshur_Pixel - PaddleOCR / PP-OCRv6 framework
License and data notes
No dataset files are redistributed in this repository. Users are responsible for checking the dataset license/provenance and downstream deployment constraints before using the model commercially or publicly.