Koshur OCRv6 β€” Kashmiri OCR Recognition Model

Koshur OCRv6 is a Kashmiri optical character recognition (OCR) text-recognition model trained with the PaddleOCR PP-OCRv6 small recognition stack. It is intended for recognizing Kashmiri text from cropped text-line images, including Perso-Arabic/Nastaliq-style Kashmiri text.

This repository contains the trained model checkpoints, exact source code/configuration used for training, training logs, metrics, and reproducibility artifacts. The training dataset itself is intentionally not included in this model repository.

Repository status

  • Training status: completed successfully
  • PBS/HPC exit status: 0
  • Training runtime: 16h 34m 27s
  • Final epoch: 200 / 200
  • Final global step: 532,800
  • Recommended checkpoint: checkpoints/best_accuracy.pdparams
  • Best epoch: 191

What is included

  • checkpoints/
    • best_accuracy.pdparams β€” best validation checkpoint by norm_edit_dis; recommended for inference/evaluation.
    • latest.pdparams, latest.pdopt, latest.states β€” latest/final training state.
    • iter_epoch_200.* β€” final epoch checkpoint.
    • iter_epoch_191.* β€” best epoch checkpoint, when available.
    • all-epoch-checkpoints/ β€” full epoch checkpoint archive is being uploaded separately when available.
  • artifacts/
    • train.log β€” full PaddleOCR training log.
    • metrics_summary.json β€” parsed training/evaluation summary.
    • eval_history.jsonl β€” validation history extracted from logs.
    • train_loss_tail.jsonl β€” tail of parsed step-level training loss records.
    • assets/ β€” pretrained/base artifacts used for the run where applicable.
  • source/
    • koshur-ocrv6/ β€” project code, scripts, configs, tests, and documentation used for this run.
    • PaddleOCR/ β€” PaddleOCR source snapshot used by the training command.

What is not included

The image dataset is not uploaded here. No train/validation/test images or label-list files are included in this model repo.

Dataset provenance and split counts are documented below for reproducibility.

Dataset

Source dataset: Omarrran/Koshur_Pixel

The dataset was prepared for OCR recognition training as cropped text-image samples with labels. It was exported into PaddleOCR recognition-list format on Pragya HPC, then split into train/validation/test partitions.

Dataset split counts used in this training run:

  • Training samples: 330,078
  • Validation samples: 10,388
  • Test samples: 6,995
  • Total samples: 347,461
  • Character dictionary size: 124 entries

Dataset exclusion policy:

  • The dataset is excluded from this repository to keep the model release separate from data distribution.
  • This repo publishes the trained OCR model, code, configs, logs, metrics, and model artifacts only.

Architecture

This run uses PaddleOCR's PP-OCRv6 small recognition architecture configured for Kashmiri OCR recognition.

High-level structure:

  • OCR task: cropped text-line recognition / image-to-text
  • Framework: PaddleOCR / PaddlePaddle
  • Recognition stack: PP-OCRv6 small recognition model
  • Losses observed in training logs:
    • CTC loss (CTCLoss)
    • NRTR loss (NRTRLoss)
    • Combined training loss (loss)
  • Main validation indicator: norm_edit_dis
  • Secondary metric: exact-match acc
  • Max text length: 160
  • Character dictionary: source/koshur-ocrv6/artifacts/charset/koshur_dict.txt

The training was initialized from:

artifacts/assets/pretrained/PP-OCRv6_small_rec_pretrained.pdparams

Methodology

Training pipeline:

  1. Build/verify Kashmiri character dictionary.
  2. Export OCR recognition examples into PaddleOCR-compatible recognition data.
  3. Train PP-OCRv6 small recognition model for 200 epochs.
  4. Evaluate periodically on the validation split using normalized edit distance and exact-match accuracy.
  5. Save epoch checkpoints and a best validation checkpoint.
  6. Preserve source code, configs, training logs, and checkpoints for reproducibility.

Training configuration highlights:

  • Epochs: 200
  • Global steps: 532,800
  • GPUs: 4 Γ— NVIDIA A100-SXM4-80GB
  • Batch size per GPU/card: 64
  • Eval batch step: [0,1000]
  • Save epoch step: 1
  • Max text length: 160
  • Mixed precision flag in best metric: is_float16=False

The exact config used for training is included under:

source/koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml

Results

Primary validation metric: norm_edit_dis (normalized edit-distance similarity; higher is better).

Best validation result:

  • Best epoch: 191
  • Best norm_edit_dis: 0.42845766519166395
  • Best exact-match acc: 0.007893723519547821
  • Validation FPS at best checkpoint: 1565.0712993756804

Final recorded validation result:

  • Final eval index: 532
  • Final norm_edit_dis: 0.42539299199315106
  • Final exact-match acc: 0.006738544467906677
  • Final validation FPS: 1567.270772256869

Final training step:

  • Epoch: 200 / 200
  • Global step: 532,800
  • Learning rate: 0.000001
  • Batch exact-match acc: 0.015625
  • Batch norm_edit_dis: 0.385448
  • CTC loss: 1.182781
  • NRTR loss: 1.165309
  • Combined loss: 2.351904
  • Throughput: 278.53996 samples/sec

Full log and parsed histories:

  • Full training log: artifacts/train.log
  • Parsed summary: artifacts/metrics_summary.json
  • Eval history: artifacts/eval_history.jsonl
  • Training-loss tail: artifacts/train_loss_tail.jsonl

Loading and using the model

This is a PaddleOCR/PaddlePaddle checkpoint release, not a standard Hugging Face Transformers model. Use the included PaddleOCR source snapshot and config.

Example setup:

git clone https://huggingface.co/Omarrran/koshur-ocrv6
cd koshur-ocrv6
python -m venv .venv
source .venv/bin/activate
pip install -r source/koshur-ocrv6/requirements.txt

Example inference/evaluation pattern:

cd source/PaddleOCR
python tools/infer_rec.py \
  -c ../koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml \
  -o Global.pretrained_model=../../checkpoints/best_accuracy \
     Global.character_dict_path=../koshur-ocrv6/artifacts/charset/koshur_dict.txt \
     Global.infer_img=/path/to/cropped_text_line_image.png

For batch evaluation, adapt:

source/koshur-ocrv6/scripts/06_eval_rec.py

Recommended checkpoint for inference/evaluation:

checkpoints/best_accuracy.pdparams

If PaddleOCR expects the checkpoint prefix rather than the full .pdparams filename, pass:

checkpoints/best_accuracy

Reproducing training

The original training command used Paddle distributed launch on Pragya HPC with 4 GPUs. The essential training arguments were:

python -m paddle.distributed.launch --devices <4 GPUs> tools/train.py \
  -c source/koshur-ocrv6/configs/rec/koshur_ppocrv6_small_rec.yml \
  -o Global.epoch_num=200 \
     Global.save_model_dir=output/koshur_ppocrv6_small_rec_full \
     Global.use_gpu=true \
     Global.distributed=true \
     Global.character_dict_path=source/koshur-ocrv6/artifacts/charset/koshur_dict.txt \
     Global.pretrained_model=artifacts/assets/pretrained/PP-OCRv6_small_rec_pretrained.pdparams \
     Global.eval_batch_step=[0,1000] \
     Global.save_epoch_step=1 \
     Global.max_text_length=160 \
     Metric.main_indicator=norm_edit_dis \
     Metric.ignore_space=false \
     Train.loader.batch_size_per_card=64 \
     Train.loader.num_workers=8 \
     Eval.loader.batch_size_per_card=64 \
     Eval.loader.num_workers=4

Dataset paths are intentionally omitted here because the dataset is not included in this model repo. Reproduction requires preparing the Omarrran/Koshur_Pixel dataset into PaddleOCR recognition-list format.

Limitations

  • Exact-match accuracy is low because full-string OCR exact match is strict, especially for long Kashmiri text strings and orthographic variation.
  • norm_edit_dis is the primary metric used for checkpoint selection.
  • This checkpoint should be evaluated on downstream page/line crops before deployment.
  • The model card reports validation metrics extracted from the training log; separate held-out test-set evaluation should be added before making benchmark or SOTA claims.
  • The release is PaddleOCR-format, not Transformers-format; use the included PaddleOCR code/configs.

Citation / attribution

If you use this model, please cite or acknowledge:

  • The Koshur OCRv6 training run and model release: Omarrran/koshur-ocrv6
  • Dataset provenance: Omarrran/Koshur_Pixel
  • PaddleOCR / PP-OCRv6 framework

License and data notes

No dataset files are redistributed in this repository. Users are responsible for checking the dataset license/provenance and downstream deployment constraints before using the model commercially or publicly.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support