Confusion model weighted by how Arabic script actually fails
Dot and skeleton confusion, word splits and merges, dropped and spurious characters
Severity sampled 4–18% per example, so one model covers clean and degraded pages
LoRA r=32 on all attention and MLP projections, bf16, no quantisation
Greedy decoding: one right answer, no sampling
Guardrail rejects wholesale rewrites and keeps the original
CER / WER evaluation by severity band against the raw-OCR baseline
Trains on a free Colab T4 with an automatic fp16 fallback