Skip navigation
Please use this identifier to cite or link to this item: https://repository.esi-sba.dz/jspui/handle/123456789/969
Title: AraMS-28k: A Weakly Supervised Pipeline for Line-Level, Layout-Aware Annotation and Restoration of Historical Arabic Manuscripts
Authors: ZELLAGUI, MOhamed DIaa EDdine
Keywords: Arabic Manuscript OCR
Weak Supervision
LLM-Assisted Annotation
Fuzzy String Alignment
Document Layout Analysis
Historical Document Processing
Manuscript Restoration
Marginal Annotation
Issue Date: 2026
Abstract: This thesis presents three integrated contributions to the computational processing of historical Arabic manuscripts: a weakly supervised annotation pipeline, the AraMS-28k dataset, and a manuscript restoration subsystem. First, the thesis develops a weakly supervised annotation pipeline that combines page segmentation, multimodal large language model (LLM) structured OCR, and fuzzy text alignment to automate line-level annotation generation while preserving complete human oversight. A central methodological contribution is the conĄdence-100 rule: we demonstrate both theoretically and empirically that when the normalized fuzzy similarity between an OCR line and its matched ground-truth span reaches its maximum value, the correspondence is correct under the adopted matching assumptionsŮa property we justify formally and then verify empiricallyŮnotwithstanding differences in diacritization. Manual review of all conĄdence-100 samples across 14 books yielded zero observed matching errors, conĄrming the reliability of the validation protocol. Second, using the developed pipeline, we construct and release AraMS-28k, a line-level annotated historical Arabic manuscript dataset that is, to our knowledge, the Ąrst to provide explicit marginal-annotation labels at this scale, and whose multi-script coverage (including Maghrebi and RuqŠah hands) is not matched by existing public resources. It comprises 14 books, 548 fully annotated pages, 2,495 partially annotated pages, 27,969 main-text line samples, and 131 marginal line samples. The dataset includes bounding boxes, main/margin labels, insertion anchors for margin lines, conĄdence scores, and full provenance information. The framework achieves a 6.5× effective speed-up over manual transcription workĆows. To establish the dataset as a usable benchmark, we train a Transformer line recogniser on it and report a stratiĄed cross-script evaluation: 6.48% character error rate on in-distribution pages, rising to 25.37% on a fully unseen Naskh book and 37.88% on unseen Maghrebi script Ů a gradient that quantiĄes, rather than merely asserts, the generalisation gap the dataset is built to expose. Third, the thesis introduces a manuscript restoration subsystem comprising a synthetic degradation model applied on the Ćy during training (four freshly-randomised degraded variants per clean crop, resampled every epoch) and a U-Net restoration network enhanced with a text-aware training objective: a frozen HATFormer recogniser reads the restored image and its recognition error against the ground-truth transcription is back-propagated into the restorer. Following the recognition-based evaluation of AutoHDR, the restorer is assessed on both image-quality metrics and the downstream OCR error of the restored output; on the held-out test set the text-aware objective improves downstream recognition over a recognition-agnostic U-Net baseline at near-equal structural Ądelity (SSIM 0.978versus 0.989), reducing character error rate from 35.1% on the degraded input to 27.1% Ů against 28.1% for the recognition-agnostic baseline Ů and recovering the corrupted text to within about one point of the 26% clean-line recognition level.*** Ce memoire presente trois contributions integrees pour le traitement informatique des manuscrits arabes historiques: un pipeline dŠannotation faiblement supervise, lŠensemble de donnees AraMS-28k, et un sous-systeme de restauration de manuscrits. Premierement, le memoire developpe un pipeline dŠannotation faiblement supervise combinant la segmentation de pages, lŠOCR structuree par grand modele de langage multimodal (LLM), et un alignment Ćou de texte pour automatiser la generation dŠannotations au niveau ligne tout en preservant une supervision humaine complete. Notre contribution methodologique centrale est la regle de conĄance-100 : nous demontrons formellement et empiriquement que lorsque la similarite Ćoue normalisee atteint sa valeur maximale, la correspondance de ligne est correcte, independamment des differences de vocalisation. LŠaccord inter-annotateurs (κ = 0,94) conĄrme la Ąabilite du protocole de validation. Deuxiemement, a lŠaide du pipeline developpe, nous construisons et diffusons AraMS- 28k: 14 ouvrages, 548 pages entierement annotees, 2,495 pages partiellement annotees, 27,969 echantillons de lignes de texte principal et 131 echantillons de lignes marginales. Le pipeline atteint une acceleration effective de 6,5× par rapport a la transcription manuelle. Troisiemement, le memoire presente un sous-systeme de restauration comprenant un modele de degradation synthetique applique a la volee pendant lŠentrainement (quatre variantes degradees re-tirees aleatoirement par imagette, re-echantillonnees a chaque epoque) et un reseau U-Net enrichi dŠun objectif sensible au texte: un reconnaisseur HATFormer Ąge lit lŠimage restauree et son erreur de reconnaissance par rapport a la transcription de reference est retropropagee dans le restaurateur. Suivant le protocole dŠevaluation par reconnaissance dŠAutoHDR, le restaurateur est evalue a la fois sur des metriques de qualite dŠimage et sur lŠerreur OCR en aval; lŠobjectif sensible au texte ameliore la reconnaissance par rapport a une base U-Net non sensible au texte, a Ądelite structurelle quasi identique (SSIM 0,978 contre 0,989), reduisant le taux dŠerreur de caracteres sur lŠensemble de test de 35,1% (entree degradee) a 27,1% (contre 28,1% pour la base non sensible au texte).
Description: Supervisor : Dr. Chaib Souleyman
URI: https://repository.esi-sba.dz/jspui/handle/123456789/969
Appears in Collections:Ingenieur

Files in This Item:
File Description SizeFormat 
Engineering-1-1.pdf55,66 kBAdobe PDFView/Open
Show full item record


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.