Skip navigation
Please use this identifier to cite or link to this item: https://repository.esi-sba.dz/jspui/handle/123456789/988
Full metadata record
DC FieldValueLanguage
dc.contributor.authorTADJ, MOhamed WAlid-
dc.date.accessioned2026-10-07T08:19:36Z-
dc.date.available2026-10-07T08:19:36Z-
dc.date.issued2026-
dc.identifier.urihttps://repository.esi-sba.dz/jspui/handle/123456789/988-
dc.descriptionSupervisor : Dr. CHAIB Souleyman /Co-Supervisor :Pr. AÏMEUR Esmaen_US
dc.description.abstractEnterprise retrieval-augmented generation (RAG) makes a sensitive document corpus conversationally accessible, and in doing so can disclose personal data to users and attackers not authorized to see it. This thesis measures that leakage rather than asserting it. It formalizes a LINDDUN-aligned metric pair (a binary leakage function L and a severity score S) evaluated through a profile-dependent engine across four CVSS-calibrated attacker profiles and a systematic adversarial campaign of eight attack families, over a deterministically generated corpus of 155 synthetic documents (2,190 entity-level annotations of personally identifiable information, PII; inter-annotator κ = 0.875), and quantifies a layered defense by ablation. The central finding is a genuine but partial reduction: through a validated language-model judge, the full stack lowers measured leakage from 28.8% to 17.3%, about forty percent, not an elimination. Redaction defeats direct PII extraction, but re-identification by linkage, a person reassembled from individually innocuous attributes scattered across documents, is the dominant unreduced residual, and is shown intractable at the model layer across three independent mechanisms: redaction, query-time kanonymity, and ingestion-time k-anonymity all fail or are foreclosed by their information cost. Linkage is therefore a data-governance problem, not a model-layer one: access control solves the external, unauthorized attacker, whereas the authorized user who brings outside knowledge leaves a residual that can be bounded, minimized and audited, but not eliminated. A methodological contribution underwrites these results: the deterministic detection engine is blind to linkage (κ = −0.10 against human judgment), while the language-model judge, reliable on undefended output (κ = 0.827), degrades on the defended distribution (κ = 0.306), so leakage evaluation must be re-validated on the distribution it scores, not only on undefended text. The absolute percentages demonstrate a measurement methodology on a controlled corpus; they are not a safety certification for any particular deployment.***en_US
dc.language.isoenen_US
dc.subjectGenerative AIen_US
dc.subjectRetrieval-Augmented Generationen_US
dc.subjectPrivacy Leakageen_US
dc.subjectReidentificationen_US
dc.subjectQuasi-Identifiersen_US
dc.subjectLLM-As-Judgeen_US
dc.subjectPrivacy-Utility Trade-Offen_US
dc.subjectData Governance.en_US
dc.titleGenerative AI, Privacy and Personal Data: Mapping and Reducing Data Leakage in Retrieval-Augmented Generation Systemsen_US
dc.typeThesisen_US
Appears in Collections:Ingenieur

Files in This Item:
File Description SizeFormat 
PFE_Report-1-1.pdf71,83 kBAdobe PDFView/Open
Show simple item record


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.