Skip navigation
Please use this identifier to cite or link to this item: https://repository.esi-sba.dz/jspui/handle/123456789/988
Title: Generative AI, Privacy and Personal Data: Mapping and Reducing Data Leakage in Retrieval-Augmented Generation Systems
Authors: TADJ, MOhamed WAlid
Keywords: Generative AI
Retrieval-Augmented Generation
Privacy Leakage
Reidentification
Quasi-Identifiers
LLM-As-Judge
Privacy-Utility Trade-Off
Data Governance.
Issue Date: 2026
Abstract: Enterprise retrieval-augmented generation (RAG) makes a sensitive document corpus conversationally accessible, and in doing so can disclose personal data to users and attackers not authorized to see it. This thesis measures that leakage rather than asserting it. It formalizes a LINDDUN-aligned metric pair (a binary leakage function L and a severity score S) evaluated through a profile-dependent engine across four CVSS-calibrated attacker profiles and a systematic adversarial campaign of eight attack families, over a deterministically generated corpus of 155 synthetic documents (2,190 entity-level annotations of personally identifiable information, PII; inter-annotator κ = 0.875), and quantifies a layered defense by ablation. The central finding is a genuine but partial reduction: through a validated language-model judge, the full stack lowers measured leakage from 28.8% to 17.3%, about forty percent, not an elimination. Redaction defeats direct PII extraction, but re-identification by linkage, a person reassembled from individually innocuous attributes scattered across documents, is the dominant unreduced residual, and is shown intractable at the model layer across three independent mechanisms: redaction, query-time kanonymity, and ingestion-time k-anonymity all fail or are foreclosed by their information cost. Linkage is therefore a data-governance problem, not a model-layer one: access control solves the external, unauthorized attacker, whereas the authorized user who brings outside knowledge leaves a residual that can be bounded, minimized and audited, but not eliminated. A methodological contribution underwrites these results: the deterministic detection engine is blind to linkage (κ = −0.10 against human judgment), while the language-model judge, reliable on undefended output (κ = 0.827), degrades on the defended distribution (κ = 0.306), so leakage evaluation must be re-validated on the distribution it scores, not only on undefended text. The absolute percentages demonstrate a measurement methodology on a controlled corpus; they are not a safety certification for any particular deployment.***
Description: Supervisor : Dr. CHAIB Souleyman /Co-Supervisor :Pr. AÏMEUR Esma
URI: https://repository.esi-sba.dz/jspui/handle/123456789/988
Appears in Collections:Ingenieur

Files in This Item:
File Description SizeFormat 
PFE_Report-1-1.pdf71,83 kBAdobe PDFView/Open
Show full item record


Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.