| DC Field | Value | Language |
| dc.contributor.author | TADJ, MOhamed WAlid | - |
| dc.date.accessioned | 2026-10-07T08:19:36Z | - |
| dc.date.available | 2026-10-07T08:19:36Z | - |
| dc.date.issued | 2026 | - |
| dc.identifier.uri | https://repository.esi-sba.dz/jspui/handle/123456789/988 | - |
| dc.description | Supervisor : Dr. CHAIB Souleyman /Co-Supervisor :Pr. AÏMEUR Esma | en_US |
| dc.description.abstract | Enterprise retrieval-augmented generation (RAG) makes a sensitive document corpus
conversationally accessible, and in doing so can disclose personal data to users
and attackers not authorized to see it. This thesis measures that leakage rather than
asserting it. It formalizes a LINDDUN-aligned metric pair (a binary leakage function
L and a severity score S) evaluated through a profile-dependent engine across
four CVSS-calibrated attacker profiles and a systematic adversarial campaign of
eight attack families, over a deterministically generated corpus of 155 synthetic documents
(2,190 entity-level annotations of personally identifiable information, PII;
inter-annotator κ = 0.875), and quantifies a layered defense by ablation. The central
finding is a genuine but partial reduction: through a validated language-model
judge, the full stack lowers measured leakage from 28.8% to 17.3%, about forty percent,
not an elimination. Redaction defeats direct PII extraction, but re-identification
by linkage, a person reassembled from individually innocuous attributes scattered
across documents, is the dominant unreduced residual, and is shown intractable
at the model layer across three independent mechanisms: redaction, query-time kanonymity,
and ingestion-time k-anonymity all fail or are foreclosed by their information
cost. Linkage is therefore a data-governance problem, not a model-layer one:
access control solves the external, unauthorized attacker, whereas the authorized
user who brings outside knowledge leaves a residual that can be bounded, minimized
and audited, but not eliminated. A methodological contribution underwrites these
results: the deterministic detection engine is blind to linkage (κ = −0.10 against human
judgment), while the language-model judge, reliable on undefended output
(κ = 0.827), degrades on the defended distribution (κ = 0.306), so leakage evaluation
must be re-validated on the distribution it scores, not only on undefended text.
The absolute percentages demonstrate a measurement methodology on a controlled
corpus; they are not a safety certification for any particular deployment.*** | en_US |
| dc.language.iso | en | en_US |
| dc.subject | Generative AI | en_US |
| dc.subject | Retrieval-Augmented Generation | en_US |
| dc.subject | Privacy Leakage | en_US |
| dc.subject | Reidentification | en_US |
| dc.subject | Quasi-Identifiers | en_US |
| dc.subject | LLM-As-Judge | en_US |
| dc.subject | Privacy-Utility Trade-Off | en_US |
| dc.subject | Data Governance. | en_US |
| dc.title | Generative AI, Privacy and Personal Data: Mapping and Reducing Data Leakage in Retrieval-Augmented Generation Systems | en_US |
| dc.type | Thesis | en_US |
| Appears in Collections: | Ingenieur
|