Privacy by Design: From Distributed Learning to Post-Deployment Risks

DSpace Repositorium (Manakin basiert)


Dateien:

Zitierfähiger Link (URI): http://hdl.handle.net/10900/182018
http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1820182
Dokumentart: Dissertation
Erscheinungsdatum: 2026-07-31
Sprache: Englisch
Fakultät: 7 Mathematisch-Naturwissenschaftliche Fakultät
Fachbereich: Informatik
Gutachter: Akgün, Mete (Dr.)
Tag der mündl. Prüfung: 2026-07-17
Freie Schlagwörter: Privacy by Design
Datenschutzfreundliches maschinelles Lernen
Verteiltes maschinelles Lernen
Genomweite Assoziationsstudien
Quantenmaschinelles Lernen
Adversarielles maschinelles Lernen
k-Anonymität
Privacy by Design
Privacy-Preserving Machine Learning
Distributed Machine Learning
Genome-Wide Association Studies
Quantum Machine Learning
Adversarial Machine Learning
k-Anonymity
Lizenz: http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en
Zur Langanzeige

Abstract:

Modern data processing activities increasingly rely on sensitive data, from medical images and genomic data to electronic health records, to enable research, discovery, diagnosis, and decision-making. Yet the very capabilities that make these systems powerful also create privacy risks: data must often be shared across institutions for generalization and accuracy, and once models and datasets are released for downstream use, they become long-lived artifacts that others can query, link, or exploit. Privacy by design, the principle that privacy should be a foundational property of any data processing activity rather than a post-incident patch, offers a response to these risks. However, translating this into concrete technical practice remains a challenge, because the right answer depends on where in the processing lifecycle one stands, what is being protected, and what assumptions about adversaries are relevant. This thesis approaches privacy by design as a technical agenda organized around two regimes: pre-deployment and post-deployment. On the pre-deployment side, we develop privacy-preserving methods for computation on distributed data under semi-honest threat models. We introduce randomized-encoding based approaches for scalable kernel learning on medical images and for multi-site genome-wide association studies on quantitative phenotypes. We further show that widely used classical kernels can be realized through quantum feature maps, and introduce a distributed secure quantum architecture for kernel computation, validated on simulated quantum hardware. On the post-deployment side, we study what deployed artifacts reveal and how that exposure can be exploited or mitigated. We introduce a targeted adversarial attack for hard-label black-box image classifiers that leverages edge information from images to accelerate attack progress under strict query budgets, consistently outperforming existing methods in the low-query regime across diverse architectures. We also develop a topological framework for adaptive k-anonymisation for dynamic datasets, enabling incremental updates to anonymised data releases without full recomputation when the underlying data changes. Taken together, the results presented in this thesis demonstrate that privacy by design for data processing activities is not a single technique but a discipline whose realization spans a composition of architectures, threat models, and lifecycle stages.

Das Dokument erscheint in: