Skip to main navigation Skip to search Skip to main content

Engineering Privacy in Practice: Privacy Engineering Solutions for GDPR-Compliant Contemporary Software Systems

Research output: ThesisPhd Thesis 1 (Research TU/e / Graduation TU/e)

13 Downloads (Pure)

Abstract

Privacy is a context-dependent property. Under this interpretation, individuals shall possess mechanisms to create different instances of their privacy, varying in kind and in magnitude, according to their preferences. In software engineering, this requires concrete controls across the personal data lifecycle for collection, processing, storage, sharing, retention and deletion, together with governance, accountability and auditable evidence. While Privacy by Design provides principles, the central challenge is to translate those principles into requirements, architectures, operational procedures and measurable properties that demonstrate that mechanisms behave as intended in each context. Data-intensive architectures and Artificial Intelligence (AI) have expanded the contexts in which software handles personal data. In parallel, regulations such as the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA) impose lifecycle obligations and formalize data-subject rights, moving privacy from policy text into engineered practice that must be planned, executed and evidenced with rigor comparable to security and reliability. Publication trends show very high volumes on federated learning and sustained but lower volumes on privacy engineering. This imbalance suggests that scholarship has focused on privacy awareness through mechanisms such as Privacy-Enhancing Technologies. In contrast, the end-to-end engineering of privacy across the development and personal data lifecycles and under operational constraints has received comparatively less attention. This dissertation addresses this gap with the following goal: to extend current knowledge and practice in privacy engineering by providing a consolidated view of the engineering space and by delivering validated methodologies and software that demonstrate privacy awareness and verifiable compliance in operational settings. In light of this objective, we formulate the following research questions:• RQ1: How has privacy engineering evolved since the introduction of privacy legislation and what aspects organize its engineering practice?• RQ2: How should privacy engineering be incorporated in decentralized data governance paradigms so that privacy aspects are reflected in executable controls and auditable evidence?• RQ3: How can Data Minimization and Purpose Limitation be formulated as measurable constraints that link privacy-related artifacts to source code?• RQ4: What design principles enable deployable text anonymization services that support analytics while meeting privacy obligations?• RQ5: How can organizations validate and evaluate third-party synthetic data in a compliant and secure manner and select datasets that preserve task utility?• RQ6: How can established software engineering practice improve federated learning under heterogeneous data in privacy-sensitive applications? To address RQ1, a Systematic Literature Review with thematic synthesis organizes prior evidence into thirteen higher-order themes that characterize privacy engineering. The synthesis makes explicit how these themes interact, where they are situated within both the software development and personal data lifecycles and how emphases vary across domains. It also highlights underrepresented primary foci, including Data Minimization and Purpose Limitation, Lifelong Management and Incident Response and Management. This mapping links policy-facing artifacts to enforcement points and evidence and provides a field-level scaffold for subsequent work. Using design science research, we develop a conceptual framework for decentralized data privacy governance within data mesh environments to address RQ2. The framework maps privacy engineering themes with data mesh components, clarifies roles and responsibilities and specifies platform capabilities that link declared policies to access control, provenance and audit, thereby enabling testable compliance during operation. To address RQ3, Data Minimization and Purpose Limitation are formalized as code level constraints using functional modeling. The approach defines document-to-code linkage, introduces quantitative measures for over-collection, under-collection and purpose alignment and provides guidance for threshold selection and adaptation so the method generalizes across heterogeneous software systems. For RQ4 and RQ5, two industrial studies applying action research in operational settings demonstrate the translation of theory into practice. First, an interoperable text anonymization service is designed, integrated and evaluated within a streaming analytics platform, documenting design principles and measurements that support GDPR compliant processing while preserving downstream analytic utility. Second, a cloud platform and a governance evaluation protocol are co-designed and deployed for externally generated synthetic data with one telecommunications provider and two vendors. The work defines reproducible, auditable procedures, implements validation checks that prevent disclosure of values of the original dataset while enabling data integrity and provides quantitative evaluation using distributional and correlation measures against a permuted baseline, followed by an offline stakeholder study aligned with an internal Interest Score. Finally, to address RQ6, an experimental study adapts an established software engineering pattern for use in federated learning in heterogeneous data settings. By organizing client-side processing and aggregation and specifying baselines and selection criteria, the study reports cross-dataset results, including healthcare, showing improved performance relative to general baselines under Non-Independent and Identically Distributed(IID) data conditions. Together, these studies move privacy engineering from a principle to a verifiable execution. They deliver a map of the engineering space and a set of methods and systems that produce the artifacts, controls and evidence required by contemporary regulation. For scholars, the results identify underrepresented areas and provide a replication-ready basis for updates and specialization. For practitioners, they show how to plan artifact sand hand-offs across lifecycles and how to make policies executable and auditable in deployed systems.
Original languageEnglish
QualificationDoctor of Philosophy
Awarding Institution
  • JADS Den Bosch
Supervisors/Advisors
  • van den Heuvel, Willem-Jan A.M., Promotor
  • Tamburri, Damian A., Copromotor
Award date23 Jun 2026
Place of PublicationEindhoven
Publisher
Print ISBNs978-90-386-6724-9
DOIs
Publication statusPublished - 23 Jun 2026

Bibliographical note

Proefschrift.

Fingerprint

Dive into the research topics of 'Engineering Privacy in Practice: Privacy Engineering Solutions for GDPR-Compliant Contemporary Software Systems'. Together they form a unique fingerprint.

Cite this