Video summary
2019-01-16 CERIAS - A Data Privacy Primer
Main summary
Key takeaways
Main Ideas & Concepts Covered
1) What “data privacy” means (and what it’s not)
Privacy is presented as more than secrecy/confidentiality. Core themes include:
- Protection: your information shouldn’t adversely affect you later.
- Control / approved use: data should only be used in ways you authorize.
- Disclosure & sharing: privacy concerns arise when information is disclosed or shared.
- Recourse: if privacy is violated, people need ways to respond/remedy.
2) Historical and legal framing
Warren & Brandeis (U.S. legal tradition)
- The “right to be let alone” is described as a foundational articulation of privacy in the U.S.
EU Charter / GDPR concept (modern legal view)
- Personal data must be protected and processed:
- Fairly
- For specified purposes
- Based on consent (or another legitimate legal basis)
- Data subjects have rights such as:
- Access to data collected about them
- Rectification (correction) of inaccurate data
3) Why “obvious” privacy definitions and consent are hard in practice
The talk emphasizes that simple rules often fail under real-world conditions, such as:
- “You consent when you click accept”
- “Privacy means your data isn’t identified”
Consent example: Facebook photo used for a billboard
- A family’s photo appeared on a billboard in the Czech Republic, apparently without their direct awareness.
- The explanation hinges on Facebook’s Terms of Service:
- Uploading content grants Facebook a broad license (non-exclusive, transferable, sublicensable, royalty-free, worldwide) to host/distribute/modify/create derivatives, consistent with privacy/app settings.
- The mismatch is framed as:
- Legally: consent exists via the Terms
- Viscerally: users may not feel they truly agreed to that use
- Deleting the account may not terminate the granted rights (described as “before April 19th,” rights in perpetuity).
“Personal information” and breach disclosure laws are limited
- The speaker discusses breach disclosure laws and how they define “personal information,” which can be:
- Narrow (triggering disclosure only if specific elements are present)
- Weak for privacy protection, because privacy harm can occur even when strict definitions aren’t met
- Example implications discussed:
- Publishing data like name/address + other sensitive details might not meet trigger conditions if it lacks specific elements (e.g., driver’s license or credit card numbers).
- Risk from “random” numbers:
- If random values could be interpreted as Social Security numbers, they might still fall under sensitive personal information definitions—even without names.
4) Re-identification and “anonymity” is brittle
A major message: removing direct identifiers is not sufficient, because data can become identifiable through linkage with external information.
AOL search data (anonymized but re-identified)
- AOL released ~20 million search queries from 650,000 users with user IDs replaced by random numbers.
- A reporter used query content to identify an individual (knocking on their door and confirming).
- Lesson: “anonymized” datasets can still enable identification.
Netflix challenge / other anonymized data
- The talk references “Netflix challenge”-style work where identities/links could be inferred from de-anonymized patterns.
Census / “theoretically possible”
- Privacy techniques have been pressured because re-identification can be possible, leading to changing practices (including an example of the Census Bureau).
Location/travel release problems (NYC taxis, Uber-like records)
- NYC released start/end points of taxi rides.
- Even without names, location traces can reveal sensitive information (e.g., home location).
- Efforts to protect privacy by sharing with third parties (e.g., Uber) still faced complications when data was released “as is.”
5) Formalizing anonymization concepts and modern definitions
k-anonymity idea
- Goal: for a released record (given quasi-identifier set), it cannot be distinguished among at least k records.
- If each combination appears at least k times, inferring individual membership is harder.
EU/other definitions of personal data
- GDPR (summarized):
- Personal data is information relating to an identified or identifiable natural person (directly or indirectly).
- Another more plain-language approach:
- Data belonging to an individual with specified or reasonably specifiable identity, including through combination with other data sources.
Privacy vs utility tradeoff (no perfect guarantee)
- Perfect privacy in this formal sense would imply zero utility.
- Identification can occur once the adversary has enough external information.
- Therefore, released information always carries some risk of becoming identifying.
Methodologies & Techniques Discussed
A) Encryption (confidentiality) and beyond
- Encryption
- Protects data confidentiality (keeps content secret).
- Digital signatures
- Support integrity (detect tampering).
- Encrypted computation
- Compute on data while it remains encrypted, so plaintext is never exposed.
- Fully Homomorphic Encryption (FHE)
- Enables processing encrypted data to perform operations (described as add/multiply/logic on encrypted values).
- Also noted: moving toward privacy of the computation itself (hiding what is being computed).
- Presented as an active research/prototype area.
B) Anonymization (generalization/suppression)
- Generalization / k-anonymity-style anonymization
- Reduce specificity of identifiers (e.g., exact birthdate → birth year).
- Separate link vs sensitive values (described project approach)
- Encrypt identifiers and the link between identifying info and sensitive info.
- Allow matching only at a group/generalized level (e.g., a color category), not at the individual level.
- Key limitation
- Anonymity is brittle: it can collapse if enough auxiliary info arrives.
C) Noise-based privacy (not as brittle)
- Noise edition / noise mechanisms
- Add randomness so external knowledge doesn’t deterministically reveal individuals.
Differential Privacy (DP) — emphasized technique
- Core goal: make the output about aggregated data nearly unchanged whether or not any single individual is included.
- Privacy parameter (ε) controls the bound on how much outputs can differ between neighboring datasets.
- Two intuitive requirements:
- Removing/adding one person should not change the answer too much.
- Outputs for neighboring datasets should have close/bounded distributions (privacy loss relates to ε).
DP mechanism examples
-
Laplace mechanism (for numeric answers)
- Compute the true statistic/query result.
- Add noise sampled from a Laplace distribution.
- Noise scale depends on:
- Sensitivity: maximum possible change in the statistic when one individual’s data changes (e.g., count changes by at most 1)
- ε: privacy parameter controlling noise magnitude
- Important nuance: sensitivity uses the worst-case maximum, not average/expected change.
-
Randomized response (for survey yes/no style data; described as DP)
- For each respondent, flip a coin to decide whether to report truth or a randomized answer.
- Over many participants, true prevalence/answer can be estimated.
- Provides plausible deniability per person; helps protect against inferring individual inclusion.
-
Exponential mechanism (for categorical / selection among outputs)
- Define:
- Utility scores for candidate outputs (how good each result is)
- Sensitivity of utility
- Select outputs with probability proportional to a function of utility and ε:
- Higher-utility outputs more likely, but incorrect ones possible.
- Useful when privacy must be maintained while “choosing the best label.”
- Define:
D) DP Properties (how guarantees behave)
- Post-processing theorem
- Any transformation after a DP mechanism preserves DP (as long as you don’t access raw original data again).
- Composition
- Multiple DP mechanisms on the same dataset accumulate privacy loss.
- Total ε becomes roughly the sum of ε values across mechanisms.
- Emphasizes graceful degradation rather than abrupt failure.
E) Interpreting ε and strengthening protection notions
- ε is not simply the probability of identification; it is a subtle bound.
- Differential identifiability (referenced concept)
- Connects DP to probability of identification under a specified adversary model/game.
- Benefit: calibrates ε to an assumed adversary strength and combines with DP’s graceful degradation.
Final Takeaways / Lessons Emphasized
- Privacy should be considered as a system of principles, not just confidentiality.
- Practical re-identification is common: removing identifiers doesn’t guarantee safety.
- Differential privacy offers a more robust framework than brittle anonymity because it uses noise calibrated to formal guarantees.
- Computer security techniques are central to meeting privacy principles in practice.
Speakers / Sources Featured
Speakers
- Chris Clifton (host/faculty member for the seminar)
Sources / Referenced Organizations / Documents
- Warren & Brandeis (Harvard Law Review article) — “right to be let alone”
- Harvard Law Review (venue/source of the Warren & Brandeis privacy articulation)
- Charter of Fundamental Rights of the European Union
- GDPR (General Data Protection Regulation) (via its interpretation of the Charter)
- Indiana Code (breach disclosure law and related definition discussion)
- AOL (AOL customer search data release example)
- The New York Times (reporter example used to re-identify from AOL queries)
- Latanya Sweeney (de-anonymization / linkage example involving hospital admissions and voter lists)
- Netflix (Netflix challenge referenced)
- Census Bureau (differential privacy/changes mentioned)
- New York City (taxi data release example)
- Uber (records request / privacy controversy referenced)
- GE Privacy / “Suresh/Serum pure Angeles emirati” (unclear exact attribution due to subtitle errors; referenced in conjunction with a k-anonymity-related solution)
- serum pure Angeles emirati / cannon enmity (subtitle appears noisy; likely referencing anonymity methods/research collaborators)
- Fair Information Practices (FIPs) (“PHIPPs” as transcribed) — privacy principles referenced
- Freedom of Information Act (FOIA) (hospital admissions records request example)
(End of summary.)