Video summary

2019-01-16 CERIAS - A Data Privacy Primer

Main summary

Key takeaways

Educational

Main Ideas & Concepts Covered

1) What “data privacy” means (and what it’s not)

Privacy is presented as more than secrecy/confidentiality. Core themes include:

  • Protection: your information shouldn’t adversely affect you later.
  • Control / approved use: data should only be used in ways you authorize.
  • Disclosure & sharing: privacy concerns arise when information is disclosed or shared.
  • Recourse: if privacy is violated, people need ways to respond/remedy.

2) Historical and legal framing

Warren & Brandeis (U.S. legal tradition)

  • The “right to be let alone” is described as a foundational articulation of privacy in the U.S.

EU Charter / GDPR concept (modern legal view)

  • Personal data must be protected and processed:
    • Fairly
    • For specified purposes
    • Based on consent (or another legitimate legal basis)
  • Data subjects have rights such as:
    • Access to data collected about them
    • Rectification (correction) of inaccurate data

3) Why “obvious” privacy definitions and consent are hard in practice

The talk emphasizes that simple rules often fail under real-world conditions, such as:

  • “You consent when you click accept”
  • “Privacy means your data isn’t identified”

Consent example: Facebook photo used for a billboard

  • A family’s photo appeared on a billboard in the Czech Republic, apparently without their direct awareness.
  • The explanation hinges on Facebook’s Terms of Service:
    • Uploading content grants Facebook a broad license (non-exclusive, transferable, sublicensable, royalty-free, worldwide) to host/distribute/modify/create derivatives, consistent with privacy/app settings.
  • The mismatch is framed as:
    • Legally: consent exists via the Terms
    • Viscerally: users may not feel they truly agreed to that use
  • Deleting the account may not terminate the granted rights (described as “before April 19th,” rights in perpetuity).

“Personal information” and breach disclosure laws are limited

  • The speaker discusses breach disclosure laws and how they define “personal information,” which can be:
    • Narrow (triggering disclosure only if specific elements are present)
    • Weak for privacy protection, because privacy harm can occur even when strict definitions aren’t met
  • Example implications discussed:
    • Publishing data like name/address + other sensitive details might not meet trigger conditions if it lacks specific elements (e.g., driver’s license or credit card numbers).
    • Risk from “random” numbers:
      • If random values could be interpreted as Social Security numbers, they might still fall under sensitive personal information definitions—even without names.

4) Re-identification and “anonymity” is brittle

A major message: removing direct identifiers is not sufficient, because data can become identifiable through linkage with external information.

AOL search data (anonymized but re-identified)

  • AOL released ~20 million search queries from 650,000 users with user IDs replaced by random numbers.
  • A reporter used query content to identify an individual (knocking on their door and confirming).
  • Lesson: “anonymized” datasets can still enable identification.

Netflix challenge / other anonymized data

  • The talk references “Netflix challenge”-style work where identities/links could be inferred from de-anonymized patterns.

Census / “theoretically possible”

  • Privacy techniques have been pressured because re-identification can be possible, leading to changing practices (including an example of the Census Bureau).

Location/travel release problems (NYC taxis, Uber-like records)

  • NYC released start/end points of taxi rides.
  • Even without names, location traces can reveal sensitive information (e.g., home location).
  • Efforts to protect privacy by sharing with third parties (e.g., Uber) still faced complications when data was released “as is.”

5) Formalizing anonymization concepts and modern definitions

k-anonymity idea

  • Goal: for a released record (given quasi-identifier set), it cannot be distinguished among at least k records.
  • If each combination appears at least k times, inferring individual membership is harder.

EU/other definitions of personal data

  • GDPR (summarized):
    • Personal data is information relating to an identified or identifiable natural person (directly or indirectly).
  • Another more plain-language approach:
    • Data belonging to an individual with specified or reasonably specifiable identity, including through combination with other data sources.

Privacy vs utility tradeoff (no perfect guarantee)

  • Perfect privacy in this formal sense would imply zero utility.
  • Identification can occur once the adversary has enough external information.
  • Therefore, released information always carries some risk of becoming identifying.

Methodologies & Techniques Discussed

A) Encryption (confidentiality) and beyond

  • Encryption
    • Protects data confidentiality (keeps content secret).
  • Digital signatures
    • Support integrity (detect tampering).
  • Encrypted computation
    • Compute on data while it remains encrypted, so plaintext is never exposed.
  • Fully Homomorphic Encryption (FHE)
    • Enables processing encrypted data to perform operations (described as add/multiply/logic on encrypted values).
    • Also noted: moving toward privacy of the computation itself (hiding what is being computed).
    • Presented as an active research/prototype area.

B) Anonymization (generalization/suppression)

  • Generalization / k-anonymity-style anonymization
    • Reduce specificity of identifiers (e.g., exact birthdate → birth year).
  • Separate link vs sensitive values (described project approach)
    • Encrypt identifiers and the link between identifying info and sensitive info.
    • Allow matching only at a group/generalized level (e.g., a color category), not at the individual level.
  • Key limitation
    • Anonymity is brittle: it can collapse if enough auxiliary info arrives.

C) Noise-based privacy (not as brittle)

  • Noise edition / noise mechanisms
    • Add randomness so external knowledge doesn’t deterministically reveal individuals.

Differential Privacy (DP) — emphasized technique

  • Core goal: make the output about aggregated data nearly unchanged whether or not any single individual is included.
  • Privacy parameter (ε) controls the bound on how much outputs can differ between neighboring datasets.
  • Two intuitive requirements:
    • Removing/adding one person should not change the answer too much.
    • Outputs for neighboring datasets should have close/bounded distributions (privacy loss relates to ε).

DP mechanism examples

  • Laplace mechanism (for numeric answers)

    • Compute the true statistic/query result.
    • Add noise sampled from a Laplace distribution.
    • Noise scale depends on:
      • Sensitivity: maximum possible change in the statistic when one individual’s data changes (e.g., count changes by at most 1)
      • ε: privacy parameter controlling noise magnitude
    • Important nuance: sensitivity uses the worst-case maximum, not average/expected change.
  • Randomized response (for survey yes/no style data; described as DP)

    • For each respondent, flip a coin to decide whether to report truth or a randomized answer.
    • Over many participants, true prevalence/answer can be estimated.
    • Provides plausible deniability per person; helps protect against inferring individual inclusion.
  • Exponential mechanism (for categorical / selection among outputs)

    • Define:
      • Utility scores for candidate outputs (how good each result is)
      • Sensitivity of utility
    • Select outputs with probability proportional to a function of utility and ε:
      • Higher-utility outputs more likely, but incorrect ones possible.
    • Useful when privacy must be maintained while “choosing the best label.”

D) DP Properties (how guarantees behave)

  • Post-processing theorem
    • Any transformation after a DP mechanism preserves DP (as long as you don’t access raw original data again).
  • Composition
    • Multiple DP mechanisms on the same dataset accumulate privacy loss.
    • Total ε becomes roughly the sum of ε values across mechanisms.
    • Emphasizes graceful degradation rather than abrupt failure.

E) Interpreting ε and strengthening protection notions

  • ε is not simply the probability of identification; it is a subtle bound.
  • Differential identifiability (referenced concept)
    • Connects DP to probability of identification under a specified adversary model/game.
    • Benefit: calibrates ε to an assumed adversary strength and combines with DP’s graceful degradation.

Final Takeaways / Lessons Emphasized

  • Privacy should be considered as a system of principles, not just confidentiality.
  • Practical re-identification is common: removing identifiers doesn’t guarantee safety.
  • Differential privacy offers a more robust framework than brittle anonymity because it uses noise calibrated to formal guarantees.
  • Computer security techniques are central to meeting privacy principles in practice.

Speakers / Sources Featured

Speakers

  • Chris Clifton (host/faculty member for the seminar)

Sources / Referenced Organizations / Documents

  • Warren & Brandeis (Harvard Law Review article) — “right to be let alone”
  • Harvard Law Review (venue/source of the Warren & Brandeis privacy articulation)
  • Charter of Fundamental Rights of the European Union
  • GDPR (General Data Protection Regulation) (via its interpretation of the Charter)
  • Indiana Code (breach disclosure law and related definition discussion)
  • AOL (AOL customer search data release example)
  • The New York Times (reporter example used to re-identify from AOL queries)
  • Latanya Sweeney (de-anonymization / linkage example involving hospital admissions and voter lists)
  • Netflix (Netflix challenge referenced)
  • Census Bureau (differential privacy/changes mentioned)
  • New York City (taxi data release example)
  • Uber (records request / privacy controversy referenced)
  • GE Privacy / “Suresh/Serum pure Angeles emirati” (unclear exact attribution due to subtitle errors; referenced in conjunction with a k-anonymity-related solution)
  • serum pure Angeles emirati / cannon enmity (subtitle appears noisy; likely referencing anonymity methods/research collaborators)
  • Fair Information Practices (FIPs) (“PHIPPs” as transcribed) — privacy principles referenced
  • Freedom of Information Act (FOIA) (hospital admissions records request example)

(End of summary.)

Original video