Video summary
James Landay at SNU Data Science Seminar: “AI For Good” Isn’t Good Enough, 2025.09.17
Main summary
Key takeaways
Main ideas and lessons
1) “AI for Good” is not enough
- Simply pursuing “good outcomes” at the technical level (or focusing only on direct users) is insufficient.
- To increase the chance of positive societal impact, AI must be designed with a multi-level view that includes:
- User
- Community
- Society
2) AI systems already affect people—often without people noticing
Generative AI and other AI tools are deeply embedded in daily life (e.g., speech recognition, translation, image generation, coding assistants). Alongside benefits, they have also produced significant harms, including:
- Manipulation of social feeds affecting mental health/body image
- Misinformation undermining democratic processes
- Addiction-like design patterns (e.g., casino-like gaming dynamics)
- Biased high-stakes decisions (e.g., bail/sentencing support shown as biased by race)
- Hallucinations in chatbots causing reputational and legal harm (example: Air Canada case)
3) Guiding principles behind Human-Centered AI research
Referencing Stanford’s Human-Centered AI institute, the talk outlines three guiding principles that steer research, policy, and education:
-
Impact-first principle (societal impact guidance)
- Because AI is too potent, consider ethical/social effects before, during, and after development—not only after deployment.
-
Inspiration principle (learn from human intelligence)
- The brain is extremely efficient; learn from human cognition/behavior to create better algorithms (e.g., data efficiency, power efficiency).
- Stanford invests in neuroscience and cognitive/behavioral sciences to inform AI design.
-
Augmentation principle (augment, don’t replace)
- AI will reshape jobs, but the goal should be to enhance human capabilities rather than replace humans in narrow tasks.
- Governments and society need preparation for workforce upheaval.
4) Why “AI for Good” efforts often fail
The speaker critiques two common approaches that, in his view, lack the interdisciplinary, end-to-end nature of true Human-Centered AI:
-
Type A: Social-science critique-first
- Useful for identifying harms, but often too late to design solutions proactively.
-
Type B: Technologists go it alone
- Applying AI naively to real social domains (health, education, COVID-related tools) may work in controlled settings but fail in real-world contexts.
- Example theme: optimistic predictions (e.g., replacing radiologists) didn’t account for real workflows and job “messiness.”
5) The methodology: Human-Centered AI at three levels
A framework is presented for designing and analyzing AI systems using three coordinated analysis/design levels:
A) User-centered AI (direct interaction + usability)
- Purpose: Build systems around the end user’s needs and abilities.
- Methods/instructions:
- Use HCI (Human-Computer Interaction) techniques.
- Involve end users early and often.
- Iterate rapidly on design.
- Validate through rigorous testing.
- Goal: Systems that function well for humans, not just “mimic humans.”
B) Community-centered AI (impacts beyond the direct user)
- Why necessary: AI side effects can harm people who aren’t the “user” but are still impacted.
- Methods/instructions:
- Identify impacted groups related to the deployment context (e.g., accused people, victims, lawyers, families—not only the judge).
- Include people affected by systemic issues (e.g., structural barriers like racism).
- Recognize that some community members may also need to be included at the user-centered stage if they become direct users.
C) Society-centered AI (systemic consequences if widely adopted)
- Why necessary: Broad deployment can create large-scale societal costs and behavior changes.
- Methods/instructions:
- Forecast outcomes if the system becomes ubiquitous.
- Study societal impacts and costs (e.g., prison costs and effects; traffic and fatalities from autonomous cars).
- Consider indirect effects on groups (e.g., public self-image impacts with Instagram; community-level consequences of surveillance or deployment patterns).
6) Implementation requirement: “True interdisciplinary teams”
Proper Human-Centered AI requires interdisciplinary teams from the start, not “ethics as an afterthought.”
-
Instruction checklist for teams (as described):
- Include technologists/AI experts (core engineering)
- Include design experts (HCI and interaction design)
- Include social sciences/humanities expertise
- Add domain experts depending on application (medicine, law, economics, environmental science, etc.)
-
Key practice rule:
- Social scientists/ethicists must be integrated partners early, not a separate late-stage “harm-check” team.
- Rationale: if profit incentives dominate, late checks won’t prevent deployment.
Example walkthroughs used to illustrate the framework
Example 1: Autonomous vehicles (Tesla/autopilot; broader “autonomous” definitions)
Autonomous driving is used to show how user-only design can miss hard problems.
User-centered focus (what research often did)
- How to replace or augment the driver (visibility, information, interfaces).
Community-centered focus (missed interactions)
- Study effects on other drivers’ behavior when autonomous cars appear.
- Study interactions with pedestrians and bicyclists.
- Early fatalities reportedly involved complexities like combined pedestrian/bicyclist scenarios, suggesting insufficient community-level modeling.
Society-centered focus (city/region consequences)
Consider whether autonomous cars increase:
- traffic,
- deaths (more miles driven),
- or demand for infrastructure expansion.
Example logic:
- people may live farther away and send vehicles to park elsewhere, increasing total driving.
“Power” framing at the society level
At the society level, the speaker frames design as a power question:
- Who decides how resources are spent?
- Power is concentrated in a few dominant companies controlling foundation models and/or autonomous systems.
Example 2: “Hybrid Physical Digital Spaces” (smart buildings / wellness sensing)
The speaker critiques his own project and explains how human-centered analysis expanded it.
Original vision (user wellness + adaptive environments)
Buildings would sense worker stress and adjust:
- lighting,
- music,
- scenery/nature cues,
- dynamic ambient murals guiding exercise or social encounters.
Initial research step (lab study to establish science)
-
Method described:
- One-year lab experiment with 400+ participants
- Manipulated building features:
- natural vs artificial materials,
- natural light/view vs artificial,
- images (including demographic diversity in displayed images)
- Measured stress via wearable sensing (EDA/electrodermal activity) and surveys.
-
Major findings (high level):
- natural materials and natural windows reduced stress
- effects on creativity and subgroup differences were observed.
Scaling outcomes (from surveys to richer measurement)
- Identify multiple outcomes (stress, belonging, creativity, physical activity, attitudes/behavior).
- Avoid focusing on productivity due to misuse concerns (example: warehouse workers fired for performance metrics).
Sensor triangulation approach (instructions)
Use three data sources:
- Self-report, but move beyond one-off surveys
- Experience Sampling Method (ESM): brief periodic check-ins during the day
- Multiple sensors:
- building sensors (occupancy, energy, recycling/water usage)
- personal wearable/phone/laptop sensors for physiological measures (with strong attention to privacy/security)
Privacy/security requirements (as stated)
- Assess participant comfort.
- Make studies opt-in.
- Provide ways to view/delete data.
- Restrict analysis/queries using need-to-know principles.
Added community/society focus via new workshop method (implication design)
The speaker argues user-centered work plus basic community inclusion can still miss higher-level risks, leading to a workshop methodology focused on risks/benefits trade-offs and communication of implications.
Workshop methodology (detailed instruction steps):
- Goal: generate “implication designs” that:
- communicate what’s happening, and
- protect against harms.
- Participants: mixed groups; often community members, sometimes experts.
- Format: card/deck-based, lightweight, iterative.
- Three rounds:
- Anticipation round
- Identify who is affected by the technology.
- Identify positive/negative trade-offs.
- Implication design round
- Propose design interventions to communicate/protect against identified threats.
- Action round (role-play)
- Simulate threats and test countermeasures collaboratively.
- Anticipation round
- Scale reported: workshops run with 150+ participants over 10 iterations.
Illustrative workshop scenario
A health-focused speech AI companion device scenario (doctor programmed; learns the user; makes recommendations). Prototypes included:
- role-based data access levels (doctor vs nurse vs security vs family/friends),
- door/entry device concepts to prevent caregiver misuse and detect suspicious changes.
Q&A themes (what additional guidance was emphasized)
1) Practical difficulty of multi-level projects
Interdisciplinary work is harder because:
- researchers must learn each other’s “language” and incentives,
- evaluation metrics differ across disciplines (papers vs journals vs timeframes).
Advice implied by the response:
- accept upfront learning cost,
- expect better research outcomes despite extra work.
2) Why companies should do it
Compared to adoption paths for user-centered design:
- it was initially uncommon,
- later became standard because it prevents problems and improves revenue.
Human-centered AI may be more expensive initially, but:
- harms avoided = cost avoided,
- better products/markets = revenue gain.
Research goal:
- build techniques/tools to “make it easier,” including simulation methods for society-level effects.
3) Training talent / liberal arts education
The speaker emphasizes maintaining liberal arts breadth alongside technical education, suggesting HCI (or similar human-centered requirements) be part of computer science degree expectations, not only electives.
4) “Team science” and credit/paper-counting
Concern: paper-counting can discourage collaboration and distort incentives.
The speaker argues top institutions emphasize broader impact (papers, startups, open source, public outcomes) rather than pure publication metrics.
Conclusion / takeaways
- AI should be designed not only to augment tasks/users, but to protect and improve outcomes at user, community, and society levels simultaneously.
- Achieving that requires:
- interdisciplinary teams,
- proactive design/analysis (not after-the-fact safety checks),
- and attention to power and resource allocation shaped by deployment.
Speakers / sources featured
Speakers
- Professor James Landay (Stanford; co-founder/co-director, Stanford Institute for Human-Centered AI)
- Professor Cha (host/moderator; thanked speaker and conducted Q&A introductions)
Referenced individuals (examples/examples-in-the-field)
- Jeff Hinton (radiology replacement quote/prediction referenced)
- Kurt Langlotz (radiology AI paper; “radiologists who use AI…” framing referenced)
Referenced systems/cases
- Compass sentencing/bail program (referenced as biased in the U.S.)
- Air Canada chatbot hallucination legal/reputational harm (referenced)
- Instagram (body image impact)
- AlphaFold (scientific discovery example)
- ChatGPT, Midjourney, DALL·E, GitHub Copilot / “co-pilot” (generative AI examples)
- Tesla Autopilot / “full self-driving” (autonomous driving example)
Institutions / organizations referenced
- Stanford University
- Seoul National University (SNU)
- Carnegie Mellon University
- University of Texas at Austin (transportation research referenced)
- ETH Zürich and EPFL (GPU infrastructure exception referenced)
- U.S. Environmental Protection Agency (EPA) (time-in-buildings statistic referenced)
- New York Times (Langlotz/Langlotz paper popularized via NYT referenced)
Data/terms mentioned
- HCI (Human-Computer Interaction)
- EDA (electrodermal activity)
- ESM (Experience Sampling Method)
- “foundation models” and “pedaflops” (compute discussion)