Video summary

2020 STAT115 Lect1.2 Big Data Challenge

Main summary

Key takeaways

Educational

Main Ideas, Concepts, and Lessons

1) Growth of sequencing scale and automation (2001 → present)

  • Around 2001: Sequencing was large-scale, labor-intensive, and highly structured.
    • A genome center could have multiple rooms:
      • Sample-prep room: ~35 people, preparing samples for 3–4 weeks to get DNA into a sequencing-ready form.
      • Sequencing room: about 74 large sequencing machines.
    • Operations / logistics:
      • ~10 people operated the machines.
      • Each instrument could run 15–40 runs per day.
    • Data output (per day):
      • Each instrument produced about 1–2 megabytes/day (as stated in subtitles).
      • The entire room could produce about 120 megabytes/day.
  • Key improvement: Modern sequencing is highly automated and dramatically better, but the major revolution occurred in 2007 with massively parallel sequencing.

2) Modern sequencing workflow: faster sample prep + long-running machines

  • 2007+ workflow changes:
    • Sample preparation: reduced to something like a station-based setup, allowing one person to prepare samples in under a day.
    • Sequencing machines:
      • The same operator can load samples and then leave.
      • Machines run for 3–5 days.
      • Each instrument can generate roughly ~0.5 gigabytes/day per machine (as described).
    • The process is described as a flow path:
      • load DNA on one part of the system,
      • DNA flows through,
      • sequence data comes out directly from the instrument.

3) “Mini machines” vs sequencing facilities: different scales, output, and reads

  • Individual lab / mini machines (example given):
    • Can run at least a full day.
    • Produce on the order of millions of reads per day (e.g., “4 million” or “25 million reads/day,” as stated).
    • Reads are typically paired-end with ~150 base pairs per read (paired structure emphasized).
  • Sequencing facility / core facility (bigger scale):
    • Machines run 1–2 days.
    • Example output:
      • ~300 gigabytes in 1–2 days
      • about 1 billion reads from one run
      • paired-end, ~150 base pair reads
    • Mention of large institutions having such core facilities (e.g., Harvard-related and cancer institute-related facilities).

4) Very large “mega” sequencing runs (flow cells, capacities, and genome throughput)

  • The “mega” sequencer is described as being machine-sized like a copier, but with extremely high throughput.
  • Continuous operation idea:
    • It can run two runs at a time by alternating flow cells, keeping the machine running much of the time.
  • Flow cell / read-length flexibility:
    • Output changes depending on sequencing length settings (shorter vs longer read lengths).
  • Data scale (as stated):
    • In dual flow-cell mode, up to ~6000 gigabytes of data per certain unit time/run configuration (subtitles mention “as high as six thousand gigabytes of data”).
  • Examples connecting read length to time and output:
    • Sequencing length (e.g., ~160 bp) relates to runtime:
      • mentioned: ~24 hours for 160 bp
      • longer lengths (e.g., 250 bp) imply longer runtime (e.g., ~60 hours mentioned)
    • Each run produces on the order of tens of gigabytes (e.g., ~40 GB stated in one example).
  • Throughput claim:
    • A single run can cover:
      • about 48 whole genomes
      • and ~500 exomes (as stated in subtitles).

5) Dramatic drop in cost and the role of computing/algorithms

  • Human Genome Project (early era):
    • Initially cost ~$30 million over 13 years (as stated).
  • Modern era improvements:
    • With enough automation and the high-throughput shift:
      • can redo the project for about ~$100 million (as stated for a hypothetical redo).
    • 2007: high-throughput sequencing availability → huge price drop.
  • Current pricing (as stated):
    • Pricing is said to hover around ~$1000 per human genome.
  • But computational burden remains:
    • Sequencing generates terabytes of raw data, requiring:
      • major storage
      • heavy processing
      • efficient algorithms to interpret results quickly.
  • Algorithmic progress using deep learning (2018 example):
    • A deep neural network approach was developed to call mutations / germline variants from raw data.
    • Motivation:
      • Individual genomes are ~99.99% similar
      • so it’s better to store/capture only differences rather than everything.
    • Outcome:
      • quickly identifies variant locations without needing to store all raw sequence detail.

6) Real-world applications: personal health and precision medicine

  • Angelina Jolie (germline risk → preventive action):
    • Chose double mastectomy due to inherited risk indicated by family history and genetic testing.
    • Demonstrates how genetic risk information can lead to preventative procedures.
  • Consumer genetic testing (23andMe example):
    • 23andMe: send DNA via mouth swab
    • Cost around $99 (sometimes on sale)
    • Provides health/disease risk predictions (as referenced in subtitles).
    • Potential downstream actions:
      • increased screening
      • preventive procedures (for high-risk individuals)
      • lifestyle changes (e.g., diet/alcohol/exercise adjustments).
  • Cancer and somatic mutation targeting (precision medicine):
    • Example institution: Dana-Farber Cancer Institute.
    • For diagnosed cancer patients:
      • DNA sequencing is performed on tumor samples
      • typically not whole genome
      • instead, sequence a few hundred genes (targeted panels)
      • use detected mutations to choose targeted therapies
    • Matching mutation → drug targeting (examples mentioned in subtitles):
      • If relevant mutations like an “…EGFR mutation” are present, use an EGFR inhibitor
      • Specific mutations correspond to specific targeted drugs (some drug name text was garbled in the source, but the intent was mutation-specific therapy).
  • Lesson / application takeaway:
    • Genomic sequencing supports both:
      • research and
      • clinical decision-making and
      • consumer health risk assessment
    • This creates a “huge data challenge,” motivating the course.

Methodology / Instructions Presented (Workflow Summary)

While no formal step-by-step “how-to” list is provided, the subtitles describe a sequencing workflow at a high level:

  • Sample preparation
    • Historically: weeks of prep with many staff
    • Modern systems: prep in <1 day
  • Load samples into sequencing system
    • Place prepared DNA into the sequencer and start the run
  • Run sequencing for extended duration
    • Modern large instruments run for 3–5 days (in one described mode)
  • Generate sequence data
    • Data output is produced directly by the sequencing hardware across long runs
  • Downstream compute / interpret variants
    • Use algorithms (including deep neural networks) to call mutations/variants
  • Clinical use for treatment decisions
    • In cancer care:
      • sequence tumor DNA using targeted gene panels (not full genome)
      • identify mutations
      • choose drugs targeted to those mutations

Speakers or Sources Featured (As Named in Subtitles)

  • Harvard (references to Harvard main campus / Harvard Medical School)
  • Dana-Farber Cancer Institute
  • 23andMe
  • Angelina Jolie
  • Human Genome Project (project as a source/event; not an individual)
  • Moore’s law (concept cited)

Original video