Video summary
2020 STAT115 Lect1.2 Big Data Challenge
Main summary
Key takeaways
Main Ideas, Concepts, and Lessons
1) Growth of sequencing scale and automation (2001 → present)
- Around 2001: Sequencing was large-scale, labor-intensive, and highly structured.
- A genome center could have multiple rooms:
- Sample-prep room: ~35 people, preparing samples for 3–4 weeks to get DNA into a sequencing-ready form.
- Sequencing room: about 74 large sequencing machines.
- Operations / logistics:
- ~10 people operated the machines.
- Each instrument could run 15–40 runs per day.
- Data output (per day):
- Each instrument produced about 1–2 megabytes/day (as stated in subtitles).
- The entire room could produce about 120 megabytes/day.
- A genome center could have multiple rooms:
- Key improvement: Modern sequencing is highly automated and dramatically better, but the major revolution occurred in 2007 with massively parallel sequencing.
2) Modern sequencing workflow: faster sample prep + long-running machines
- 2007+ workflow changes:
- Sample preparation: reduced to something like a station-based setup, allowing one person to prepare samples in under a day.
- Sequencing machines:
- The same operator can load samples and then leave.
- Machines run for 3–5 days.
- Each instrument can generate roughly ~0.5 gigabytes/day per machine (as described).
- The process is described as a flow path:
- load DNA on one part of the system,
- DNA flows through,
- sequence data comes out directly from the instrument.
3) “Mini machines” vs sequencing facilities: different scales, output, and reads
- Individual lab / mini machines (example given):
- Can run at least a full day.
- Produce on the order of millions of reads per day (e.g., “4 million” or “25 million reads/day,” as stated).
- Reads are typically paired-end with ~150 base pairs per read (paired structure emphasized).
- Sequencing facility / core facility (bigger scale):
- Machines run 1–2 days.
- Example output:
- ~300 gigabytes in 1–2 days
- about 1 billion reads from one run
- paired-end, ~150 base pair reads
- Mention of large institutions having such core facilities (e.g., Harvard-related and cancer institute-related facilities).
4) Very large “mega” sequencing runs (flow cells, capacities, and genome throughput)
- The “mega” sequencer is described as being machine-sized like a copier, but with extremely high throughput.
- Continuous operation idea:
- It can run two runs at a time by alternating flow cells, keeping the machine running much of the time.
- Flow cell / read-length flexibility:
- Output changes depending on sequencing length settings (shorter vs longer read lengths).
- Data scale (as stated):
- In dual flow-cell mode, up to ~6000 gigabytes of data per certain unit time/run configuration (subtitles mention “as high as six thousand gigabytes of data”).
- Examples connecting read length to time and output:
- Sequencing length (e.g., ~160 bp) relates to runtime:
- mentioned: ~24 hours for 160 bp
- longer lengths (e.g., 250 bp) imply longer runtime (e.g., ~60 hours mentioned)
- Each run produces on the order of tens of gigabytes (e.g., ~40 GB stated in one example).
- Sequencing length (e.g., ~160 bp) relates to runtime:
- Throughput claim:
- A single run can cover:
- about 48 whole genomes
- and ~500 exomes (as stated in subtitles).
- A single run can cover:
5) Dramatic drop in cost and the role of computing/algorithms
- Human Genome Project (early era):
- Initially cost ~$30 million over 13 years (as stated).
- Modern era improvements:
- With enough automation and the high-throughput shift:
- can redo the project for about ~$100 million (as stated for a hypothetical redo).
- 2007: high-throughput sequencing availability → huge price drop.
- With enough automation and the high-throughput shift:
- Current pricing (as stated):
- Pricing is said to hover around ~$1000 per human genome.
- But computational burden remains:
- Sequencing generates terabytes of raw data, requiring:
- major storage
- heavy processing
- efficient algorithms to interpret results quickly.
- Sequencing generates terabytes of raw data, requiring:
- Algorithmic progress using deep learning (2018 example):
- A deep neural network approach was developed to call mutations / germline variants from raw data.
- Motivation:
- Individual genomes are ~99.99% similar
- so it’s better to store/capture only differences rather than everything.
- Outcome:
- quickly identifies variant locations without needing to store all raw sequence detail.
6) Real-world applications: personal health and precision medicine
- Angelina Jolie (germline risk → preventive action):
- Chose double mastectomy due to inherited risk indicated by family history and genetic testing.
- Demonstrates how genetic risk information can lead to preventative procedures.
- Consumer genetic testing (23andMe example):
- 23andMe: send DNA via mouth swab
- Cost around $99 (sometimes on sale)
- Provides health/disease risk predictions (as referenced in subtitles).
- Potential downstream actions:
- increased screening
- preventive procedures (for high-risk individuals)
- lifestyle changes (e.g., diet/alcohol/exercise adjustments).
- Cancer and somatic mutation targeting (precision medicine):
- Example institution: Dana-Farber Cancer Institute.
- For diagnosed cancer patients:
- DNA sequencing is performed on tumor samples
- typically not whole genome
- instead, sequence a few hundred genes (targeted panels)
- use detected mutations to choose targeted therapies
- Matching mutation → drug targeting (examples mentioned in subtitles):
- If relevant mutations like an “…EGFR mutation” are present, use an EGFR inhibitor
- Specific mutations correspond to specific targeted drugs (some drug name text was garbled in the source, but the intent was mutation-specific therapy).
- Lesson / application takeaway:
- Genomic sequencing supports both:
- research and
- clinical decision-making and
- consumer health risk assessment
- This creates a “huge data challenge,” motivating the course.
- Genomic sequencing supports both:
Methodology / Instructions Presented (Workflow Summary)
While no formal step-by-step “how-to” list is provided, the subtitles describe a sequencing workflow at a high level:
- Sample preparation
- Historically: weeks of prep with many staff
- Modern systems: prep in <1 day
- Load samples into sequencing system
- Place prepared DNA into the sequencer and start the run
- Run sequencing for extended duration
- Modern large instruments run for 3–5 days (in one described mode)
- Generate sequence data
- Data output is produced directly by the sequencing hardware across long runs
- Downstream compute / interpret variants
- Use algorithms (including deep neural networks) to call mutations/variants
- Clinical use for treatment decisions
- In cancer care:
- sequence tumor DNA using targeted gene panels (not full genome)
- identify mutations
- choose drugs targeted to those mutations
- In cancer care:
Speakers or Sources Featured (As Named in Subtitles)
- Harvard (references to Harvard main campus / Harvard Medical School)
- Dana-Farber Cancer Institute
- 23andMe
- Angelina Jolie
- Human Genome Project (project as a source/event; not an individual)
- Moore’s law (concept cited)