Video summary

2020 STAT115 Lect1.1 Bioinfo History

Main summary

Key takeaways

Educational

Main ideas / lessons conveyed

  • Bioinformatics and computational biology evolved in waves, driven by new biological data types and the need for new computational methods to interpret them.
  • Core theme: biology advances → data becomes available → computational algorithms become necessary to:
    • compare
    • assemble
    • classify
    • predict biological function and/or structure

Timeline of major “waves” and key methods

1) Protein sequence & structure (early “wave”)

Key milestones and methods

  • 1965 – Protein sequencing

    • Frederick Sanger invented protein sequencing methods.
    • Enabled the key goal: determining whether a new protein sequence is similar to previously known sequences.
  • Sequence alignment algorithms

    • Need: rapidly compare a new sequence against all sequences in existence.
    • Smith–Waterman: a famous sequence alignment algorithm.
  • 1973 – Protein structure database

    • Creation of the PDB (Protein Data Bank).
    • Rationale: protein function depends on 3D structure, not only sequence.
    • As crystallography structures accumulate, PDB stores them for later querying.
  • 1990 – Faster similarity search: BLAST

    • Instead of exhaustive alignment, BLAST searches a pre-existing protein database.
    • Result: much quicker retrieval of similar sequences.
  • 1994 – Motif discovery / sequence patterns

    • With enough sequences, researchers identify sequence motifs (recurring patterns).
    • Motifs can suggest function (e.g., motifs linked to ATP binding or catalytic activity).
    • A “Blocks” database is mentioned as summarizing motif/function relationships from known sequences.

Protein structure prediction competitions

  • CASP (Critical Assessment of protein Structure Prediction)

    • Motivation: if many groups publish “best” predictors, the field needs objective evaluation.
    • Runs every two years.
    • Predictions are made without knowing the true structures during the competition.
    • After a delay, experimentally solved structures become gold standards used to score methods.
  • Impact of CASP

    • Raised the bar: methods can’t rely only on claims of being “the best.”
    • Some computational biologists left the field after seeing weaknesses against gold-standard evaluation.
  • 2018 (CASP13 era) / Deep learning breakthrough

    • AlphaFold (DeepMind; “alphago” referenced) achieved strong performance:
      • Best predictions for 25 out of 43 proteins (as described)
      • Second place for 3 out of 43 proteins (as described)
      • Also described as predicting the first structure within hours
    • The story emphasizes that competitions can validate breakthroughs and accelerate progress.

2) Microarrays & gene expression (second “wave”)

Concept and motivation

  • A professor introduced microarrays to the class as an exciting idea.
  • Core concept: measure gene expression in bulk, rather than one RNA/gene at a time.

Microarray development and workflow

  • Early phase:
    • Spotted arrays on glass slides
    • Initially manual, later improved with robotic spotting
  • Scaling up:
    • From hundreds/thousands of targets to commercial systems (e.g., Affymetrix)
    • With millions of probes on small arrays

Operational workflow (as described):

  1. Apply a sample mixture of RNA from cells to the microarray.
  2. RNA hybridizes to matching probes.
  3. Wash away non-specific binding.
  4. Detect light signals to infer which genes are expressed.

Computational analysis (1995–2010)

  • Many microarray analysis methods were developed roughly between 1995 and 2010.

Example application: leukemia diagnosis/classification

  • Distinguishing similar leukemia types:
    • ALL vs AML
    • Similar by pathology but different in treatment/outcome
  • Method described:
    • Use microarray expression profiles across ~20,000 genes
    • Identify gene sets/signatures consistently higher in one type vs the other
  • For a new patient:
    • run the same microarray
    • use the signature gene-set expression to predict diagnosis
    • guide treatment selection

3) DNA sequencing & genomics (third “wave” and moving focus)

Enabling concepts

DNA sequencing became feasible through:

  • understanding DNA (double helix context)
  • recombinant DNA (cutting/stitching)
  • Sanger sequencing
  • PCR to amplify DNA when starting material is limited

Databases and reuse: NCBI + BLAST

  • NCBI (National Center for Biotechnology Information) built/maintained sequence databases.
  • BLAST used repeatedly to determine whether newly sequenced fragments are:
    • already known, or
    • novel (then deposited into NCBI)

Speedup and sequencing “race”

  • Anecdote: sequencing one gene once could take years (e.g., yeast gene/7S RNA described as a long PhD effort).
  • Modern sequencing/deposition/review can be described as extremely fast (e.g., “three seconds”).

Human Genome Project (public effort)

  • Began around 1990, with target completion by 2005 (as described).
  • Analogy:
    • 23 chromosome pairs = 23 “books”
    • divide work into smaller pieces (chapters/pages), then assemble later
  • Early decade emphasis:
    • first ~10 years focused on correct mapping/division relationships before deep sequencing

Celera Genomics (private effort; “whole-genome shotgun”)

  • Mentioned as led by Craig Venter.
  • Strategic shift:
    • avoid extensive division/mapping (“don’t divide into books”)
    • make copies → shred into small pieces
    • assemble computationally using overlaps

Informatics as a decisive factor

  • Even if public work used structured pieces, informatics still mattered.
  • For shotgun sequencing, informatics becomes even more critical:
    • it generates many short reads (e.g., ~1000 base pairs as described)
    • these must be assembled computationally into longer contiguous genomes

Joint milestone announcement

  • White House / Clinton helped coordinate public and private efforts to avoid destructive dynamics.
  • Joint announcement:
    • Spring of 2000: “draft human genome” issued (with gaps)
    • additional work continued for ~3 years to fill remaining gaps

Broader lesson / outcome

  • Major beneficiaries: scientists and the broader community via data access.
  • Long-term impact:
    • stimulated gene prediction and downstream research/drug discovery
  • Current practice (as described):
    • most genome sequencing uses shotgun methods
    • clone-by-clone is mostly historical
    • much of today’s reference genome content traces back to the public effort

Speakers or sources featured (named in the subtitles)

  • Frederick Sanger
  • Smith–Waterman (algorithm referenced; not named as a person in the subtitles)
  • Todd Golub (Dana-Farber / Broad Institute mentioned in the leukemia study context)
  • Craig Venter (Celera Genomics)
  • NCBI (National Center for Biotechnology Information) (institution)
  • AlphaFold / DeepMind (referred to as “alphago company”)
  • CASP organizers / experts (institutional group)
  • Clinton (Bill Clinton, mentioned regarding coordination)
  • Jerry Rubin (example early sequencing effort / PhD work referenced)
  • Affymetrix (company mentioned for commercial microarrays)

Original video