Video summary
2020 STAT115 Lect1.1 Bioinfo History
Main summary
Key takeaways
Main ideas / lessons conveyed
- Bioinformatics and computational biology evolved in waves, driven by new biological data types and the need for new computational methods to interpret them.
- Core theme: biology advances → data becomes available → computational algorithms become necessary to:
- compare
- assemble
- classify
- predict biological function and/or structure
Timeline of major “waves” and key methods
1) Protein sequence & structure (early “wave”)
Key milestones and methods
-
1965 – Protein sequencing
- Frederick Sanger invented protein sequencing methods.
- Enabled the key goal: determining whether a new protein sequence is similar to previously known sequences.
-
Sequence alignment algorithms
- Need: rapidly compare a new sequence against all sequences in existence.
- Smith–Waterman: a famous sequence alignment algorithm.
-
1973 – Protein structure database
- Creation of the PDB (Protein Data Bank).
- Rationale: protein function depends on 3D structure, not only sequence.
- As crystallography structures accumulate, PDB stores them for later querying.
-
1990 – Faster similarity search: BLAST
- Instead of exhaustive alignment, BLAST searches a pre-existing protein database.
- Result: much quicker retrieval of similar sequences.
-
1994 – Motif discovery / sequence patterns
- With enough sequences, researchers identify sequence motifs (recurring patterns).
- Motifs can suggest function (e.g., motifs linked to ATP binding or catalytic activity).
- A “Blocks” database is mentioned as summarizing motif/function relationships from known sequences.
Protein structure prediction competitions
-
CASP (Critical Assessment of protein Structure Prediction)
- Motivation: if many groups publish “best” predictors, the field needs objective evaluation.
- Runs every two years.
- Predictions are made without knowing the true structures during the competition.
- After a delay, experimentally solved structures become gold standards used to score methods.
-
Impact of CASP
- Raised the bar: methods can’t rely only on claims of being “the best.”
- Some computational biologists left the field after seeing weaknesses against gold-standard evaluation.
-
2018 (CASP13 era) / Deep learning breakthrough
- AlphaFold (DeepMind; “alphago” referenced) achieved strong performance:
- Best predictions for 25 out of 43 proteins (as described)
- Second place for 3 out of 43 proteins (as described)
- Also described as predicting the first structure within hours
- The story emphasizes that competitions can validate breakthroughs and accelerate progress.
- AlphaFold (DeepMind; “alphago” referenced) achieved strong performance:
2) Microarrays & gene expression (second “wave”)
Concept and motivation
- A professor introduced microarrays to the class as an exciting idea.
- Core concept: measure gene expression in bulk, rather than one RNA/gene at a time.
Microarray development and workflow
- Early phase:
- Spotted arrays on glass slides
- Initially manual, later improved with robotic spotting
- Scaling up:
- From hundreds/thousands of targets to commercial systems (e.g., Affymetrix)
- With millions of probes on small arrays
Operational workflow (as described):
- Apply a sample mixture of RNA from cells to the microarray.
- RNA hybridizes to matching probes.
- Wash away non-specific binding.
- Detect light signals to infer which genes are expressed.
Computational analysis (1995–2010)
- Many microarray analysis methods were developed roughly between 1995 and 2010.
Example application: leukemia diagnosis/classification
- Distinguishing similar leukemia types:
- ALL vs AML
- Similar by pathology but different in treatment/outcome
- Method described:
- Use microarray expression profiles across ~20,000 genes
- Identify gene sets/signatures consistently higher in one type vs the other
- For a new patient:
- run the same microarray
- use the signature gene-set expression to predict diagnosis
- guide treatment selection
3) DNA sequencing & genomics (third “wave” and moving focus)
Enabling concepts
DNA sequencing became feasible through:
- understanding DNA (double helix context)
- recombinant DNA (cutting/stitching)
- Sanger sequencing
- PCR to amplify DNA when starting material is limited
Databases and reuse: NCBI + BLAST
- NCBI (National Center for Biotechnology Information) built/maintained sequence databases.
- BLAST used repeatedly to determine whether newly sequenced fragments are:
- already known, or
- novel (then deposited into NCBI)
Speedup and sequencing “race”
- Anecdote: sequencing one gene once could take years (e.g., yeast gene/7S RNA described as a long PhD effort).
- Modern sequencing/deposition/review can be described as extremely fast (e.g., “three seconds”).
Human Genome Project (public effort)
- Began around 1990, with target completion by 2005 (as described).
- Analogy:
- 23 chromosome pairs = 23 “books”
- divide work into smaller pieces (chapters/pages), then assemble later
- Early decade emphasis:
- first ~10 years focused on correct mapping/division relationships before deep sequencing
Celera Genomics (private effort; “whole-genome shotgun”)
- Mentioned as led by Craig Venter.
- Strategic shift:
- avoid extensive division/mapping (“don’t divide into books”)
- make copies → shred into small pieces
- assemble computationally using overlaps
Informatics as a decisive factor
- Even if public work used structured pieces, informatics still mattered.
- For shotgun sequencing, informatics becomes even more critical:
- it generates many short reads (e.g., ~1000 base pairs as described)
- these must be assembled computationally into longer contiguous genomes
Joint milestone announcement
- White House / Clinton helped coordinate public and private efforts to avoid destructive dynamics.
- Joint announcement:
- Spring of 2000: “draft human genome” issued (with gaps)
- additional work continued for ~3 years to fill remaining gaps
Broader lesson / outcome
- Major beneficiaries: scientists and the broader community via data access.
- Long-term impact:
- stimulated gene prediction and downstream research/drug discovery
- Current practice (as described):
- most genome sequencing uses shotgun methods
- clone-by-clone is mostly historical
- much of today’s reference genome content traces back to the public effort
Speakers or sources featured (named in the subtitles)
- Frederick Sanger
- Smith–Waterman (algorithm referenced; not named as a person in the subtitles)
- Todd Golub (Dana-Farber / Broad Institute mentioned in the leukemia study context)
- Craig Venter (Celera Genomics)
- NCBI (National Center for Biotechnology Information) (institution)
- AlphaFold / DeepMind (referred to as “alphago company”)
- CASP organizers / experts (institutional group)
- Clinton (Bill Clinton, mentioned regarding coordination)
- Jerry Rubin (example early sequencing effort / PhD work referenced)
- Affymetrix (company mentioned for commercial microarrays)