Video summary
Internship Program 9th Class | Target and Ligand Identification
Main summary
Key takeaways
Main Ideas / Lessons Conveyed
1) Where “Target and Ligand Identification” Fits in the Drug Discovery Pipeline
This session focuses on an early decision-making step in the drug discovery pipeline:
- Target identification
- Ligand (compound) selection
These choices are foundational for later computational stages such as:
- Molecular docking
- Virtual screening
- Development of therapeutic compounds
A high-level workflow overview:
- Use multiple databases to identify targets
- Retrieve 3D protein structures (e.g., from PDB)
- Identify active sites (using CASTp)
- Download ligand libraries from ligand databases (e.g., ChEMBL)
- Apply filters (including ADMET and other funnel steps) before final docking
2) What Makes a “Good/Ideal Drug Target”
Target selection is framed through several attributes:
Druggability
Key question: Can small molecules bind and modulate the protein?
Pocket/binding site considerations discussed:
- Binding pockets should accommodate ligands
- Pocket volume guidance:
- Suggested as at least ~300 ų, with an upper range later mentioned for CASTp usage
- Pocket characteristics:
- Hydrophobic pockets for anchoring/stability
- Conformational flexibility to support induced fit
- Ability to support hydrogen bonding via proper residue arrangement
Tools mentioned for pocket detection/analysis:
- CASTp
- Other generic pocket tools
- Checks using “similar protein/pocket” approaches (subtitles partially unclear; examples like “Site reports / similar” were referenced)
Disease Relevance
Key question: Is the target causally linked to the disease mechanism/pathways?
Evidence types described:
- Pathogenic mutations and genetic evidence via:
- OMIM
- NCBI Gene
- Upregulation/downregulation and pathway involvement
- Experimental support such as:
- knockout/knockdown phenotypes
- observed therapeutic benefit from modulating the target
Selectivity
Drugs should preferentially act on the target while avoiding harm to:
- healthy/normal tissues
- critical pathways
Structural Availability
Requirement: the target ideally has a high-confidence 3D structure for docking.
Preferred/allowed structure sources:
- X-ray crystallography
- NMR
- cryo-EM, with stated resolution guidance:
- < 2.5 Å preferred
- 2–3.5 Å acceptable for cryo-EM models (as stated)
- If experimental structures are missing:
- Use AI-based structure prediction tools:
- AlphaFold (described as a last resort, not first priority)
- Other tools were mentioned but were unclear in subtitles
- Use AI-based structure prediction tools:
Co-crystallized Ligand Presence (Active Site Definition)
If a PDB structure contains a co-crystallized ligand:
- It helps identify realistic binding residues
- It supports choosing the correct active site for docking
3) Databases Used for Target Validation and Enrichment (Conceptual Map)
The session emphasizes using different databases for evidence at different levels:
- GeneCards / NCBI Gene
- gene-level information and curated links to related resources
- OMIM
- genetic disease/variant context and inheritance-related information
- ChEMBL / UniProt / PubMed / ClinicalTrials (generally mentioned)
- biochemical evidence, assays, activity, and supporting literature
- ClinicalTrials for clinical evidence of target modulation (as described)
4) Protein Structure Retrieval from PDB (Practical Criteria + Interpretation)
Core retrieval rules emphasized:
-
Species/organism must match the target context
- Example: for human targets, prefer Homo sapiens structures over rat/mouse
-
Methods prioritized
- X-ray crystallography and NMR prioritized
- Resolution guidance reiterated:
- < 2.5 Å preferred (stated)
- ~2–3 Å tolerated if needed (stated)
- Quality metrics:
- R factor and R free expected to be < 0.25 (as stated)
What to look for in PDB files:
- PDB header content:
- structure classification
- deposition date
- resolution and atomic content
- ATOMS / HETATM
- ligands, ions, water, cofactors appear under hetero atoms
- Connectivity records for non-standard bonds
Apo vs. Holo vs. AlphaFold:
- Apo: protein without bound ligand
- Holo: protein with ligand bound
- Docking guidance:
- Holo preferred for active pocket targeting/inhibition
- Apo can still be used (e.g., blind/global docking or exploring pockets)
- If no experimental structures exist:
- Use AlphaFold
- AlphaFold quality threshold:
- PLDDT > 70
- For multi-chain proteins:
- analyze only chain(s) with score > 70 for domain-specific analysis
5) Active Site Identification Using CASTp (Instruction-like Criteria)
Active site detection guidance:
- Use CASTp to identify binding pockets/active sites
Pocket volume constraints (as stated):
- Target pockets between ~300 and 1000 ų
- Avoid pockets too small (stated emphasis not to go below ~300 ų; subtitles also mention “not below 100” and “not below 300”)
Pocket composition expectations:
- Hydrophobic elements for non-polar anchoring
- Polar/charge regions aligned for hydrogen bonding
- Proper residue arrangement
- Additional confidence concept:
- evolutionary sequence conservation
6) Ligand Sourcing and Bioactivity Selection (Databases + Numeric Assay Meanings)
Ligand database options explicitly mentioned:
- ChEMBL (preferred for bioactivity data)
- ZINC20 (for purchasable compounds; very large collection)
- DrugBank (drug repurposing and approved drugs)
- PAPKA / “PubChem” mentioned (SMILES formats referenced; subtitle specifics unclear)
- PDBQT referenced for docking input formats (covered later)
From ChEMBL, assay/property handling discussed:
- Selecting target-specific bioactive ligands with experimental activity
Common potency/effect terms:
- IC50: concentration causing 50% inhibition
- EC50: effective concentration for 50% activation (agonist context)
- KD: dissociation constant
- Ki: inhibition constant
Rule of thumb:
- Smaller values generally indicate stronger potency (as described)
7) Virtual Screening “Funnel” Concept (Pipeline Logic)
A funnel-style workflow:
- Start with compounds from ligand databases (e.g., ChEMBL / PubChem actives)
- Apply:
- Lipinski Rule of Five filters
- Safety / ADMET-related filters
- Optional diversity filter (possibly combined with others)
- Perform final docking
- Then proceed to molecular docking (practical follow-up indicated)
8) Lipinski Rule of Five (Explicit Criteria)
Drug-likeness filtering criteria:
- Molecular weight (MW) ≤ 500 Da
- logP ≤ 5
- Hydrogen bond donors ≤ 5
- Hydrogen bond acceptors ≤ 10
Rationale mentioned:
- Called the “Rule of Five” because thresholds relate to multiples of 5
9) ADMET Prediction Tools (And Stated Plan)
ADMET profiling is performed after ligand and target sourcing.
Tools mentioned:
- ADMETlab 3.0 (explicitly stated as the one they will use)
- ADMETlab 2 / Deep PK also referenced
The session notes:
- Today emphasizes ADMET filtering conceptually
- More details may be covered in a next session
Practical Workflow Demonstrated (Step-by-Step Bullets)
Practical session: Using Gene Cards → UniProt → PDB → Structure Preparation
-
Literature review / target selection
- Identify a target gene/protein based on:
- literature review
- literature gap search (optional)
- Instructor suggested literature summarization tools:
- Consensus AI
- also mentioned Google Scholar and PubMed
- Identify a target gene/protein based on:
-
Gene Cards (target gene exploration)
- Open GeneCards (Human Gene Database)
- Search by gene/protein symbol (example used: ATM)
- Review key sections:
- gene overview
- identifiers/summaries
- genomics location
- protein details
- function and pathways
- expression and subcellular localization
- disease/clinical annotations and variants
- orthologs/paralogs explanations
- interaction network visualization
- Use GeneCards as a starting point for linked resources (e.g., NCBI Gene, PDB, UniProt, related links)
-
UniProt (protein-specific annotation)
- Copy/save key IDs from GeneCards (e.g., UniProt IDs)
- Open UniProt using the linked protein record
- Review:
- protein function and taxonomy
- subcellular location
- disease and variants
- post-translational modifications / processing
- expression data
- reaction/catalytic activity descriptions (if present)
- sequence availability
- Note: UniProt may provide previews/links, but PDB is used for full experimental structures
-
Retrieve experimental structure from RCSB PDB
- Use UniProt information to find relevant PDB structures
- Choose a structure based on:
- resolution (prefer around ~2.5 Å or comparable)
- experimental method (e.g., cryo-EM if available)
- apo/holo context (co-crystallized ligands noted)
- Review PDB entry details:
- macromolecule name, organism
- deposition/release information
- method and resolution
- chains present
- bound ligands/ions and cofactors
-
Structure preparation for docking/visualization (PyMOL + preparation logic)
- Download legacy PDB format for docking pipeline compatibility
- Load in PyMOL (and possibly Discovery Studio)
- PyMOL tasks:
- identify ligands/ions (e.g., magnesium, zinc) and waters
- optionally delete unwanted components for docking accuracy
- water guidance:
- remove most waters
- retain only structurally bridged waters when they support binding without harming docking
- water guidance:
- Coloring/representation:
- color by secondary structure (alpha helices, beta sheets, loops) and other selection methods
- Protein preparation approach noted:
- for docking workflow, the instructor prefers AutoDock tools for cleaning + generating PDBQT and then grids
-
Assignment/output expectation (report creation)
- Create a report including:
- screenshots from GeneCards analysis
- screenshots from UniProt analysis
- screenshots from PDB entry exploration
- optionally, a 2D structure image of the target
- Submit weekly; assignments correspond to tools covered in each session
- Create a report including:
Speakers / Sources Featured (As Stated or Clearly Implied)
Speakers
- Miss Adiba Fatima (also spelled Adeeba Fatima / Adiba Fatima in subtitles) — Founder, Biotic Catalyst; computation biology mentor; delivered the session.
- Host/Moderator (unnamed in subtitles) — opened session, handled announcements and attendance, introduced speaker, and delivered some closing remarks.
Sources / Organizations / Databases Mentioned
- WhiteNOVA International Society for Sciences
- GeneCards
- UniProt
- RCSB PDB / Protein Data Bank (PDB)
- ChEMBL
- ZINC20
- DrugBank
- OMIM
- NCBI Gene
- UniProt/knowledge base (UniProtKB)
- Pfam / InterPro (domain resources mentioned)
- AlphaFold
- ADMETlab 3.0
- Lipinski Rule of Five (Lipinski et al., 2001)
- CASTp
- AutoDock Tools
- PyMOL
- Discovery Studio
- PubMed
- ClinicalTrials.gov (referred to as “Clinical Trials Core website”)
- Google Scholar and Consensus AI (literature search aids)