Video summary

Internship Program 9th Class | Target and Ligand Identification

Main summary

Key takeaways

Educational

Main Ideas / Lessons Conveyed

1) Where “Target and Ligand Identification” Fits in the Drug Discovery Pipeline

This session focuses on an early decision-making step in the drug discovery pipeline:

  • Target identification
  • Ligand (compound) selection

These choices are foundational for later computational stages such as:

  • Molecular docking
  • Virtual screening
  • Development of therapeutic compounds

A high-level workflow overview:

  • Use multiple databases to identify targets
  • Retrieve 3D protein structures (e.g., from PDB)
  • Identify active sites (using CASTp)
  • Download ligand libraries from ligand databases (e.g., ChEMBL)
  • Apply filters (including ADMET and other funnel steps) before final docking

2) What Makes a “Good/Ideal Drug Target”

Target selection is framed through several attributes:

Druggability

Key question: Can small molecules bind and modulate the protein?

Pocket/binding site considerations discussed:

  • Binding pockets should accommodate ligands
  • Pocket volume guidance:
    • Suggested as at least ~300 ų, with an upper range later mentioned for CASTp usage
  • Pocket characteristics:
    • Hydrophobic pockets for anchoring/stability
    • Conformational flexibility to support induced fit
    • Ability to support hydrogen bonding via proper residue arrangement

Tools mentioned for pocket detection/analysis:

  • CASTp
  • Other generic pocket tools
  • Checks using “similar protein/pocket” approaches (subtitles partially unclear; examples like “Site reports / similar” were referenced)

Disease Relevance

Key question: Is the target causally linked to the disease mechanism/pathways?

Evidence types described:

  • Pathogenic mutations and genetic evidence via:
    • OMIM
    • NCBI Gene
  • Upregulation/downregulation and pathway involvement
  • Experimental support such as:
    • knockout/knockdown phenotypes
    • observed therapeutic benefit from modulating the target

Selectivity

Drugs should preferentially act on the target while avoiding harm to:

  • healthy/normal tissues
  • critical pathways

Structural Availability

Requirement: the target ideally has a high-confidence 3D structure for docking.

Preferred/allowed structure sources:

  • X-ray crystallography
  • NMR
  • cryo-EM, with stated resolution guidance:
    • < 2.5 Å preferred
    • 2–3.5 Å acceptable for cryo-EM models (as stated)
  • If experimental structures are missing:
    • Use AI-based structure prediction tools:
      • AlphaFold (described as a last resort, not first priority)
      • Other tools were mentioned but were unclear in subtitles

Co-crystallized Ligand Presence (Active Site Definition)

If a PDB structure contains a co-crystallized ligand:

  • It helps identify realistic binding residues
  • It supports choosing the correct active site for docking

3) Databases Used for Target Validation and Enrichment (Conceptual Map)

The session emphasizes using different databases for evidence at different levels:

  • GeneCards / NCBI Gene
    • gene-level information and curated links to related resources
  • OMIM
    • genetic disease/variant context and inheritance-related information
  • ChEMBL / UniProt / PubMed / ClinicalTrials (generally mentioned)
    • biochemical evidence, assays, activity, and supporting literature
    • ClinicalTrials for clinical evidence of target modulation (as described)

4) Protein Structure Retrieval from PDB (Practical Criteria + Interpretation)

Core retrieval rules emphasized:

  • Species/organism must match the target context

    • Example: for human targets, prefer Homo sapiens structures over rat/mouse
  • Methods prioritized

    • X-ray crystallography and NMR prioritized
    • Resolution guidance reiterated:
      • < 2.5 Å preferred (stated)
      • ~2–3 Å tolerated if needed (stated)
    • Quality metrics:
      • R factor and R free expected to be < 0.25 (as stated)

What to look for in PDB files:

  • PDB header content:
    • structure classification
    • deposition date
    • resolution and atomic content
  • ATOMS / HETATM
    • ligands, ions, water, cofactors appear under hetero atoms
  • Connectivity records for non-standard bonds

Apo vs. Holo vs. AlphaFold:

  • Apo: protein without bound ligand
  • Holo: protein with ligand bound
  • Docking guidance:
    • Holo preferred for active pocket targeting/inhibition
    • Apo can still be used (e.g., blind/global docking or exploring pockets)
  • If no experimental structures exist:
    • Use AlphaFold
    • AlphaFold quality threshold:
      • PLDDT > 70
    • For multi-chain proteins:
      • analyze only chain(s) with score > 70 for domain-specific analysis

5) Active Site Identification Using CASTp (Instruction-like Criteria)

Active site detection guidance:

  • Use CASTp to identify binding pockets/active sites

Pocket volume constraints (as stated):

  • Target pockets between ~300 and 1000 ų
  • Avoid pockets too small (stated emphasis not to go below ~300 ų; subtitles also mention “not below 100” and “not below 300”)

Pocket composition expectations:

  • Hydrophobic elements for non-polar anchoring
  • Polar/charge regions aligned for hydrogen bonding
  • Proper residue arrangement
  • Additional confidence concept:
    • evolutionary sequence conservation

6) Ligand Sourcing and Bioactivity Selection (Databases + Numeric Assay Meanings)

Ligand database options explicitly mentioned:

  • ChEMBL (preferred for bioactivity data)
  • ZINC20 (for purchasable compounds; very large collection)
  • DrugBank (drug repurposing and approved drugs)
  • PAPKA / “PubChem” mentioned (SMILES formats referenced; subtitle specifics unclear)
  • PDBQT referenced for docking input formats (covered later)

From ChEMBL, assay/property handling discussed:

  • Selecting target-specific bioactive ligands with experimental activity

Common potency/effect terms:

  • IC50: concentration causing 50% inhibition
  • EC50: effective concentration for 50% activation (agonist context)
  • KD: dissociation constant
  • Ki: inhibition constant

Rule of thumb:

  • Smaller values generally indicate stronger potency (as described)

7) Virtual Screening “Funnel” Concept (Pipeline Logic)

A funnel-style workflow:

  • Start with compounds from ligand databases (e.g., ChEMBL / PubChem actives)
  • Apply:
    • Lipinski Rule of Five filters
    • Safety / ADMET-related filters
    • Optional diversity filter (possibly combined with others)
  • Perform final docking
  • Then proceed to molecular docking (practical follow-up indicated)

8) Lipinski Rule of Five (Explicit Criteria)

Drug-likeness filtering criteria:

  • Molecular weight (MW) ≤ 500 Da
  • logP ≤ 5
  • Hydrogen bond donors ≤ 5
  • Hydrogen bond acceptors ≤ 10

Rationale mentioned:

  • Called the “Rule of Five” because thresholds relate to multiples of 5

9) ADMET Prediction Tools (And Stated Plan)

ADMET profiling is performed after ligand and target sourcing.

Tools mentioned:

  • ADMETlab 3.0 (explicitly stated as the one they will use)
  • ADMETlab 2 / Deep PK also referenced

The session notes:

  • Today emphasizes ADMET filtering conceptually
  • More details may be covered in a next session

Practical Workflow Demonstrated (Step-by-Step Bullets)

Practical session: Using Gene Cards → UniProt → PDB → Structure Preparation

  1. Literature review / target selection

    • Identify a target gene/protein based on:
      • literature review
      • literature gap search (optional)
    • Instructor suggested literature summarization tools:
      • Consensus AI
      • also mentioned Google Scholar and PubMed
  2. Gene Cards (target gene exploration)

    • Open GeneCards (Human Gene Database)
    • Search by gene/protein symbol (example used: ATM)
    • Review key sections:
      • gene overview
      • identifiers/summaries
      • genomics location
      • protein details
      • function and pathways
      • expression and subcellular localization
      • disease/clinical annotations and variants
      • orthologs/paralogs explanations
      • interaction network visualization
    • Use GeneCards as a starting point for linked resources (e.g., NCBI Gene, PDB, UniProt, related links)
  3. UniProt (protein-specific annotation)

    • Copy/save key IDs from GeneCards (e.g., UniProt IDs)
    • Open UniProt using the linked protein record
    • Review:
      • protein function and taxonomy
      • subcellular location
      • disease and variants
      • post-translational modifications / processing
      • expression data
      • reaction/catalytic activity descriptions (if present)
      • sequence availability
    • Note: UniProt may provide previews/links, but PDB is used for full experimental structures
  4. Retrieve experimental structure from RCSB PDB

    • Use UniProt information to find relevant PDB structures
    • Choose a structure based on:
      • resolution (prefer around ~2.5 Å or comparable)
      • experimental method (e.g., cryo-EM if available)
      • apo/holo context (co-crystallized ligands noted)
    • Review PDB entry details:
      • macromolecule name, organism
      • deposition/release information
      • method and resolution
      • chains present
      • bound ligands/ions and cofactors
  5. Structure preparation for docking/visualization (PyMOL + preparation logic)

    • Download legacy PDB format for docking pipeline compatibility
    • Load in PyMOL (and possibly Discovery Studio)
    • PyMOL tasks:
      • identify ligands/ions (e.g., magnesium, zinc) and waters
      • optionally delete unwanted components for docking accuracy
        • water guidance:
          • remove most waters
          • retain only structurally bridged waters when they support binding without harming docking
    • Coloring/representation:
      • color by secondary structure (alpha helices, beta sheets, loops) and other selection methods
    • Protein preparation approach noted:
      • for docking workflow, the instructor prefers AutoDock tools for cleaning + generating PDBQT and then grids
  6. Assignment/output expectation (report creation)

    • Create a report including:
      • screenshots from GeneCards analysis
      • screenshots from UniProt analysis
      • screenshots from PDB entry exploration
      • optionally, a 2D structure image of the target
    • Submit weekly; assignments correspond to tools covered in each session

Speakers / Sources Featured (As Stated or Clearly Implied)

Speakers

  • Miss Adiba Fatima (also spelled Adeeba Fatima / Adiba Fatima in subtitles) — Founder, Biotic Catalyst; computation biology mentor; delivered the session.
  • Host/Moderator (unnamed in subtitles) — opened session, handled announcements and attendance, introduced speaker, and delivered some closing remarks.

Sources / Organizations / Databases Mentioned

  • WhiteNOVA International Society for Sciences
  • GeneCards
  • UniProt
  • RCSB PDB / Protein Data Bank (PDB)
  • ChEMBL
  • ZINC20
  • DrugBank
  • OMIM
  • NCBI Gene
  • UniProt/knowledge base (UniProtKB)
  • Pfam / InterPro (domain resources mentioned)
  • AlphaFold
  • ADMETlab 3.0
  • Lipinski Rule of Five (Lipinski et al., 2001)
  • CASTp
  • AutoDock Tools
  • PyMOL
  • Discovery Studio
  • PubMed
  • ClinicalTrials.gov (referred to as “Clinical Trials Core website”)
  • Google Scholar and Consensus AI (literature search aids)

Original video