Video summary

The SEO Audit Process I'd Use in 2026

Main summary

Key takeaways

Educational

Main ideas / lessons (SEO audit “battle plan”)

  • Agencies often charge a lot for SEO audits, but you can follow a structured, step-by-step process to create your own data-driven plan.
  • The audit emphasizes starting with crawl/indexing issues, then moving through internal linking, performance, and technical/content hygiene.
  • It repeatedly blends:
    • Quantitative data from crawling and analytics tools, and
    • Qualitative/intent-based judgment when diagnosing content and opportunities.
  • Later steps expand from classic SEO fixes into topic dominance and AI/LLM retrieval optimization (getting mentioned/covered across many platforms and improving citations AI uses).

Methodology: step-by-step SEO audit process (detailed)

1) Find crawling and indexing opportunities

Create a Google Sheet

  • Add a first tab named “crawl”.
  • (A template is mentioned as existing in “Gotcha SEO Academy,” but the approach demonstrates building from scratch.)

Run a crawl with Screaming Frog

  • Strongly recommended for the audit (speaker has used it for ~a decade).

Configure Screaming Frog (quickest, useful crawl)

  • Crawl configuration
    • Uncheck/disable options that add heavy load (e.g., resource links).
    • Keep remaining settings aligned with the speaker’s “simple” configuration.
  • Extraction
    • Use standard settings.
    • Structured data auditing is optional and site-dependent (more useful for local/ecommerce).
  • Duplicates / analysis enhancements
    • Enable near duplicates
    • Enable spell & grammar check
    • Enable embedding-related options:
      • semantic similarity
      • low relevance

Set up API access (minimum set recommended = 3)

  • Google Analytics 4
  • Google Search Console
  • PageSpeed Insights
  • (Optionally integrate AI platforms later; starting without them is acceptable.)

Run the crawl

  • Export/import results into the Google Sheet.
  • Keep the crawl open for later visualizations and steps.

2) Find internal linking opportunities

  • From the crawl export in the Google Sheet:
    • Freeze the header row
    • Add filters
    • Focus on indexable content
  • Prioritize columns for internal linking:
    • Crawl depth
      • Flag pages deeper than 3 clicks
      • Why: deeper pages are harder to crawl/index → can’t rank if not indexed.
      • Definition: crawl depth = number of clicks to reach a URL.
    • Unique in links (internal links count)
      • Flag pages with fewer than ~5 internal links (example threshold used)

Interpret patterns

  • Many deeply buried pages + low internal links suggests:
    • missing internal-link “injection” on relevant pages, and/or
    • weak topic support (content not strongly connected to the site’s theme)

3) Find page loading speed opportunities

  • Use the Performance Score column from the crawl results.
  • Mark pages with Performance Score < 80.

Why it matters

Even if it’s not a direct ranking factor:

  • Bad UX increases bounce / “pogo-sticking” risk.
  • AI/LLM crawlers have very limited time (speaker references < ~3 seconds).
  • Better speed improves crawling, indexing, and UX/conversions.

Practical diagnosis

  • Use Lighthouse (Chrome tooling/extension).
  • Test ~5–10 pages.
  • If desktop looks fine, test mobile (mobile is often the real issue).

Optimization expectation

  • Fixing one page can improve performance sitewide.

4) Find 404s and broken links

  • Review status codes in the crawl.

Not all 404s are bad

  • A 404 is acceptable if the page is intentionally removed and should disappear from the index.

“Bad” 404s

  • 404 pages that still have positive value signals.

Action criteria used

  • Filter to 404 pages
  • Check whether they have KPIs such as:
    • traffic
    • Search Console impressions
    • clicks
    • engagement/event metrics

If a 404 has positive KPIs

  • Decide whether to redirect it or rework/restore content to recapture demand.

Impact concept

  • Leaving valuable pages at 404 can cause large impression losses.

5) Find thin content

  • Use word count as the initial flag.
  • Flag pages with < 500 words.

Important clarification

  • Flagging ≠ automatically adding junk content.
  • Put pages into investigate buckets for:
    • improvement
    • deletion
    • restructuring

6) Find duplicate content

Definition

  • Same exact content appears on more than one page.

Detection approach

  • Use Screaming Frog’s near-duplicate/duplicate match data.
    • If configured properly, it may show few/no near duplicates (example: none found on their site).
  • Verification method:
    • Run a Siteliner crawl for duplicate detection.
    • If match percentage < 30%, it’s likely not a big concern.
    • Sort by match percentage and check for:
      • exact duplicates
      • highly similar (“near-duplicate”) pages

Potential action

  • If similarity is high but not exact, make pages more unique.

7) Find keyword cannibalization

Definition

  • Two pages compete for the same keywords with the same intent.

Clarification

  • Pages can target similar keywords (singular/plural, same seed topic) without cannibalizing if intent differs.

Intent-based example

  • A commercial “sales” page and an investigative “buyer guide” page can work together rather than compete—depending on keyword variants and intent.

Simple identification approach

  • In the crawl:
    • filter/search by title text contains a seed term
    • check whether multiple pages target the same modifiers/intent and thus truly compete

Rule of thumb

  • You can share seed keywords if intent modifiers make pages genuinely different.
  • Still recommended: keep one core topic per page.

8) Find irrelevant content

Goal

  • Keep topic focus and “subject matter expertise” tight.
  • Avoid drifting into topics that don’t belong to the site theme (harder to build strong topic clusters).

Detection method

  • Search/filter by title tags that do not contain the core theme term (example used: “SEO”).

Evaluate candidates

  • Even if such pages rank and get traffic, consider:
    • redirecting
    • deleting
    • reworking to become relevant

If irrelevant pages have positive KPIs

  • Prioritize preserving value while improving relevance.

9) Find weak clusters

Cluster definition

  • A group of pages around one topic, with support assets linking back to a core page.

Process

  • Identify both strong and weak clusters.
  • If a topic underperforms, examine how strong the support assets are.

Example logic

  • One cluster might have ~20 supporting pages (moderately strong).
  • Another pillar + extensive supporting reviews can reach ~75–100 assets (very strong).

Key takeaway

  • Often improves SEO to strengthen one cluster vs applying many scattered fixes.

10) Find content quality opportunities (qualitative diagnosis using data)

  • Use organic performance metrics to select pages to inspect.
  • Filter by low Google Search Console impressions.
  • Highlight the top 10–15 weakest pages by impressions.

Define “low quality / not worth keeping”

General criteria:

  • no traffic
  • no impressions
  • no backlinks
  • If older pages still have none of these signals: consider deletion or a strategy reset.

If the page has modest signals but is low-quality

  • Consider “beefing up” with more relevant information (e.g., bios, interviews, transcripts), if appropriate.

11) Find on-page SEO opportunities

Bare-minimum checklist for organic pages

Ensure the keyword appears in:

  • title tag
  • meta description
  • URL slug
  • H1
  • first sentence

NLP/related-topic coverage

  • Use NLP concepts to ensure coverage of relevant subtopics.

12) Find content optimization opportunities

  • For pages needing improvement (often those ranking around ~6 in the example):

Workflow

  • Run the keyword through Rankability (content optimizer).
  • Use “import URL content” + run optimizer.

Use the output to:

  • identify unused topics
  • add missing subtopics covered by competitors/top results

Timing concept

  • Revisit/refresh assets as they age (competitors and topical coverage change).

Expected impact

  • Content rework can improve ranking by ~2–4 spots (as claimed).

13) Find low-hanging fruits (keyword footprint sweet spot)

  • Use keyword position data to find pages near top results.
  • Focus on pages ranking approximately positions 2–15.

Process

  • Sort by average position and filter by relevant cluster/keyword set.
  • Verify in incognito/private mode to avoid personalization bias.

Prioritize fixes (usually)

  1. Rankability optimization (relevance)
  2. internal linking
  3. loading speed
  4. If still not improving: move to backlinks as the next tier

14) Find clustering opportunities (position 50+ established demand signals)

  • Define a bucket: keywords/pages ranking 50 and beyond.
    • Ignore newly published pages; the approach targets pages with demand signals.
  • In GSC:
    • sort/filter so you see cases with impressions but insufficient relevance/optimization.

Example insight

  • If there’s demand for “technical SEO training” but you only have academy content (no dedicated page):
    • build a dedicated landing page.

How to build

  • Send keyword to Rankability
  • Extract NLP keywords and create an outline
  • Decide AI vs manual based on competition:
    • lower competition → more AI
    • higher competition (e.g., SEO industry) → more manual tailoring

Core principle

  • After indexing, relevance is emphasized as the most important factor.

15) Find topic domination opportunities (cover a topic across many “surfaces” + AI retrieval)

Definition

  • Cover a topic on your site and also across other platforms that Google and AI systems index/retrieve from.

Goal

  • Influence both:
    • traditional first-page rankings
    • AI/ChatGPT-style retrieval outputs

Process

  • Use Search Console to find a keyword you already do well for.
  • Check which assets appear (book page, blog, etc.).
  • Identify missing coverage and improve the best candidate asset.

Next steps

  • Check YouTube:
    • if missing, create video content (YouTube matters because Google owns it)
  • Examine which external sources rank (e.g., Reddit, industry publications)
  • Track recurring subreddits and set alerts for mentions to respond/promote
  • Pitch inclusion on lists when feasible

16) Find critical citation opportunities (reverse-engineer AI citations)

Concept

  • AI answers often use RAG (retrieval) from sources it can cite.

Process

  • For the target topic/keyword:
    • run the query on multiple AI platforms:
      • ChatGPT
      • Perplexity
      • Claude
      • Grock
    • check each platform’s citations
    • identify gaps (where your brand/site isn’t cited)
    • outreach to sources missing your citation

Why multi-platform

  • Each system retrieves slightly differently, producing unique citation gaps.

17) Find competitor’s top performing pages (use proof, then replicate framework)

Rationale

  • AI-generated ideas are less reliable than pages with measurable link attraction.

Process

  • Use Ahrefs or Semrush
  • Go to “best by links” for a competitor
  • Identify page frameworks correlated with high link acquisition

Noted pattern

  • Often statistics-driven pages attract links.

Strategy

  • Replicate the proven framework, but keep it unique.
  • Consider transferring successful templates from other industries to differentiate.

18) Find misinformation about your brand

Goal

  • Detect inaccurate or inconsistent brand info that can harm trust, retrieval, or recommendations.

Quick method

  • Use a GPT/plugin approach (built by the speaker) to probe what the model knows from training data only (no web search).
  • Input the brand name.

Cautions

  • Not perfect; models can hallucinate if the brand isn’t well represented.

Action principle

  • If results show weak/incorrect knowledge, improve brand clarity and accuracy.

Speakers / sources featured

  • Speaker: Nathan Gotch
  • Tools/platforms mentioned:
    • Screaming Frog
    • Google Sheets
    • Gotcha SEO Academy (templates mentioned)
    • Google Analytics 4 (GA4)
    • Google Search Console (GSC)
    • PageSpeed Insights
    • Lighthouse (Chrome/extension)
    • Siteliner
    • Rankability (content optimizer)
    • Rankability / “detailed Chrome extension” (intent/optimization workflow)
    • Ahrefs / Semrush (competitor “best by links” research)
    • ChatGPT
    • Perplexity
    • Claude
    • Grock
    • YouTube
    • Reddit
    • Search Engine Journal (example: pitching lists)
    • A “free GPT”/probe tool built by the speaker (brand misinformation probing; name not clearly specified)

Original video