Video summary

ChatGPT vs Claude vs Gemini: Which AI Is Actually Best for Language Learning?

Main summary

Key takeaways

Educational

Main ideas / lessons

  • LLMs can help as “language tutors,” but they’re not reliable teachers by default. They’re optimized to sound helpful, not to ensure learning progress.
  • AI has a “fluency ceiling” if you only practice with it. Real conversation pressure and human interaction are hard to simulate with machines, so it should supplement real practice.
  • Most people misuse LLMs for language learning. The video argues that the goal of these tools is partly engagement/product use, not helping you learn correctly.
  • Evaluation matters: the creator tests multiple LLMs on concrete language-learning tasks (accuracy, vocabulary quality, transliteration consistency, volume, formatting, images, usefulness, and grammar explanations).
  • Use-case alignment: the “best” model depends on what you want it to do inside your learning system.
  • Overall conclusion of this test: Claude is the overall winner, with Gemini and Perplexity strong in different areas; ChatGPT is the weakest overall in this evaluation.

Methodology / evaluation criteria

The speaker compares ChatGPT, Claude, Gemini, and Perplexity as language-learning tools (not general chat or writing assistants). The evaluation focuses on tasks commonly useful for language learners:

  1. Factual information / hallucination risk

    • Question: “Does it provide factual information?”
    • Proxy test: whether it produces correct or incorrect language facts when the user might not know enough to verify.
    • Example trap task: request Hebrew verbs with a specific Nefal conjugation condition involving four roots (rare/unrealistic for standard Hebrew).
  2. Vocabulary quality & sample sentences (general usefulness)

    • Question: “Is the vocabulary useful and correct?”
    • Test idea: ask for relevant vocabulary and example sentences in a way that should produce learner-useful material.
    • Reference point: contrasts with Duolingo-style sentences that are “correct but not very helpful” (e.g., trivial sentences).
  3. Transliteration quality & consistency

    • Question: “Is transliteration accurate and consistent?”
    • Test emphasis: whether the model uses systems that match typical expectations for the learner.
    • Problem case: mismatched transliteration to pronunciation or inconsistent conventions.
  4. Volume / coverage

    • Question: “How many items does it return?”
    • Concern: insufficient coverage can be a problem even if items are individually correct.
    • Note: results can be repetitive, requiring deduping when importing into Anki.
  5. Formatting for learning workflows

    • Question: “Can it output in formats you can directly use?”
    • Specific focus: CSV export for spaced repetition workflows (Anki).
    • Examples of failure/variation:
      • ChatGPT might produce tables with heavy formatting/emojis.
      • Gemini provides better integration with Google Sheets.
      • Claude/others may provide CSV but require extra steps.
  6. Image generation (if relevant)

    • Question: “Can it create images, and how good/useful are they?”
    • Note: Claude is characterized as refusing or not producing images in this test, while others generate “labeled” images that may contain fake/unusable text.
  7. Usefulness of vocabulary (pragmatic usefulness)

    • Question: “Are the words actually helpful in real contexts?”
    • Test includes whether it provides more “useful real-world phrases” beyond the obvious textbook basics.
  8. Explanation of difficult grammar concepts

    • Bonus criterion: can it clearly and correctly explain challenging grammar?
    • Example difficult topics posed:
      • Deontic modality (with examples)
      • Hebrew vs Italian verb distinctions tied to stativitiy and historical change
      • The subjunctive
    • Additional observation: explanations vary—asking twice can yield different (but similar) answers per model.

What happened in the tests (main outcomes)

1) Factual information test (Hebrew verbs with four roots in Nefal-related context)

  • ChatGPT: produced plausible-sounding but incorrect answers (including category/conjugation confusion). The response looked believable unless the user already knew the grammar.
  • Perplexity: also produced incorrect outputs.
  • Gemini: hallucinated verbs (appeared to invent non-attested forms rather than merely confusing conjugation class).
  • Claude: initially returned some likely incorrect items, but it also corrected itself later and hedged where certainty was lacking.
  • Overall: ChatGPT and Perplexity failed; Gemini failed; Claude performed better due to self-correction/hedging, though still not perfect.

2) Vocabulary & sample sentences

  • ChatGPT: passed; responses were “fine” and usable, but not standout.
  • Gemini: similar “pass” level.
  • Perplexity: strong when it provides sources/links for examples; better when prompts combine domains (e.g., color + clothing + specific tenses).
  • Claude: especially strong at combining domains and producing useful structured outputs.

3) Transliteration

  • ChatGPT: acceptable/typical.
  • Gemini: problematic transliteration for Hebrew (misaligned with pronunciation; sometimes used conventions resembling reconstructed biblical Hebrew).
  • Claude & Perplexity: preferred/expected transliteration approach (as described by the creator).
  • All models: occasional odd artifacts (e.g., Cyrillic letters, inconsistent transliteration of certain characters).

4) Volume (quantity of items)

  • Gemini: lowest volume (often ~25 entries), even paid version.
  • ChatGPT: around ~35 average.
  • Perplexity: around ~43.
  • Claude: highest volume (often 70+ and sometimes 200+), with some repetition that required deduping (up to ~5%).

5) Formatting

  • Gemini: best for spreadsheet workflow—clear tables that connect to Google Sheets.
  • ChatGPT: often messy outputs (tables with many emojis).
  • Claude: provided CSV but in a way that might require manual saving/copying.
  • Perplexity: produced well-formatted CSV.

6) Images

  • Claude: no useful image generation in this test (refuses; creator describes it as avoiding a “losing game”).
  • ChatGPT / Perplexity / Gemini: produced labeled images, but:
    • labeled Hebrew text could be gibberish / fake
    • images became “uncanny” or contextually odd
  • Claude judged best here specifically by refusing (i.e., avoiding misleading content).

7) Usefulness of vocabulary

  • ChatGPT: useful but mostly obvious/basic items.
  • Gemini: similar “basic and fine.”
  • Claude: exceptionally useful expansions beyond textbook content; good phrases and contextually relevant examples.
  • Perplexity: best in providing more interesting real-world phrasing (creator gives examples like situational phrases in bathroom-related vocabulary).

8) Grammar explanations

  • All four generally handled difficult grammar explanations “well enough.”
  • ChatGPT: decent technical explanations and helpful memorization strategy suggestions.
  • Gemini: strong technical explanation; offered clear visual tables and memorization aids.
  • Perplexity: concise semantic-focused answers and included links to where to learn more.
  • Claude: technical explanation was good but not especially standout in the creator’s particular comparison.

Final ranking / conclusions

  • Overall winner: Claude
  • Runner-up: Gemini / Perplexity are positioned as strong depending on priorities, with:
    • Gemini strong at integration with Google products and grammar tables, plus image generation if labels aren’t important.
    • Perplexity strong at source citation and more interesting vocabulary.
  • Overall weakness: ChatGPT (weaker across the evaluation overall, including accuracy concerns).

How the creator says to use LLMs (high-level, non-spoiler framing)

Use an LLM inside a broader structured learning system, not as the sole tutor.

The speaker repeatedly signals that the “how to implement” details are covered in:

  • a next video
  • and a planned course on self-directed language study and LLM usage.

Speakers / sources featured

  • Speaker: Dr. Taylor Jones (linguist; PhD from the University of Pennsylvania)
  • Companies/models compared: ChatGPT, Claude, Gemini, Perplexity
  • Referenced tools/resources: Duolingo, Lingod, Anki, Google Sheets, WordHippo, Reverso Context, “tatttoo corpus of sentences” (as mentioned), Google (image generation), languagejones.com (course/guide site), Patreon (support page).

Original video