Video summary
ChatGPT vs Claude vs Gemini: Which AI Is Actually Best for Language Learning?
Main summary
Key takeaways
Main ideas / lessons
- LLMs can help as “language tutors,” but they’re not reliable teachers by default. They’re optimized to sound helpful, not to ensure learning progress.
- AI has a “fluency ceiling” if you only practice with it. Real conversation pressure and human interaction are hard to simulate with machines, so it should supplement real practice.
- Most people misuse LLMs for language learning. The video argues that the goal of these tools is partly engagement/product use, not helping you learn correctly.
- Evaluation matters: the creator tests multiple LLMs on concrete language-learning tasks (accuracy, vocabulary quality, transliteration consistency, volume, formatting, images, usefulness, and grammar explanations).
- Use-case alignment: the “best” model depends on what you want it to do inside your learning system.
- Overall conclusion of this test: Claude is the overall winner, with Gemini and Perplexity strong in different areas; ChatGPT is the weakest overall in this evaluation.
Methodology / evaluation criteria
The speaker compares ChatGPT, Claude, Gemini, and Perplexity as language-learning tools (not general chat or writing assistants). The evaluation focuses on tasks commonly useful for language learners:
-
Factual information / hallucination risk
- Question: “Does it provide factual information?”
- Proxy test: whether it produces correct or incorrect language facts when the user might not know enough to verify.
- Example trap task: request Hebrew verbs with a specific Nefal conjugation condition involving four roots (rare/unrealistic for standard Hebrew).
-
Vocabulary quality & sample sentences (general usefulness)
- Question: “Is the vocabulary useful and correct?”
- Test idea: ask for relevant vocabulary and example sentences in a way that should produce learner-useful material.
- Reference point: contrasts with Duolingo-style sentences that are “correct but not very helpful” (e.g., trivial sentences).
-
Transliteration quality & consistency
- Question: “Is transliteration accurate and consistent?”
- Test emphasis: whether the model uses systems that match typical expectations for the learner.
- Problem case: mismatched transliteration to pronunciation or inconsistent conventions.
-
Volume / coverage
- Question: “How many items does it return?”
- Concern: insufficient coverage can be a problem even if items are individually correct.
- Note: results can be repetitive, requiring deduping when importing into Anki.
-
Formatting for learning workflows
- Question: “Can it output in formats you can directly use?”
- Specific focus: CSV export for spaced repetition workflows (Anki).
- Examples of failure/variation:
- ChatGPT might produce tables with heavy formatting/emojis.
- Gemini provides better integration with Google Sheets.
- Claude/others may provide CSV but require extra steps.
-
Image generation (if relevant)
- Question: “Can it create images, and how good/useful are they?”
- Note: Claude is characterized as refusing or not producing images in this test, while others generate “labeled” images that may contain fake/unusable text.
-
Usefulness of vocabulary (pragmatic usefulness)
- Question: “Are the words actually helpful in real contexts?”
- Test includes whether it provides more “useful real-world phrases” beyond the obvious textbook basics.
-
Explanation of difficult grammar concepts
- Bonus criterion: can it clearly and correctly explain challenging grammar?
- Example difficult topics posed:
- Deontic modality (with examples)
- Hebrew vs Italian verb distinctions tied to stativitiy and historical change
- The subjunctive
- Additional observation: explanations vary—asking twice can yield different (but similar) answers per model.
What happened in the tests (main outcomes)
1) Factual information test (Hebrew verbs with four roots in Nefal-related context)
- ChatGPT: produced plausible-sounding but incorrect answers (including category/conjugation confusion). The response looked believable unless the user already knew the grammar.
- Perplexity: also produced incorrect outputs.
- Gemini: hallucinated verbs (appeared to invent non-attested forms rather than merely confusing conjugation class).
- Claude: initially returned some likely incorrect items, but it also corrected itself later and hedged where certainty was lacking.
- Overall: ChatGPT and Perplexity failed; Gemini failed; Claude performed better due to self-correction/hedging, though still not perfect.
2) Vocabulary & sample sentences
- ChatGPT: passed; responses were “fine” and usable, but not standout.
- Gemini: similar “pass” level.
- Perplexity: strong when it provides sources/links for examples; better when prompts combine domains (e.g., color + clothing + specific tenses).
- Claude: especially strong at combining domains and producing useful structured outputs.
3) Transliteration
- ChatGPT: acceptable/typical.
- Gemini: problematic transliteration for Hebrew (misaligned with pronunciation; sometimes used conventions resembling reconstructed biblical Hebrew).
- Claude & Perplexity: preferred/expected transliteration approach (as described by the creator).
- All models: occasional odd artifacts (e.g., Cyrillic letters, inconsistent transliteration of certain characters).
4) Volume (quantity of items)
- Gemini: lowest volume (often ~25 entries), even paid version.
- ChatGPT: around ~35 average.
- Perplexity: around ~43.
- Claude: highest volume (often 70+ and sometimes 200+), with some repetition that required deduping (up to ~5%).
5) Formatting
- Gemini: best for spreadsheet workflow—clear tables that connect to Google Sheets.
- ChatGPT: often messy outputs (tables with many emojis).
- Claude: provided CSV but in a way that might require manual saving/copying.
- Perplexity: produced well-formatted CSV.
6) Images
- Claude: no useful image generation in this test (refuses; creator describes it as avoiding a “losing game”).
- ChatGPT / Perplexity / Gemini: produced labeled images, but:
- labeled Hebrew text could be gibberish / fake
- images became “uncanny” or contextually odd
- Claude judged best here specifically by refusing (i.e., avoiding misleading content).
7) Usefulness of vocabulary
- ChatGPT: useful but mostly obvious/basic items.
- Gemini: similar “basic and fine.”
- Claude: exceptionally useful expansions beyond textbook content; good phrases and contextually relevant examples.
- Perplexity: best in providing more interesting real-world phrasing (creator gives examples like situational phrases in bathroom-related vocabulary).
8) Grammar explanations
- All four generally handled difficult grammar explanations “well enough.”
- ChatGPT: decent technical explanations and helpful memorization strategy suggestions.
- Gemini: strong technical explanation; offered clear visual tables and memorization aids.
- Perplexity: concise semantic-focused answers and included links to where to learn more.
- Claude: technical explanation was good but not especially standout in the creator’s particular comparison.
Final ranking / conclusions
- Overall winner: Claude
- Runner-up: Gemini / Perplexity are positioned as strong depending on priorities, with:
- Gemini strong at integration with Google products and grammar tables, plus image generation if labels aren’t important.
- Perplexity strong at source citation and more interesting vocabulary.
- Overall weakness: ChatGPT (weaker across the evaluation overall, including accuracy concerns).
How the creator says to use LLMs (high-level, non-spoiler framing)
Use an LLM inside a broader structured learning system, not as the sole tutor.
The speaker repeatedly signals that the “how to implement” details are covered in:
- a next video
- and a planned course on self-directed language study and LLM usage.
Speakers / sources featured
- Speaker: Dr. Taylor Jones (linguist; PhD from the University of Pennsylvania)
- Companies/models compared: ChatGPT, Claude, Gemini, Perplexity
- Referenced tools/resources: Duolingo, Lingod, Anki, Google Sheets, WordHippo, Reverso Context, “tatttoo corpus of sentences” (as mentioned), Google (image generation), languagejones.com (course/guide site), Patreon (support page).