Video summary
The Ultimate AI Showdown: ChatGPT vs Claude vs Gemini
Main summary
Key takeaways
Overview
A research team compares major large language models—ChatGPT, Claude, and Gemini—to determine whether they can be trusted to provide accurate academic references, particularly for research use where citations must:
- Exist (the referenced work is real)
- Actually support the cited claim (the source contains the information being attributed to it)
What the team tested
They evaluated two types of “hallucinations”:
- First-order hallucinations: the model provides a citation/reference that does not exist.
- Second-order hallucinations: even when a reference exists, the model may cite it without the claim being supported by the source (i.e., the citation does not contain the information it claims to support).
Stress-test prompt style
The team used “stress test” prompts requiring responses that include:
- Real quotations
- Citations in APA bibliography format
They also emphasize that paying for models does not guarantee accuracy.
Results
1) References that exist (first-order)
Overall performance on providing real/valid references:
- ChatGPT: over 60% correct
- Claude: about 56%
- Gemini: only about 20% correct
Best-performing variants (real references)
- ChatGPT 5 (Thinking) + Web search + Deep research performed best.
- Claude Sonnet 4 + Research reportedly achieved 100% success for real existing references.
- Some Claude options were very poor (example: Opus 4.1), where references did not exist.
Gemini variants were notably worst
- Some Gemini configurations reportedly had 0% real references, including in paid tiers.
2) Citations that match the claim (second-order)
Across all models, the likelihood that the cited paper actually supported the claim was much lower:
- ChatGPT: just under 50%
- Claude: just over 40%
- Gemini: reported 0% (it did not provide citations where the paper supported the claim)
Top variants for second-order success
- ChatGPT 5 (Thinking) + Deep research / Web search were again the best.
- ChatGPT 5 agent reportedly had 0% success in this test.
Key takeaways and advice
- Don’t assume “paying” improves trustworthiness. The test suggests accuracy can still fail badly (especially for Gemini in this evaluation).
- A common failure mode: models may cite papers that mention a point only because the paper cited something else—a “double extraction” chain rather than primary support.
- The models behave like “plausibility machines”: outputs can sound convincing even when references are wrong or unsupported.
- Safe workflow: generate content, then manually trace every claim to the PDF and page.
Alternatives for finding academic references
The presenter recommends specialized academic tools instead of relying on general-purpose LLMs for reference accuracy:
- Elicit (checks papers in the background)
- Semantic Scholar / SISApace (paper search + literature review built on real references)
- Consensus (field-level yes/no confidence using research evidence)
Presenters / contributors
- The presenter/researcher is not named in the subtitles, referred to only as “my team” and “I.”