Video summary

The Ultimate AI Showdown: ChatGPT vs Claude vs Gemini

Main summary

Key takeaways

News and Commentary

Overview

A research team compares major large language models—ChatGPT, Claude, and Gemini—to determine whether they can be trusted to provide accurate academic references, particularly for research use where citations must:

  1. Exist (the referenced work is real)
  2. Actually support the cited claim (the source contains the information being attributed to it)

What the team tested

They evaluated two types of “hallucinations”:

  1. First-order hallucinations: the model provides a citation/reference that does not exist.
  2. Second-order hallucinations: even when a reference exists, the model may cite it without the claim being supported by the source (i.e., the citation does not contain the information it claims to support).

Stress-test prompt style

The team used “stress test” prompts requiring responses that include:

  • Real quotations
  • Citations in APA bibliography format

They also emphasize that paying for models does not guarantee accuracy.

Results

1) References that exist (first-order)

Overall performance on providing real/valid references:

  • ChatGPT: over 60% correct
  • Claude: about 56%
  • Gemini: only about 20% correct

Best-performing variants (real references)

  • ChatGPT 5 (Thinking) + Web search + Deep research performed best.
  • Claude Sonnet 4 + Research reportedly achieved 100% success for real existing references.
  • Some Claude options were very poor (example: Opus 4.1), where references did not exist.

Gemini variants were notably worst

  • Some Gemini configurations reportedly had 0% real references, including in paid tiers.

2) Citations that match the claim (second-order)

Across all models, the likelihood that the cited paper actually supported the claim was much lower:

  • ChatGPT: just under 50%
  • Claude: just over 40%
  • Gemini: reported 0% (it did not provide citations where the paper supported the claim)

Top variants for second-order success

  • ChatGPT 5 (Thinking) + Deep research / Web search were again the best.
  • ChatGPT 5 agent reportedly had 0% success in this test.

Key takeaways and advice

  • Don’t assume “paying” improves trustworthiness. The test suggests accuracy can still fail badly (especially for Gemini in this evaluation).
  • A common failure mode: models may cite papers that mention a point only because the paper cited something else—a “double extraction” chain rather than primary support.
  • The models behave like “plausibility machines”: outputs can sound convincing even when references are wrong or unsupported.
  • Safe workflow: generate content, then manually trace every claim to the PDF and page.

Alternatives for finding academic references

The presenter recommends specialized academic tools instead of relying on general-purpose LLMs for reference accuracy:

  1. Elicit (checks papers in the background)
  2. Semantic Scholar / SISApace (paper search + literature review built on real references)
  3. Consensus (field-level yes/no confidence using research evidence)

Presenters / contributors

  • The presenter/researcher is not named in the subtitles, referred to only as “my team” and “I.”

Original video