Video summary

I Tested Claude/Codex/Cursor $20/mo Plan Limits: 10 Coding Prompts

Main summary

Key takeaways

Technology

Summary

The video compares how quickly three $20-per-month AI coding subscriptions use their included limits—not just what equivalent API usage might cost. The host ran the same 10 uncached coding prompts across two PHP/Laravel projects, with five runs per project. The tasks focused on hardening existing code: a sync API and a CSV contact importer.

The experiment tracked subscription-limit usage, completion time, API-equivalent cost where reported, and test results. The host cautions that the comparison is limited: it used 10 separate requests rather than long-running sessions, and usage can vary depending on the project, prompts, context, and caching.

Results

  • Claude Opus, medium effort: Used about 41% of the five-hour allowance and 4% of the weekly allowance. The 10 runs took 27 minutes in total, with an estimated API cost of $6.96. All tests passed across both projects.
  • Claude Sonnet, high effort: Used about 27% of the five-hour allowance and 2% of the weekly allowance in a subsequent test. It was described as cheaper than Opus, but some tests failed, including two on one project.
  • Codex, medium effort: The captions identify the model as “GPT 6.1 Soul.” The test used about 25% of the five-hour allowance and 4% of the weekly allowance. All tests passed, but the runs were slower than Opus, with some taking nine or ten minutes.
  • Cursor with Grok 4.6, high effort: Cursor uses a monthly allowance rather than the same five-hour and weekly limits. After more than an hour of runs, reported Grok usage had increased by 2%. Results were inconsistent, with a failed test and other errors.

Conclusions

In this small test, Opus and Codex delivered similar test quality and used roughly the same share of their weekly allowances. Opus used the five-hour allowance faster, while Codex was slower. Sonnet and Grok produced less consistent test results.

The host emphasizes that these findings apply only to this test and may not predict usage on other projects or in longer sessions. The video also cites a separate experiment by Pavel, which reportedly found different relative subscription value—another example of how results can depend on test design and workload.

Review or Guide Covered

This is a hands-on subscription-limit comparison, rather than a coding tutorial. It demonstrates a repeatable evaluation approach: run identical prompts from fresh CLI sessions, then compare limit usage, elapsed time, and test outcomes.

Main Speaker and Sources

  • Main speaker: The AI Coding Daily host, reporting results from their own leaderboard script and subscription tests.
  • Additional source: Pavel, whose separate subscription-usage research is briefly discussed.

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video