Video summary

Why Is My AI So Slow? Go Faster with EvoX AI Harness

Main summary

Key takeaways

Product Review

Product reviewed: EvoX AI Harness

A “coding and productivity assistant” desktop-style app (shown running on Mac in the video) that uses cloud AI models plus optional local workflows. It’s positioned as faster than local generation, with tools for:

  • Coding agents
  • Self-evolution / learning from mistakes
  • Swarm-style autonomous task execution

Key features mentioned

Coding + productivity assistant (multiple work modes)

The app includes multiple work modes/sections:

  • Chat
  • Co-work (more integrated tasks)
  • Dedicated coding section

Cloud execution with a free credit system

  • Starts with $15 worth of credits
  • Creator reports 1,500 credits available at signup

Model choice (including “thinking modes”)

Models mentioned include:

  • DeepSeek V4 Flash
  • GPT Sol
  • Luna
  • Sol
  • Terror
  • Grok 4.6 (SpaceX)
  • Kimi K3

Thinking effort modes:

  • extra high
  • high
  • minimal
  • off

The video mainly compares low and extra high, noting that low can reduce quality.

Cloud vs local workflow

The assistant can operate in two ways:

  • Local system mode: generate files on your machine
  • Remote execution mode: generate/operate remotely “into a computer”

It can also:

  • Create and organize project folders
  • Write generated files into those folders

Safety/automation controls for code changes

The app includes controls around file writing and edits, such as:

  • Prompting for approval to write files
  • Options like:
    • auto-accept edits
    • confirm steps
    • run smart accept edits
  • “Similar commands” option to reduce repeated confirmations

Beta features

  • Beta swarm agent (autonomous background development/testing)
  • One-click migration of projects

Skills system

The assistant can load specialized skills depending on the task (e.g., interactive visual demos/games). The creator claims it can:

  • save/reuse memory/skills to improve future outputs

Self-evolution / self-improvement

The video shows a “self-evolution” progress indicator, e.g.:

“94% away from self-evolution”

This implies the system tracks progress toward evolving/improving its future capability and skills.

Reasoning visibility differs by model

The creator states:

  • With closed-weight models, reasoning is hidden
  • With Kimi K3, some reasoning can be seen

What the reviewer tested (and outcomes)

1) Speed comparison: local vs cloud

  • Local Kimi K3 was described as extremely slow:
    • Example: 7,000 tokens at ~6.62k/sec
  • Cloud EvoX using Kimi K3 was “a lot faster” for a similar task

2) “Solar system” HTML generation (Kimi K3 vs GPT Sol)

  • Kimi K3: produced a working, good-looking interactive solar system
  • GPT Sol:
    • Initially produced a blank screen
    • Then it fixed the code after adjusting imports and logic

Conclusion: Kimi K3 looked better visually, while GPT Sol required fixes (but recovered successfully).


3) Interactive game-like task: spaceship + lasers + explosions concept

The creator attempted a more advanced interactive feature.

  • GPT Sol:
    • Created a flyable spaceship and working laser shooting
    • Laser visuals had a direction issue (described as going the “wrong direction” / vertically)
    • Still playable and had no runtime errors
  • Kimi K3:
    • Lasers/planets appeared, but the spaceship wasn’t properly visible/controllable

Conclusion: GPT Sol “won” for gameplay/control responsiveness.


4) Extra-high benchmark: photo-realistic real-time face (cloud vs local)

  • Cloud (GPT Sol, extra high):
    • Finished in about 1 minute
    • Iterated after defects (e.g., missing mouth)
    • Later corrected to include mouth/lips with improved rendering/shaders
  • Local comparison (GLM 5.3 quantized):
    • At low thinking: 77,000 tokens and ~2 hours
    • Also showed a WebGL2 requirement bug message (though rendering may have still happened)

Explicit conclusion: EvoX/cloud was dramatically faster (1 minute vs 2 hours) with visible iteration to fix defects.


Swarm agent test (autonomous coding/testing)

  • Swarm agent enabled with Ultra cool set to auto
  • Prompted it to generate an implementation for a new model: DeepSseek V4.1
  • The swarm reportedly:
    • Checked out existing implementations/skills
    • Compared against prior models’ code (e.g., older DeepSeek V4 and Quen 4)
    • Requested permission to clone MLX source
    • Ran a background plan including:
      • writing/reading files
      • creating tests
      • comparing diffs
      • producing implementation files locally
  • Result: creator reports implementation files were successfully generated
  • UX detail: notifications keep the user informed so they can “relax and watch YouTube.”

Pros (as stated/shown)

  • Much faster than local generation
    • Example cited: face render ~1 minute cloud vs ~2 hours local
  • Strong code-writing workflow
    • Creates files, updates code, and corrects errors (e.g., blank-screen fix, face-mouth fix)
  • Useful autonomy via swarm
    • Background planning, testing, cloning dependencies, generating implementations
  • Self-improvement concept (“self-evolution”) with a visible progression indicator
  • Model flexibility
    • Switch between models and thinking modes; reasoning visibility varies by model
  • Skills system
    • Loads task-appropriate capabilities (e.g., game/visual skills)

Cons / negatives mentioned

  • Closed-weight model reasoning is hidden
  • Quality can drop at low thinking effort
    • Example: GPT Sol at low produced a blank screen first
  • Local execution can be extremely slow
    • Example cited: ~2 hours run
  • Some output artifacts/bugs can appear
    • Laser direction visual artifact (lasers moving vertically)
    • Face render initial missing mouth + occasional shader/render quirks
    • Local GLM showed a WebGL2 requirement message

Comparisons made

EvoX cloud vs local generation speed

Multiple tasks showed drastic time savings—especially the face render (1 min vs 2 hrs).

Model comparisons within EvoX

Kimi K3 vs GPT Sol:

  • Solar system: Kimi K3 looked better; GPT Sol needed fixes
  • Spaceship/lasers: GPT Sol performed better (flyable spaceship)
  • Face render at extra high: GPT Sol produced a strong result quickly

Reasoning visibility

  • Kimi K3: more reasoning visible
  • Closed-weight models: reasoning hidden

Unique points list (distinct claims/features/outcomes mentioned)

  1. EvoX is a latest coding & productivity assistant.
  2. Designed to learn from mistakes and be self-evolving.
  3. Provides free API credits ($15, reported 1,500 credits).
  4. Supports multiple models including DeepSeek V4 Flash, GPT Sol, Kimi K3, Grok 4.6, etc.
  5. Works across channels: communicates via Slack and Telegram (claimed).
  6. Includes beta swarm and one-click project migration.
  7. Desktop app UX includes Chat, Co-work, and Dedicated coding sections.
  8. Supports working with local files and remote computer access.
  9. Creates project folders and writes outputs to them.
  10. Supports model choice including thinking modes (extra high/high/minimal/off).
  11. Local Kimi K3 performance was extremely slow (example: 7,000 tokens, 6.62k/sec).
  12. Kimi K3 cloud run is “much faster.”
  13. GPT Sol may hide reasoning; Kimi K3 can show reasoning.
  14. GPT Sol at low thinking initially produced blank screen, then fixed code.
  15. Kimi K3 produced a working solar system; GPT Sol needed import/logic correction.
  16. GPT Sol can fix issues after code edits (responsive iteration).
  17. Spaceship/lasers test: - GPT Sol produced flyable spaceship and lasers - Kimi K3 attempt lacked visible/controllable spaceship
  18. GPT Sol had a laser direction artifact but was playable with no runtime errors.
  19. “Extra high” face render with GPT Sol completed in about 1 minute.
  20. Local face render using GLM 5.3 quantized took 77,000 tokens and ~2 hours.
  21. Local GLM had a WebGL2 bug message (though rendering possibly happened).
  22. Face render initially missing mouth; user prompt led to correction.
  23. Self-evolution indicator shown: “94% away from self-evolution.”
  24. Self-evolution implies it saves/uses a “gene” and improves next runs.
  25. Swarm agent test involved generating DeepSeek V4.1 implementation.
  26. Swarm autonomously did planning, cloned deps (MLX), checked/used other models (DeepSeek V4, Quen 4), and generated tests/implementation.
  27. Swarm produced implementation files in the user’s folder.
  28. Swarm sends notifications so user doesn’t need to actively watch.
  29. Overall verdict in the video: GPT Sol “wins” several higher-quality comparisons; extra-high cloud face generation is best.

Speaker-specific views (attributed)

  • Single main speaker/reviewer (video narrator):
    • Provides comparisons, benchmarks, and qualitative judgments.
    • Claims about reasoning visibility, speed differences, and self-evolution come from their demonstrations.
  • No other distinct speakers are clearly present in the subtitles.

Original video