Video summary

Nex N2.5 Mini tested - 16GB Local LLM setup

Main summary

Key takeaways

Technology

Summary

The video tests Nex N2.5 Mini in GGUF Q4_K_M quantization, focusing on performance and practical coding abilities. The presenter evaluates the model on a 16 GB GPU-oriented setup running Ubuntu and llama.cpp, with 32 GB of system RAM. The model was loaded with a 128K context; a separate memory test filled the context to 256K.

Performance and evaluation results

  • Speed: Uncached prefill reached about 233 tokens/sec on a short sequence and 265 tokens/sec on a 32K sequence. Decode speed was about 55–56 tokens/sec. Cached prefill was substantially faster.
  • Long-context memory: The model retrieved inserted information correctly in all 15 runs, across five context depths from 0% to 100%. The presenter considered this one of its strongest results.
  • Reasoning: 36/48 (75%). It handled logic reasonably well but missed about half of the difficult and expert-level questions.
  • OpenAI HumanEval: 87% overall. The model answered 154 of 164 tasks, with a 93% success rate on answered questions. The presenter cautioned that the benchmark is old and may appear in the model’s training data.

Coding and tool-based tasks

  • Kanban frontend test — failed: The app had poorly arranged columns and generated errors. Despite multiple prompts, the model repeatedly fell into a coding loop and could not fix the layout.
  • Sand physics — mixed, mostly successful: It created working sand, water, walls, and acid interactions, but the water behavior had a visible issue. The task consumed many tokens.
  • Dungeon crawler — good result: It generated a functional-looking dungeon with raycasting, lighting, doors, and a fog-of-war map. The presenter noted that it completed this without prior training or context compression.
  • Blender lantern — failed: The model produced an incomplete object and could not continue generating or repairing the code.
  • Godot 3D game — poor result: The environment was difficult to navigate, controls behaved strangely, and the game was incomplete and prone to crashing. The session used nearly all available context after repeated compression.

Overall assessment

The model’s strongest areas were generation speed and long-context retrieval. Its reasoning and HumanEval results were respectable, but its performance on larger, iterative coding tasks was inconsistent. The Kanban, Blender, and Godot tasks were all unsuccessful.

The presenter’s overall recommendation was not to spend time using this quantized version, while noting that a higher quantization such as Q6 or Q8 might perform differently.

Main speaker and sources

  • Main speaker: The presenter from Luke’s Dev Lab; the auto-generated subtitles render the channel name inconsistently.
  • Sources and tests mentioned: The Nex N2.5 Mini model card, llama.cpp, the presenter’s Luke’s Dev Lab evaluation site, OpenAI HumanEval, and the presenter’s Kanban, sand-physics, dungeon-crawler, Blender, and Godot tests.

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video