{"id":1774,"date":"2026-07-19T08:40:00","date_gmt":"2026-07-19T08:40:00","guid":{"rendered":"https:\/\/nassimstudio.com\/blog\/?p=1774"},"modified":"2026-07-22T15:26:07","modified_gmt":"2026-07-22T15:26:07","slug":"nine-models-tested-three-rejected-32-benchmarks","status":"publish","type":"post","link":"https:\/\/nassimstudio.com\/blog\/nine-models-tested-three-rejected-32-benchmarks\/","title":{"rendered":"9 Models Tested, 3 Rejected: What I Learned From 32 Benchmarks"},"content":{"rendered":"<p>Not every model that claims to be &#8220;good for coding&#8221; actually is. I tested 9 different LLMs on my RTX 3060 Ti, ran 32 benchmark configurations, and rejected 3 of them outright. The failures taught me more about what matters in local AI than the successes.<\/p>\n<p>Here&#8217;s the honest breakdown &#8211; what worked, what didn&#8217;t, and why.<\/p>\n<h2>The Full Roster<\/h2>\n<table>\n<tr>\n<td>Model<\/td>\n<td>Size<\/td>\n<td>Type<\/td>\n<td>Quantization<\/td>\n<td>Verdict<\/td>\n<\/tr>\n<tr>\n<td>Gemma 4 E4B<\/td>\n<td>4B active<\/td>\n<td>MoE<\/td>\n<td>Q4_K_M<\/td>\n<td>Kept &#8211; fast worker<\/td>\n<\/tr>\n<tr>\n<td>Qwen3-4B Thinking<\/td>\n<td>4B<\/td>\n<td>Dense<\/td>\n<td>Q4_1<\/td>\n<td>Kept &#8211; simple planner<\/td>\n<\/tr>\n<tr>\n<td>Qwen3.5 9B<\/td>\n<td>9B<\/td>\n<td>Dense<\/td>\n<td>Q4_K_M<\/td>\n<td>Maybe &#8211; too heavy for speed<\/td>\n<\/tr>\n<tr>\n<td>Gemma 4 26B-A4B<\/td>\n<td>4B active<\/td>\n<td>MoE<\/td>\n<td>Q4_K_M \/ UD-IQ4_NL<\/td>\n<td>Kept &#8211; reasoning model<\/td>\n<\/tr>\n<tr>\n<td>GLM-4.6V-Flash<\/td>\n<td>&#8211;<\/td>\n<td>Dense<\/td>\n<td>Q4_K_M + mmproj<\/td>\n<td><strong>Rejected<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3-Coder 30B-A3B<\/td>\n<td>3B active<\/td>\n<td>MoE<\/td>\n<td>IQ4_XS<\/td>\n<td>Maybe &#8211; not worth the complexity<\/td>\n<\/tr>\n<tr>\n<td>Qwen2.5-Coder 7B<\/td>\n<td>7B<\/td>\n<td>Dense<\/td>\n<td>Q4_K_M<\/td>\n<td>Kept &#8211; daily coder<\/td>\n<\/tr>\n<tr>\n<td>Qwen3.6 35B-A3B MTP<\/td>\n<td>3B active<\/td>\n<td>MoE + MTP<\/td>\n<td>UD-IQ4_NL<\/td>\n<td>Kept &#8211; strong planner<\/td>\n<\/tr>\n<tr>\n<td>Ornith 35B-A3B<\/td>\n<td>3B active<\/td>\n<td>MoE + MTP<\/td>\n<td>UD-IQ4_NL<\/td>\n<td>Kept &#8211; senior planner<\/td>\n<\/tr>\n<\/table>\n<p>Out of 9 models, 6 earned a permanent spot, 2 are &#8220;maybe&#8221; (useful in specific scenarios), and 1 was rejected completely.<\/p>\n<h2>The 3 Rejections<\/h2>\n<h3>GLM-4.6V-Flash: Dead on Arrival<\/h3>\n<p>I tried GLM-4.6V-Flash three times. Three rejections.<\/p>\n<p>The model is Zhipu&#8217;s multimodal offering &#8211; it has vision capabilities through the mmproj (multi-modal projector) component. On paper, it sounds interesting. In practice, it was unusable.<\/p>\n<p><strong>Run 1<\/strong> (Q4_K_M + mmproj, 32k context): 9.75 t\/s generation. Output was garbled Chinese text despite English prompts. Tokenizer warnings in the server log.<\/p>\n<p><strong>Run 2<\/strong>: 12.31 t\/s. Slightly better but still producing mixed Chinese\/English output. The model clearly wasn&#8217;t trained to handle the prompt format I was using.<\/p>\n<p><strong>Run 3<\/strong>: 18.91 t\/s. Speed improved but output quality was still broken. Repeated phrases, incoherent code suggestions, and occasional character soup.<\/p>\n<p>The lesson: model speed means nothing if the output is garbage. GLM-4.6V-Flash is optimized for Chinese language tasks. Running it for English coding was the wrong model for the job. I should have checked the training data language distribution before downloading.<\/p>\n<h3>Qwen3-Coder 30B: Too Much Complexity for Too Little Gain<\/h3>\n<p>Qwen3-Coder 30B-A3B was my attempt at a large dedicated coding model. 30 billion parameters, 3 billion active per token, with IQ4_XS quantization to fit in VRAM.<\/p>\n<p>The model worked. The speed was acceptable (14-19 t\/s depending on MoE layer count). The code quality was good. But it never justified its complexity over simpler alternatives.<\/p>\n<p><strong>The MoE tuning nightmare<\/strong>: Getting acceptable speed required tuning the <code style=\"background:#1e293b;color:#e2e8f0;padding:2px 6px;border-radius:4px;font-size:0.9em\">-ncmoe<\/code> parameter through 8 different values (20, 25, 30, 36, 37, 40, 41, 37 at 128k). At ncmoe 20-25, speed was 10-14 t\/s &#8211; too slow. At ncmoe 37, it hit 18.57 t\/s &#8211; acceptable but not impressive.<\/p>\n<p><strong>The context scaling problem<\/strong>: At 32k context, Qwen3-Coder ran at 18.57 t\/s. At 128k context, it dropped to 16.71 t\/s. The speed degradation with context size was worse than Ornith, which maintained 30-36 t\/s at 128k.<\/p>\n<p><strong>The comparison problem<\/strong>: At 37 MoE layers, Qwen3-Coder hit 18.57 t\/s. Ornith at ncmoe 34 hits 30-36 t\/s. Ornith is faster, has more parameters, supports MTP, and requires less tuning. There&#8217;s no reason to use Qwen3-Coder when Ornith exists.<\/p>\n<p>The lesson: a model that works isn&#8217;t the same as a model that&#8217;s worth the effort. Qwen3-Coder 30B is a perfectly functional model, but the maintenance overhead of tuning it for 8GB VRAM doesn&#8217;t justify its existence when better options are available.<\/p>\n<h3>Qwen3.5 9B: Technically Good, Practically Slow<\/h3>\n<p>Qwen3.5 9B is a solid general-purpose model with vision capabilities. It ran correctly, produced good output, and handled both text and image inputs. I kept it as a &#8220;maybe&#8221; because the speed never justified daily use.<\/p>\n<p>At 125k context with Q4_K_M quantization, the model generated at 22.58 t\/s. Not slow in absolute terms, but slow compared to the alternatives. Gemma E4B does the same quick tasks at 83 t\/s. Qwen2.5-Coder handles coding tasks at 37 t\/s. The 9B model sits in an awkward middle ground &#8211; too slow for quick tasks, not deep enough for hard problems.<\/p>\n<p>The vision capability is interesting, but I don&#8217;t need local vision enough to justify the speed penalty. If I need image analysis, I use a cloud API.<\/p>\n<p>The lesson: a model that&#8217;s &#8220;good at everything&#8221; is often &#8220;not the best at anything.&#8221; Specialization wins on limited hardware.<\/p>\n<h2>What the 6 Keepers Have in Common<\/h2>\n<p>The models I kept share three characteristics:<\/p>\n<p><strong>1. Clear role assignment<\/strong>: Each model does something specific well. Gemma is fast. Qwen Coder is code-focused. Ornith is deep. None of them tries to be everything.<\/p>\n<p><strong>2. Acceptable speed for its role<\/strong>: A &#8220;fast worker&#8221; needs to be genuinely fast (83 t\/s). A &#8220;daily coder&#8221; needs to be interactive (37 t\/s). A &#8220;senior planner&#8221; needs to be fast enough to not interrupt flow (30+ t\/s). Models below 15 t\/s get demoted to &#8220;maybe.&#8221;<\/p>\n<p><strong>3. Minimal tuning overhead<\/strong>: If a model requires 8 rounds of ncmoe tuning to be usable, it&#8217;s not worth the effort. The models that earned a permanent spot worked with reasonable defaults.<\/p>\n<h2>The Surprises<\/h2>\n<h3>MTP Changes Everything for Large Models<\/h3>\n<p>Before testing MTP-enabled models, I thought speculative decoding was a minor optimization &#8211; maybe 10-20% speed improvement. The reality: MTP nearly doubles generation speed on models that support it.<\/p>\n<p>The Qwen3.6 35B without MTP runs at ~18 t\/s. With MTP n=1, it hits 27.19 t\/s. That&#8217;s a 50% improvement from a single flag change. For Ornith with optimized MTP settings, the improvement is even larger &#8211; from an estimated 18-19 t\/s base to 30-36 t\/s with MTP.<\/p>\n<p>If your model supports MTP and you&#8217;re not using it, you&#8217;re running at half speed.<\/p>\n<h3>The 4B MoE Model Outperforms the 7B Dense Model on Speed<\/h3>\n<p>Gemma 4 E4B at 83 t\/s versus Qwen2.5-Coder 7B at 37 t\/s. The smaller MoE model is nearly twice as fast. This makes sense &#8211; MoE activates fewer parameters per token &#8211; but the quality gap is smaller than you&#8217;d expect. For simple coding tasks, Gemma&#8217;s output is 80-90% as good as Qwen&#8217;s, at nearly twice the speed.<\/p>\n<h3>TurboQuant+ KV Compression Is Not Optional<\/h3>\n<p>The first configs I tested used symmetric q8_0\/q8_0 KV cache. Maximum context was 32k on 8GB VRAM. After switching to asymmetric q8_0\/turbo2, context jumped to 128k. The KV cache is the bottleneck, not the model weights. This optimization alone is worth more than any model swap.<\/p>\n<h2>The Final Scorecard<\/h2>\n<table>\n<tr>\n<td>Model<\/td>\n<td>Speed<\/td>\n<td>Quality<\/td>\n<td>Tuning Effort<\/td>\n<td>Role<\/td>\n<td>Verdict<\/td>\n<\/tr>\n<tr>\n<td>Gemma 4 E4B<\/td>\n<td>83 t\/s<\/td>\n<td>Good<\/td>\n<td>Low<\/td>\n<td>Fast worker<\/td>\n<td>Daily driver<\/td>\n<\/tr>\n<tr>\n<td>Qwen3-4B<\/td>\n<td>86 t\/s<\/td>\n<td>Basic<\/td>\n<td>Low<\/td>\n<td>Quick planner<\/td>\n<td>Useful<\/td>\n<\/tr>\n<tr>\n<td>Qwen3.5 9B<\/td>\n<td>22 t\/s<\/td>\n<td>Good<\/td>\n<td>Low<\/td>\n<td>Vision\/general<\/td>\n<td>Maybe<\/td>\n<\/tr>\n<tr>\n<td>Gemma 4 26B<\/td>\n<td>18 t\/s<\/td>\n<td>Good<\/td>\n<td>Medium<\/td>\n<td>Reasoning<\/td>\n<td>Useful<\/td>\n<\/tr>\n<tr>\n<td>GLM-4.6V<\/td>\n<td>9-18 t\/s<\/td>\n<td>Broken<\/td>\n<td>N\/A<\/td>\n<td>Vision<\/td>\n<td><strong>Rejected<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3-Coder 30B<\/td>\n<td>18 t\/s<\/td>\n<td>Good<\/td>\n<td>High<\/td>\n<td>Coding<\/td>\n<td>Maybe<\/td>\n<\/tr>\n<tr>\n<td>Qwen2.5-Coder 7B<\/td>\n<td>37 t\/s<\/td>\n<td>Excellent<\/td>\n<td>Low<\/td>\n<td>Code specialist<\/td>\n<td>Useful<\/td>\n<\/tr>\n<tr>\n<td>Qwen3.6 35B MTP<\/td>\n<td>27 t\/s<\/td>\n<td>Excellent<\/td>\n<td>Medium<\/td>\n<td>Planning<\/td>\n<td>Keep<\/td>\n<\/tr>\n<tr>\n<td>Ornith 35B<\/td>\n<td>30-36 t\/s<\/td>\n<td>Excellent<\/td>\n<td>Low<\/td>\n<td>Daily coder + senior planner<\/td>\n<td>Daily driver<\/td>\n<\/tr>\n<\/table>\n<p>The local AI space moves fast. New models drop every month. But the methodology stays the same: test on your hardware, measure what matters, reject what doesn&#8217;t work, and build a stack around roles instead of chasing the single &#8220;best&#8221; model.<\/p>\n<p>32 test runs taught me that the answer isn&#8217;t one model. It&#8217;s three models, each doing what it does best.<\/p>\n<p>For the complete breakdown of how I use these three models in daily work, see <a href=\"https:\/\/nassimstudio.com\/blog\/three-model-local-ai-stack-fast-coder-planner\/\" target=\"_blank\" rel=\"noopener\">My 3-Model Local AI Stack: Fast Worker + Coder + Senior Planner<\/a>.<\/p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Not every model that claims to be &#8220;good for coding&#8221; actually is. I tested 9 different LLMs on my RTX 3060 Ti, ran 32 benchmark configurations, and rejected 3 of them outright. The failures taught me more about what matters in local AI than the successes. Here&#8217;s the honest breakdown &#8211; what worked, what didn&#8217;t, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":1783,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"rank_math_title":"","rank_math_description":"My 3-model local AI stack: fast worker for quick tasks, coder for daily development, senior planner for complex architecture decisions. All running locally.","rank_math_focus_keyword":"3 model local ai stack fast worker","rank_math_canonical_url":"https:\/\/nassimstudio.com\/blog\/my-3-model-local-ai-stack-fast-worker\/","_jetpack_feature_clip_id":0,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[42],"tags":[34,32,8],"class_list":["post-1774","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-llm","tag-productivity","tag-tools","tag-wordpress"],"blocksy_meta":[],"jetpack_featured_media_url":"https:\/\/nassimstudio.com\/blog\/wp-content\/uploads\/2026\/07\/8_thumb.jpg","jetpack_sharing_enabled":true,"_links":{"self":[{"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/posts\/1774","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/comments?post=1774"}],"version-history":[{"count":15,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/posts\/1774\/revisions"}],"predecessor-version":[{"id":1919,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/posts\/1774\/revisions\/1919"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/media\/1783"}],"wp:attachment":[{"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/media?parent=1774"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/categories?post=1774"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/nassimstudio.com\/blog\/wp-json\/wp\/v2\/tags?post=1774"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}