Thoughts on Clean Code, Automation, and Web Architecture.
Practical deep dives into building high-performance applications, refining development workflows, and maintaining scalable codebases.
Featured
Freelancing
⏱ 6 min read
I’ve Lost 60% of My Leads in 7 Months — Here’s What I’m Doing Wrong
AI & Local LLMs
⏱ 8 min read
Quantization for 8GB VRAM: Q4_K_M vs IQ4_XS vs UD-IQ4_NL
AI & Local LLMs
⏱ 7 min read
9 Models Tested, 3 Rejected: What I Learned From 32 Benchmarks
AI & Local LLMs
⏱ 7 min read
My 3-Model Local AI Stack: Fast Worker + Coder + Senior Planner
AI & Local LLMs
⏱ 8 min read
Gemma 4 vs Qwen vs Ornith: Which Local LLM Codes Best on 8GB
AI & Local LLMs
⏱ 7 min read
turbo2 vs turbo3 vs turbo4: Which KV Cache Quantization Fits Your VRAM
AI & Local LLMs
⏱ 6 min read
Speculative Decoding Explained: How MTP Turns 18 t/s Into 36 t/s
AI & Local LLMs
⏱ 9 min read
Asymmetric KV Cache: Why Compressing V More Than K Unlocks 128k
AI & Local LLMs
⏱ 8 min read
Ornith 35B on 8GB VRAM: 128k Context at 30 t/s
Web Design
⏱ 9 min read
I Fit a 35B Model Into 8GB VRAM. Heres the Full Setup.
Web Design
⏱ 10 min read
My Full Local AI Benchmark: 32 Runs, 9 Models, 8GB VRAM
Development
⏱ 5 min read





