#local-llm
- One token a second, and the SSDs are the bottleneck I didn't expect
The number I keep staring at isn't 1.00 tok/s. It's the drive ladder: one SSD gives ~52% of four-drive speed, two ~73%, three ~90%. That's…
- 512GB of Unified Memory Is the Whole Story
The number I keep staring at is 512GB of unified memory at 1.2TB/s. That's the line where running frontier-class open-weight models locally…
- M6 and M5 Ultra: buy for the memory, not the cores
Every M-series launch I do the same math and land the same place. M6 gets a Dual 16-core Neural Engine, 170GB/s of bandwidth, Neural…
- Vomit pipes Claude's output through a second local model
I run Claude Code all day and I've stopped reading half of what scrolls past. A tool that pipes that output through a second local model…
- Free speed for local LLMs, if you pick the right model
A speedup with no extra hardware cost and no downside is the rarest thing in local inference. Multi Token Prediction gave heise between 25…