SØNDAG
2026-09-13

Too many projects, too many ideas, too few hours — one learning a day anyway

One token a second, and the SSDs are the bottleneck I didn't expect

The number I keep staring at isn’t 1.00 tok/s. It’s the drive ladder: one SSD gives ~52% of four-drive speed, two ~73%, three ~90%. That’s not bandwidth scaling, that’s the slowest of each layer’s 16 reads setting the pace. Same shape as every marine NMEA bus I’ve debugged — the network runs at the speed of its worst talker, not its total capacity.

So if I were rebuilding my rack for this, I’d stop shopping for one fast drive and start buying four boring identical ones. And I’d read PREFILL.md before anything else: ~6.3 minutes to first token on a 512-token prompt, cause identified (prefill re-reads each layer’s experts 8×), fix planned but not built. Honest unfinished work beats a polished claim. Nothing here is quantized down — that’s the whole point.


The story — ARGODRIVE’s Deltafin fork runs the full 2.8T-parameter Kimi K3 MoE, 1.45 TB of expert weights unpruned, on a single M5 Max MacBook Pro with 128 GB, streaming experts from four SSDs. Measured 2026-09-08: 1.00 tok/s steady decode over 512 tokens, 1.13 over 128, 0.96 on the public 17-token prompt where upstream reported 0.68. Every number is one cold run, logs published. (Source)