SØNDAG
2026-09-13

Too many projects, too many ideas, too few hours — one learning a day anyway

The 17 GB version is the one I'd actually run

I have been talking myself out of a bigger GPU for months, and this settles it. Q4_K_M of Qwen3.8 27B is 17 GB, matches full BF16 on Terminal-Bench 2.1, and leaves roughly 64k of context on a 24 GB card. The 55 GB BF16 build was never going to live in my rack. Fit the best model you can alongside the context you actually need, and stop treating quantization as a compromise.

Two things I am taking to my own setup. One, do not chase 1-bit. At UD-IQ1_S it scores around random on GPQA Diamond, and more reasoning effort makes it worse — it burns the budget and returns nothing. Two, effort level moved scores more than quantization did. Tune that first. It is free.


The story — Quesma benchmarked Unsloth’s GGUF quantizations of Qwen3.8 27B on GPQA Diamond, IFBench, and Terminal-Bench 2.1, spending about $3,000 on Modal GPUs. The 17 GB Q4_K_M matched BF16; 2-bit UD-Q2_K_XL held on IFBench and fell noticeably but not uselessly on Terminal-Bench, writing about a quarter more tokens. The 1-bit quants collapsed to random-guess level. (Source)