#llama-cpp
POSTS
- Free speed for local LLMs, if you pick the right model
A speedup with no extra hardware cost and no downside is the rarest thing in local inference. Multi Token Prediction gave heise between 25…
Too many projects, too many ideas, too few hours — one learning a day anyway
A speedup with no extra hardware cost and no downside is the rarest thing in local inference. Multi Token Prediction gave heise between 25…