local-llm
-
The Reasoning Knob That Silently Did Nothing
A per-request reasoning budget that measurably changed nothing, the setting that actually worked instead, and why a thinking model can return a completely empty answer.
-
The VRAM Cliff — Finding the Real Context Ceiling for a 27B Model on 16GB
Bisecting the point where llama.cpp falls off a performance cliff on a 16GB card, why it looks like a model problem when it is a platform one, and why idle free VRAM does not predict it.