notes.ludex-gg.com

About these notes

This site publishes measurements taken while running local language models, n8n workflows, and agent systems on ordinary hardware. Every post is written after the work, from notes taken during it.

How the numbers are produced

Everything here is measured on a single machine — a six-core desktop CPU and one 16GB consumer GPU — under the workload described in the post. No numbers are estimated, extrapolated from a spec sheet, or copied from a vendor's documentation.

The method behind most posts is the same, and it is deliberately boring:

Exact configuration — flags, quantisation, context size, model — is included in the post whenever it affects the result, so a reader can tell whether a finding transfers to their setup or does not.

What gets published

The failed hypotheses, alongside the working answer. Several posts here lead with an explanation that turned out to be wrong, because a plausible rule that survives unchallenged is the one that costs you time later.

The costs. Where a post recommends something, it says what that choice gives up — speed, memory headroom, quality on some class of task. A recommendation without a stated tradeoff is advertising.

Corrections, in place. When a later measurement contradicts an earlier post, the post is updated rather than quietly left standing.

What does not get published

No client names, no account details, no private data. Where a post describes a system that shipped a bug, the system is described generically unless naming it serves the reader.

Nothing here is sponsored, and no vendor has reviewed any post before publication. If a link on this site ever earns a commission, the page carrying it will say so, in the page itself, not in a footer.

Corrections

If a measurement here does not reproduce on your hardware, that is worth knowing and the post should say so. The configuration in each post is complete enough to try.