skip to content
exer//build · idea

a very small model on a very small box

llama.cpp · rust · arm64

Local inference on hardware that has no business running it. The interesting number is not tokens per second — it is what you stop sending to an API.

[ placeholder content ] — shipped with the site so the layout can be reviewed. Not a published project.

The premise

Everyone benchmarks local models on tokens per second. That is the least interesting number. The interesting number is how much of your traffic stops leaving the building — because that is the number that changes a data protection conversation.

Where it stands

Idea stage. The quantised model runs. The box gets hot. Neither of those is a finding yet.

This is placeholder content shipped with the site so the layout can be reviewed with realistic text. Replace it with a real project.