Pillar · AI Computing
Running local AI properly
The hardware pages tell you what to buy. This section is what happens next: the serving stack, the tool connections, and agents doing real work. Everything here runs on my own bench.
From the bench
Serving
My vLLM setup: a benched model stable on a 48GB 4090
The working Docker compose, how every flag was benched and re-benched, the measured tok/s per model, and the houtini-lm bridge that lets Claude delegate to it all.
Read →Coming next
Building MCP servers (lessons from shipping thirteen), Open WebUI with a vLLM backend, and the dual-GPU tensor-parallel write-up once the second card lands.