The engineering log19 posts
Frontier CUDA on a consumer card: the build log for imp and axo, and how fast inference actually works, down to the bits. Newest first.
Built from scratchThe projects, 10 posts
The engines themselves, imp and axo: what they do, how they get built, and the numbers they put up.
How imp runs Qwen3.8-Flash-Next on a 32 GB card
56 GiB of experts, 32 GB of VRAM: all experts on the host, a GPU-managed cache in VRAM, and decode from 5.8 to 75-80 tokens per second on real text.The oracle you build yourself
Dependabot opened five PRs against this site, three of them major. No test suite, so I built the check: render the site with both versions and diff it.No oracle for a meadow
One sentence to Claude Opus 5, 1,922 lines of WebGL, a honeybee in backlight, no number to check. What stands in for a test when the only judge is your eye?axo: learning without backprop
A from-scratch spiking neural network in C++/CUDA on one RTX 5090 that learns with local rules, no backprop, and grows into a creature that learns to hunt.Bumping a dependency you can't read
Updating CUTLASS under imp's 4-bit maths. In numerical CUDA a bad bump doesn't crash, it returns wrong numbers. You don't review it, you make it verifiable.Decoding GGUF faster than llama.cpp on a 5090
imp decodes dense GGUF 37 to 72% faster than llama.cpp on a 5090. Not because llama.cpp is naive: imp is built for one chip. And where it loses.Three kernels for a chip the ecosystem skipped
FlashAttention-2 and NVFP4 GEMM tuned from scratch for the RTX 5090, and what the profiler taught me when every textbook optimisation was a red herring.How 97,000 lines of CUDA got written by an AI agent
Both of my imp posts end with 'every line was written by Claude Code'. The question that always follows: how does that actually work, and how do you trust it?Serving 30B models at 300 tok/s on a single RTX 5090
Why no existing inference engine fully exploits consumer Blackwell, what NVFP4 changes, and the numbers from building one that does.How imp turns a model file into words
A plain-language tour of what happens inside an AI engine: how a model gets loaded, and the steps every message runs through to come back as text.How inference works, at the metalThe concepts, 9 posts
The transferable ideas underneath, from number formats to attention, tuned for one consumer chip.