Raphael Friedmann

The engineering log19 posts

Frontier CUDA on a consumer card: the build log for imp and axo, and how fast inference actually works, down to the bits. Newest first.

Built from scratchThe projects, 10 posts

The engines themselves, imp and axo: what they do, how they get built, and the numbers they put up.

How imp runs Qwen3.8-Flash-Next on a 32 GB card

56 GiB of experts, 32 GB of VRAM: all experts on the host, a GPU-managed cache in VRAM, and decode from 5.8 to 75-80 tokens per second on real text.
Expert27 Sep 2026, 14 min

The oracle you build yourself

Dependabot opened five PRs against this site, three of them major. No test suite, so I built the check: render the site with both versions and diff it.
Ops11 Aug 2026, 9 min

No oracle for a meadow

One sentence to Claude Opus 5, 1,922 lines of WebGL, a honeybee in backlight, no number to check. What stands in for a test when the only judge is your eye?
Advanced10 Aug 2026, 15 min

axo: learning without backprop

A from-scratch spiking neural network in C++/CUDA on one RTX 5090 that learns with local rules, no backprop, and grows into a creature that learns to hunt.
Advanced17 Jun 2026, 7 min

Bumping a dependency you can't read

Updating CUTLASS under imp's 4-bit maths. In numerical CUDA a bad bump doesn't crash, it returns wrong numbers. You don't review it, you make it verifiable.
Advanced17 Jun 2026, 8 min

Decoding GGUF faster than llama.cpp on a 5090

imp decodes dense GGUF 37 to 72% faster than llama.cpp on a 5090. Not because llama.cpp is naive: imp is built for one chip. And where it loses.
Expert16 Jun 2026, 4 min

Three kernels for a chip the ecosystem skipped

FlashAttention-2 and NVFP4 GEMM tuned from scratch for the RTX 5090, and what the profiler taught me when every textbook optimisation was a red herring.
Expert11 Jun 2026, 8 min

How 97,000 lines of CUDA got written by an AI agent

Both of my imp posts end with 'every line was written by Claude Code'. The question that always follows: how does that actually work, and how do you trust it?
Advanced9 Jun 2026, 6 min

Serving 30B models at 300 tok/s on a single RTX 5090

Why no existing inference engine fully exploits consumer Blackwell, what NVFP4 changes, and the numbers from building one that does.
Advanced2 Jun 2026, 6 min

How imp turns a model file into words

A plain-language tour of what happens inside an AI engine: how a model gets loaded, and the steps every message runs through to come back as text.
Beginner28 Apr 2026, 6 min

How inference works, at the metalThe concepts, 9 posts

The transferable ideas underneath, from number formats to attention, tuned for one consumer chip.

Making a model emit valid JSON, every time

Asking nicely gets you valid JSON almost always, and almost is a broken integration. The fix is in the sampler: delete every invalid token before you sample.
Expert16 Sep 2026, 6 min

Continuous batching: how one GPU serves a crowd

A single decode stream wastes most of the GPU. Serve many requests at once, share each pass over the weights, and stop waiting for the whole batch to finish.
Advanced24 Jun 2026, 4 min

Speculative decoding: let a small model do the guessing

Decode is bandwidth-bound, so the GPU's maths units sit mostly idle. Speculative decoding spends that compute verifying guessed tokens: same output, faster.
Expert22 Jun 2026, 5 min

Online softmax, and the register that can't move

FlashAttention on the 5090, in depth: how online softmax avoids the giant score matrix, and why register-resident output forces raw mma.sync over WMMA.
Expert17 Jun 2026, 4 min

What a consumer RTX 5090 is missing next to a datacenter GPU

The 5090 and the B200 are both Blackwell, but the consumer chip lacks whole capabilities. What is gone, what it costs, and why data centre code won't port.
Advanced17 Jun 2026, 6 min

NVFP4 at the bit level

Every other post says '4-bit' and moves on. Here's NVFP4 down to the four bits and two scales, and the two Blackwell instructions (cvt, mma) that make it fly.
Expert15 Jun 2026, 3 min

When the model isn't a transformer: GDN and Mamba2

Not every LLM is built on attention. Gated DeltaNet and Mamba2 swap the score matrix for a recurrence: constant memory, and one stubborn precision question.
Expert13 Jun 2026, 4 min

The KV cache, and what really limits long context

The beginner version calls it short-term memory. In reality the KV cache, not the weights, decides how long your context gets and what runs you out of memory.
Advanced19 May 2026, 4 min

Prefill, decode, and the roofline that explains everything

An inference engine doesn't have one speed but two, and they obey opposite laws. The roofline is the most useful lens for reasoning about LLM performance.
Advanced5 May 2026, 4 min