Raphael Friedmann
← The log

How imp runs Qwen3.8-Flash-Next on a 32 GB card

56 GiB of experts, 32 GB of VRAM: all experts on the host, a GPU-managed cache in VRAM, and decode from 5.8 to 75-80 tokens per second on real text.

imp is an inference engine in C++23 and CUDA, built for exactly one chip: NVIDIA Blackwell sm_120a, the RTX 5090 with 32 GB of GDDR7. No portability layer, no FP16 fallback in the hot path. This post covers how imp runs a model that does not fit that card on paper: Qwen3.8-Flash-Next, listed in its checkpoint as architecture qwen4_exp, a Mixture-of-Experts model with 512 experts.

Where it stands: on real text, imp decodes this model at 75 to 80 tokens per second with every expert living in host RAM. The starting point was 5.8.

Anatomy of the model

NVIDIA’s NVFP4 checkpoint is 123 GiB on disk. Per the porting notes, the parameters split into 125 billion in the main model, 51 billion in an n-gram table and 4 billion in the MTP head. About 6 billion are active per token.

config.json describes a hybrid:

Property Value
Layers 48
of which Gated DeltaNet (linear_attention) 36
of which full attention (full_attention) 12, every fourth layer
Experts per MoE layer 512, 10 active per token, plus 1 shared expert
Expert intermediate size 640
Hidden size 2560
Attention: Q heads / KV heads / head dim 24 / 2 / 256
Vocabulary 248320
Max context 262144
Residual streams (hc_count) 4

The Gated DeltaNet layers are known territory for imp. Three other traits are not.

Gated residual instead of a plain residual. The residual stream is four times wide, [T, 4*2560], initialised as the embedding repeated four times. Before each block a small network (low-rank dimension 320) mixes the four streams into one input via sigmoid weights; afterwards the block output is injected into each stream with its own weight. The model has no input_layernorm and no final norm: the pre-norm lives inside the gated residual, and at the end the mixer feeds the LM head directly. imp crashed on this twice because rmsnorm() got a null weight. The fix: a null weight now means identity, handled centrally in rmsnorm().

The PLE n-gram table. Layer 1 carries an extra embedding block. A hash over the last tokens yields a set of n-gram IDs: 16 heads, 8 for bigrams and 8 for trigrams. Each ID indexes a row in a huge FP8 table. The server log:

ngram table: layer 1, 128 shards, 320001536 rows x 160, 16 heads x ngram 3,
scale 0.00019932, eos 248044, 47.7 GiB host-mapped (page cache, not resident)

The 16 lookups produce a key and a value, a gate against the current residual stream, and a causal dilated depthwise convolution (kernel 4, dilation 3). This is not decoration: an ablation without PLE took prose perplexity from 14.3 to 24.8. The table carries roughly half the quality.

More than 256 experts. Sounds like one number among many. It exposed a bug, covered below.

Why the experts don’t fit

Only the experts are NVFP4; everything else is BF16. The load log:

NVFP4 experts (56.25 GiB) do not fit the 13.68 GiB budget: all 48 MoE layer(s)
stay host-resident so the expert cache gets the VRAM (a partial upload starves it).

56 GiB of experts against a 32 GB card. The obvious move is to load as many expert layers into VRAM as fit and fetch the rest from the host. imp does the opposite on purpose. A resident expert layer buys a 100 % hit rate on 1/48 of the work and costs 1.17 GiB. The same VRAM as cache gives about 27 slots per layer on all 48 layers. Partial residency loses as long as the cache has to serve the rest. Rule: if not everything fits, everything stays on the host.

Host experts, device cache, zero-copy

The expert path has three levels.

  1. Pinned slabs in host RAM. The experts of one projection sit back to back in a pinned block mapped into the GPU’s address space. One expert is two byte ranges: packed FP4 weights and FP8 micro-scales.
  2. An LRU cache in VRAM. A fixed number of 0.88 MiB slots per MoE layer. In the run shown here: 48 layers × 221 slots × 0.88 MiB = 9323.48 MiB.
  3. A cache managed entirely on the GPU (src/exec/expert_cache_device.{h,cu}). A resolve kernel reads the routed expert selection, looks it up in LRU tables on the GPU, picks victims for the misses and writes the slot indices. A gather kernel copies the missing experts zero-copy straight from mapped host memory over PCIe, measured at 51 GB/s. No round trip to the host, no sync per layer.

That last point is the precondition for everything after it. The older host-LRU path had to copy each layer’s routing decision to the host, do the bookkeeping there, then launch the copies. Head to head it does 30.3 tokens per second; the device cache does 74.5 to 79.6.

One subtlety sits in the scales: the kernels read the per-tensor scale with the same index as the weight. That index is now a slot, not an expert number, so the cache keeps a per-slot mirror of the scales, updated whenever a slot changes owner.

How big the cache may get is a measurement question. Sweep on real text (5 prose prompts, 200 tokens each, median):

Share of free VRAM Slots/layer VRAM free after Decode
45 % 186 4995 MiB 74.1 / 76.5 tok/s
55 % 227 3242 MiB 77.5 / 77.2 tok/s
62 % 256 2170 MiB hangs

At 62 % the server is not slow, it hangs: the GDN state for a sequence can no longer be committed, requests are accepted and never answered. The default for host-resident NVFP4 experts is therefore 55 %. The hang now returns HTTP 503 instead of silence.

The ceiling also shows from the other side: in an earlier run, 60 % instead of 45 % raised the hit rate from 98.8 to 99.2 %, but decode fell from 102.42 to 78.26 tokens per second. The allocation succeeded, the bytes spilled. Under WSL2 a successful cudaMalloc does not prove the memory is actually on the card.

Two obvious optimisations were measured and dropped:

  • Smarter eviction. A simulator replays routing traces and reproduces the measured 62.7 % hit rate at 186 slots. Pinning hot experts loses to LRU at every budget, 59.4 vs 62.0 %. Routing is too flat: layer 0 touches 384 of 512 experts in 300 tokens, and the top 64 carry 41 %.
  • Computing misses on the CPU. On the Ryzen 9800X3D the host reads scattered 0.88 MiB blocks at 58 to 60 GB/s; PCIe does 52. That is 13 % headroom, not 2x.

Decode: from 5.8 to 75-80 tokens per second

The speedup came in steps, each with a named cause. Measured with 96 generated tokens on the same prompt; the output text was byte-identical across all steps:

Step tg96 tok/s What changed
Baseline (mmap, 15 % budget) 5.81 3 cudaMemcpyAsync per miss; a 4-byte scale from the stack syncs the stream
Pinned host experts 8.10 DMA-capable source
Scale and index as kernel parameters 14.35 no sync per miss
N-gram speculation off, 45 % budget 20.6 186 slots/layer, 71 % hits
cudaMemcpyBatchAsync per layer 27.2 31 to 56 µs host time per single copy gone
Device cache 33.2 resolve and gather on the GPU, no D2H per layer
Graph replay 53.0 about 2600 kernel launches per token were the host floor

Graph replay had one obstacle: the PLE block needs a host step per token (compute the hash, gather rows from the mapped table). That step now runs before each replay in prepare_decode_step_host, outside the graph.

With the 55 % cache budget and automatic configuration, decode landed at the 74 to 80 tokens per second from the sweep.

Where the time goes, from a profile of one 16.6 ms decode step: 7.2 ms in expert_cache_gather (PCIe misses), 3.1 ms in GEMV kernels, 1.1 ms in the resolve kernel. The resolve kernel originally had thread 0 scan 256 candidates serially for every miss. An 8-step tree reduction took it from 1147.0 to 888.7 µs per token at an identical hit rate.

The misses cannot hide behind compute. The forward pass is strictly sequential: attention or GDN, then gated residual, then MoE. Routing knows its experts only after attention, and cross-layer prediction fails because the misses are exactly the unpredictable ones.

Prefill: CUTLASS in expert blocks

In prefill many experts are active at once, and a cache helps little. imp copies a layer’s experts into a staging buffer and runs them through a grouped CUTLASS GEMM on the FP4 tensor cores. Three steps made this fast.

Staging for 512 experts. The grouped-GEMM helper kernels were sized for fewer experts and needed changes (multi-block scan, limit 4096). After that: a 1353-token prompt went from 4.11 to 1.96 s.

Copy only the experts that were hit. The first staging path copied all 512 experts with all three projections per layer, 63 GB per prefill. That made time to first token nearly independent of prompt length. Now a gather kernel reads the routing offsets and fetches only experts that received tokens:

Prompt tokens TTFT before TTFT after
67 1267 ms 499 ms
328 1287 ms 811 ms
1022 1328 ms 1038 ms

Routing is skewed: at 67 tokens the copy drops to 39 %, not the 73 % a uniform distribution predicts.

Block staging. A staging buffer for a whole layer takes slots away from the expert cache. The current path stages gate and up in blocks, runs the GEMMs, then stages down in blocks. In the log: MoE layer staging buffer: 225.00 MiB (128 experts x 2 projection slots, 4 expert block(s) per layer). A config bug had disabled this path for a while, so a 13.9k prompt ran as 837k single copies from unpinned memory. Measured after the fix on a 3692-token prompt:

Variant Staging buffer Slots/layer Prefill tok/s Decode tg512
before (legacy path) - 227 202.67 / 196.39 60.09 / 58.23
1 block 1350 MiB 191 1108 / 1084 54.58 / 55.06
4 blocks 225 MiB 221 1143 / 1131 / 1054 59.80 / 55.81 / 57.43

With four blocks the cache stays almost as large as before and prefill is more than five times faster. The v0.45.0 release notes call it “Qwen3.8-Flash-Next prefill 5.6x”.

Bugs on the way

Experts from 256 up were never chosen. After the port imp produced coherent text, but perplexity on a self-written, unseen prose text was 14.25; llama.cpp got 5.54, both via llama-perplexity on the same window. Per-layer dumps showed: router logits match, the selection does not. topk_gating_kernel gave each of its 256 threads exactly one expert, so experts with index 256 and up were never candidates. No earlier model had more than 256 experts. After the fix, which gives each thread several slots: 6.00 vs 5.54 for llama.cpp, and 1.422 vs 1.4238 on a Gutenberg excerpt. A regression test now covers it. The model was not visibly broken before; only the comparison against an independent reference showed the bug.

The QSA indexer is off by default. The twelve attention layers ship a learned block top-k indexer (budget 2048, compression 4). Below 2051 tokens of context its mask is all true and attention is exactly dense. imp implemented it, then set the default to off. The first justification, a rounding difference flipping MoE routing, did not survive review: no test in the suite reached 2051 tokens, so the indexer had never been active. The measurement, one server per variant:

Probe QSA off QSA on
Needle at 11283 tokens found found
1024-token decode at 11283 context 18.14 tok/s 15.88 tok/s (-12.5 %)
Degeneration suite 50/50 47 to 48/50

The structural reason: only 12 of 48 layers have attention, 2 KiB of KV per token each, so 24 KiB per token. At 12k context that is 288 MB, by calculation (not measured) about 0.19 ms of read time against a 63 ms decode step that belongs to the host experts. QSA has almost nothing to save, but adds an index GEMM, selection, gather and a second attention kernel per layer. The selection kernel runs one thread block per query row and gets more expensive as context grows. Two real defects in the QSA path were found and fixed along the way; the default stays off.

Auto-defaults that blocked startup. On default settings the model at first did not start at all: the slot path needs at least 30 slots per layer and has 0. Four defaults blocked each other:

Defect Effect
Expert layers greedily loaded into VRAM 13 of 48 in VRAM, cache with 0 slots
Cache budget flat 15 % 3 to 4 times too small for host-resident NVFP4
Auto max_seq_len ignored the expert cache KV cache ate its VRAM
Pinning tied to one switch no device view, host-LRU path, no graphs

The weight estimate had a bug too: it counted every expert as if it lived in VRAM. That gave 78.5 GB instead of 3.5 GB and capped automatic context at 4096 tokens. After the fix imp picks 131072. Current log:

max_seq_len: auto → 131072 (model=262144, vram_cap=221810, auto_cap=131072,
kv=24576 B/tok, attn_layers=12/48)

Plus a container problem: inside a container /proc/meminfo shows the host’s number. A container with --memory=8g on the 78 GB host read 76 GB free. imp now takes the minimum of MemAvailable and the cgroup headroom, minus page cache.

Tool calls cut off. With n-gram speculation, the verify step adopts a snapshot of the GDN state when zero tokens are accepted. The default scan kernels never wrote that snapshot, so uninitialised memory was adopted. In 3 of 8 fresh server processes, tool calls ended mid-argument. After the fix: 0 of 12.

Batch size 1

Auto-config would pick 32 parallel sequences for this model. imp sets it to 1, and the log says why:

max_batch_size 32 -> 1: this model's PLE block holds one n-gram context, so batched
decode would answer every sequence from the first one's context.

The PLE block holds the last two tokens and nine rows of convolution state for exactly one sequence. Before this fix the code logged “output is wrong” and kept computing anyway. Eight parallel streams reliably crashed the server. Parallel requests are now served one after another, each decode step for one sequence. That rules out continuous batching for this model for now.

What memory looks like in the end

From the server log of a current run:

Item Size
Expert cache in VRAM 9323.5 MiB (221 slots/layer)
Device cache tables 2.1 MiB
Prefill staging buffer 225 MiB
KV cache at start 1750 blocks, 28000 tokens, 656.25 MiB, growing to 131072 tokens
Experts in host RAM 57600 MiB
N-gram table 47.7 GiB mapped, page cache, not resident
GPU total 29128 of 32579 MiB used

The PLE table is mapped with MADV_RANDOM and never read in full. Over 1300 decode tokens, 25 MiB came off the NVMe; in decode it costs practically nothing. RAM is the tight part: 56 GiB of pinned experts plus PLE on a 78 GB host leave 9 to 16 GB free after loading. If RAM is too short to pin, imp skips pinning rather than risk the process. The price is steep: with a 40 GB container limit decode drops to 4.3 tokens per second.

Honest limits

  • One sequence at a time. Until the PLE context is kept per sequence there is no batched decode. A multi-user server has the throughput of a single stream.
  • Hit rate depends on the text. Benchmark mode generates repetitive text and hits 98.8 % in the cache; real text hits 62.7 %. A 120-token request ran at 54 tokens per second, the same build in the benchmark at 102. Benchmark numbers overstate this model’s real speed by a wide margin.
  • PCIe is the floor. Almost half of a decode step is gather over PCIe. Neither better eviction nor CPU compute promises more than a few percent.
  • The cache budget has a hard ceiling. More slots take the GDN state’s room. 55 % is measured, not derived.
  • Not done: the MTP head (1 layer, in the FP8 shard) does not run. Prefix-cache resumes for the PLE context are missing too.
  • QSA gains nothing on this model. Measured up to 23k context; 131k is untested.

The core of the approach is unspectacular: all experts on the host, VRAM for a cache that works without the host, and every step checked against real text instead of the benchmark.