How imp runs Qwen3.8-Flash-Next on a 32 GB card
56 GiB of experts, 32 GB of VRAM: all experts on the host, a GPU-managed cache in VRAM, and decode from 5.8 to 75-80 tokens per second on real text.
imp is an inference engine in C++23 and CUDA, built
for exactly one chip: NVIDIA Blackwell sm_120a, the RTX 5090 with 32 GB of GDDR7.
No portability layer, no FP16 fallback in the hot path. This post covers how imp
runs a model that does not fit that card on paper: Qwen3.8-Flash-Next, listed in
its checkpoint as architecture qwen4_exp, a Mixture-of-Experts
model with 512 experts.
Where it stands: on real text, imp decodes this model at 75 to 80 tokens per second with every expert living in host RAM. The starting point was 5.8.
Anatomy of the model
NVIDIA’s NVFP4 checkpoint is 123 GiB on disk. Per the porting notes, the parameters split into 125 billion in the main model, 51 billion in an n-gram table and 4 billion in the MTP head. About 6 billion are active per token.
config.json describes a hybrid:
| Property | Value |
|---|---|
| Layers | 48 |
of which Gated DeltaNet (linear_attention) |
36 |
of which full attention (full_attention) |
12, every fourth layer |
| Experts per MoE layer | 512, 10 active per token, plus 1 shared expert |
| Expert intermediate size | 640 |
| Hidden size | 2560 |
| Attention: Q heads / KV heads / head dim | 24 / 2 / 256 |
| Vocabulary | 248320 |
| Max context | 262144 |
Residual streams (hc_count) |
4 |
The Gated DeltaNet layers are known territory for imp. Three other traits are not.
Gated residual instead of a plain residual. The residual stream is four times
wide, [T, 4*2560], initialised as the embedding repeated four times. Before each
block a small network (low-rank dimension 320) mixes the four streams into one input
via sigmoid weights; afterwards the block output is injected into each stream with
its own weight. The model has no input_layernorm and no final norm: the pre-norm
lives inside the gated residual, and at the end the mixer feeds the LM head directly.
imp crashed on this twice because rmsnorm() got a null weight. The fix: a null
weight now means identity, handled centrally in rmsnorm().
The PLE n-gram table. Layer 1 carries an extra embedding block. A hash over the last tokens yields a set of n-gram IDs: 16 heads, 8 for bigrams and 8 for trigrams. Each ID indexes a row in a huge FP8 table. The server log:
ngram table: layer 1, 128 shards, 320001536 rows x 160, 16 heads x ngram 3,
scale 0.00019932, eos 248044, 47.7 GiB host-mapped (page cache, not resident)
The 16 lookups produce a key and a value, a gate against the current residual stream, and a causal dilated depthwise convolution (kernel 4, dilation 3). This is not decoration: an ablation without PLE took prose perplexity from 14.3 to 24.8. The table carries roughly half the quality.
More than 256 experts. Sounds like one number among many. It exposed a bug, covered below.
Why the experts don’t fit
Only the experts are NVFP4; everything else is BF16. The load log:
NVFP4 experts (56.25 GiB) do not fit the 13.68 GiB budget: all 48 MoE layer(s)
stay host-resident so the expert cache gets the VRAM (a partial upload starves it).
56 GiB of experts against a 32 GB card. The obvious move is to load as many expert layers into VRAM as fit and fetch the rest from the host. imp does the opposite on purpose. A resident expert layer buys a 100 % hit rate on 1/48 of the work and costs 1.17 GiB. The same VRAM as cache gives about 27 slots per layer on all 48 layers. Partial residency loses as long as the cache has to serve the rest. Rule: if not everything fits, everything stays on the host.
Host experts, device cache, zero-copy
The expert path has three levels.
- Pinned slabs in host RAM. The experts of one projection sit back to back in a pinned block mapped into the GPU’s address space. One expert is two byte ranges: packed FP4 weights and FP8 micro-scales.
- An LRU cache in VRAM. A fixed number of 0.88 MiB slots per MoE layer. In the
run shown here:
48 layers × 221 slots × 0.88 MiB = 9323.48 MiB. - A cache managed entirely on the GPU (
src/exec/expert_cache_device.{h,cu}). A resolve kernel reads the routed expert selection, looks it up in LRU tables on the GPU, picks victims for the misses and writes the slot indices. A gather kernel copies the missing experts zero-copy straight from mapped host memory over PCIe, measured at 51 GB/s. No round trip to the host, no sync per layer.
That last point is the precondition for everything after it. The older host-LRU path had to copy each layer’s routing decision to the host, do the bookkeeping there, then launch the copies. Head to head it does 30.3 tokens per second; the device cache does 74.5 to 79.6.
One subtlety sits in the scales: the kernels read the per-tensor scale with the same index as the weight. That index is now a slot, not an expert number, so the cache keeps a per-slot mirror of the scales, updated whenever a slot changes owner.
How big the cache may get is a measurement question. Sweep on real text (5 prose prompts, 200 tokens each, median):
| Share of free VRAM | Slots/layer | VRAM free after | Decode |
|---|---|---|---|
| 45 % | 186 | 4995 MiB | 74.1 / 76.5 tok/s |
| 55 % | 227 | 3242 MiB | 77.5 / 77.2 tok/s |
| 62 % | 256 | 2170 MiB | hangs |
At 62 % the server is not slow, it hangs: the GDN state for a sequence can no longer be committed, requests are accepted and never answered. The default for host-resident NVFP4 experts is therefore 55 %. The hang now returns HTTP 503 instead of silence.
The ceiling also shows from the other side: in an earlier run, 60 % instead of 45 %
raised the hit rate from 98.8 to 99.2 %, but decode fell from 102.42 to 78.26 tokens
per second. The allocation succeeded, the bytes spilled. Under WSL2 a successful
cudaMalloc does not prove the memory is actually on the card.
Two obvious optimisations were measured and dropped:
- Smarter eviction. A simulator replays routing traces and reproduces the measured 62.7 % hit rate at 186 slots. Pinning hot experts loses to LRU at every budget, 59.4 vs 62.0 %. Routing is too flat: layer 0 touches 384 of 512 experts in 300 tokens, and the top 64 carry 41 %.
- Computing misses on the CPU. On the Ryzen 9800X3D the host reads scattered 0.88 MiB blocks at 58 to 60 GB/s; PCIe does 52. That is 13 % headroom, not 2x.
Decode: from 5.8 to 75-80 tokens per second
The speedup came in steps, each with a named cause. Measured with 96 generated tokens on the same prompt; the output text was byte-identical across all steps:
| Step | tg96 tok/s | What changed |
|---|---|---|
| Baseline (mmap, 15 % budget) | 5.81 | 3 cudaMemcpyAsync per miss; a 4-byte scale from the stack syncs the stream |
| Pinned host experts | 8.10 | DMA-capable source |
| Scale and index as kernel parameters | 14.35 | no sync per miss |
| N-gram speculation off, 45 % budget | 20.6 | 186 slots/layer, 71 % hits |
cudaMemcpyBatchAsync per layer |
27.2 | 31 to 56 µs host time per single copy gone |
| Device cache | 33.2 | resolve and gather on the GPU, no D2H per layer |
| Graph replay | 53.0 | about 2600 kernel launches per token were the host floor |
Graph replay had one obstacle: the PLE block needs a host step per token (compute the
hash, gather rows from the mapped table). That step now runs before each replay in
prepare_decode_step_host, outside the graph.
With the 55 % cache budget and automatic configuration, decode landed at the 74 to 80 tokens per second from the sweep.
Where the time goes, from a profile of one 16.6 ms decode step: 7.2 ms in
expert_cache_gather (PCIe misses), 3.1 ms in GEMV kernels, 1.1 ms in the resolve
kernel. The resolve kernel originally had thread 0 scan 256 candidates serially for
every miss. An 8-step tree reduction took it from 1147.0 to 888.7 µs per token at an
identical hit rate.
The misses cannot hide behind compute. The forward pass is strictly sequential: attention or GDN, then gated residual, then MoE. Routing knows its experts only after attention, and cross-layer prediction fails because the misses are exactly the unpredictable ones.
Prefill: CUTLASS in expert blocks
In prefill many experts are active at once, and a cache helps little. imp copies a layer’s experts into a staging buffer and runs them through a grouped CUTLASS GEMM on the FP4 tensor cores. Three steps made this fast.
Staging for 512 experts. The grouped-GEMM helper kernels were sized for fewer experts and needed changes (multi-block scan, limit 4096). After that: a 1353-token prompt went from 4.11 to 1.96 s.
Copy only the experts that were hit. The first staging path copied all 512 experts with all three projections per layer, 63 GB per prefill. That made time to first token nearly independent of prompt length. Now a gather kernel reads the routing offsets and fetches only experts that received tokens:
| Prompt tokens | TTFT before | TTFT after |
|---|---|---|
| 67 | 1267 ms | 499 ms |
| 328 | 1287 ms | 811 ms |
| 1022 | 1328 ms | 1038 ms |
Routing is skewed: at 67 tokens the copy drops to 39 %, not the 73 % a uniform distribution predicts.
Block staging. A staging buffer for a whole layer takes slots away from the
expert cache. The current path stages gate and up in blocks, runs the GEMMs, then
stages down in blocks. In the log: MoE layer staging buffer: 225.00 MiB (128 experts x 2 projection slots, 4 expert block(s) per layer). A config bug had disabled this
path for a while, so a 13.9k prompt ran as 837k single copies from unpinned memory.
Measured after the fix on a 3692-token prompt:
| Variant | Staging buffer | Slots/layer | Prefill tok/s | Decode tg512 |
|---|---|---|---|---|
| before (legacy path) | - | 227 | 202.67 / 196.39 | 60.09 / 58.23 |
| 1 block | 1350 MiB | 191 | 1108 / 1084 | 54.58 / 55.06 |
| 4 blocks | 225 MiB | 221 | 1143 / 1131 / 1054 | 59.80 / 55.81 / 57.43 |
With four blocks the cache stays almost as large as before and prefill is more than five times faster. The v0.45.0 release notes call it “Qwen3.8-Flash-Next prefill 5.6x”.
Bugs on the way
Experts from 256 up were never chosen. After the port imp produced coherent text,
but perplexity on a self-written, unseen prose text was 14.25; llama.cpp got 5.54,
both via llama-perplexity on the same window. Per-layer dumps showed: router logits
match, the selection does not. topk_gating_kernel gave each of its 256 threads
exactly one expert, so experts with index 256 and up were never candidates. No
earlier model had more than 256 experts. After the fix, which gives each thread
several slots: 6.00 vs 5.54 for llama.cpp, and 1.422 vs 1.4238 on a Gutenberg
excerpt. A regression test now covers it. The model was not visibly broken before;
only the comparison against an independent reference showed the bug.
The QSA indexer is off by default. The twelve attention layers ship a learned block top-k indexer (budget 2048, compression 4). Below 2051 tokens of context its mask is all true and attention is exactly dense. imp implemented it, then set the default to off. The first justification, a rounding difference flipping MoE routing, did not survive review: no test in the suite reached 2051 tokens, so the indexer had never been active. The measurement, one server per variant:
| Probe | QSA off | QSA on |
|---|---|---|
| Needle at 11283 tokens | found | found |
| 1024-token decode at 11283 context | 18.14 tok/s | 15.88 tok/s (-12.5 %) |
| Degeneration suite | 50/50 | 47 to 48/50 |
The structural reason: only 12 of 48 layers have attention, 2 KiB of KV per token each, so 24 KiB per token. At 12k context that is 288 MB, by calculation (not measured) about 0.19 ms of read time against a 63 ms decode step that belongs to the host experts. QSA has almost nothing to save, but adds an index GEMM, selection, gather and a second attention kernel per layer. The selection kernel runs one thread block per query row and gets more expensive as context grows. Two real defects in the QSA path were found and fixed along the way; the default stays off.
Auto-defaults that blocked startup. On default settings the model at first did
not start at all: the slot path needs at least 30 slots per layer and has 0. Four
defaults blocked each other:
| Defect | Effect |
|---|---|
| Expert layers greedily loaded into VRAM | 13 of 48 in VRAM, cache with 0 slots |
| Cache budget flat 15 % | 3 to 4 times too small for host-resident NVFP4 |
Auto max_seq_len ignored the expert cache |
KV cache ate its VRAM |
| Pinning tied to one switch | no device view, host-LRU path, no graphs |
The weight estimate had a bug too: it counted every expert as if it lived in VRAM. That gave 78.5 GB instead of 3.5 GB and capped automatic context at 4096 tokens. After the fix imp picks 131072. Current log:
max_seq_len: auto → 131072 (model=262144, vram_cap=221810, auto_cap=131072,
kv=24576 B/tok, attn_layers=12/48)
Plus a container problem: inside a container /proc/meminfo shows the host’s number.
A container with --memory=8g on the 78 GB host read 76 GB free. imp now takes the
minimum of MemAvailable and the cgroup headroom, minus page cache.
Tool calls cut off. With n-gram speculation, the verify step adopts a snapshot of the GDN state when zero tokens are accepted. The default scan kernels never wrote that snapshot, so uninitialised memory was adopted. In 3 of 8 fresh server processes, tool calls ended mid-argument. After the fix: 0 of 12.
Batch size 1
Auto-config would pick 32 parallel sequences for this model. imp sets it to 1, and the log says why:
max_batch_size 32 -> 1: this model's PLE block holds one n-gram context, so batched
decode would answer every sequence from the first one's context.
The PLE block holds the last two tokens and nine rows of convolution state for exactly one sequence. Before this fix the code logged “output is wrong” and kept computing anyway. Eight parallel streams reliably crashed the server. Parallel requests are now served one after another, each decode step for one sequence. That rules out continuous batching for this model for now.
What memory looks like in the end
From the server log of a current run:
| Item | Size |
|---|---|
| Expert cache in VRAM | 9323.5 MiB (221 slots/layer) |
| Device cache tables | 2.1 MiB |
| Prefill staging buffer | 225 MiB |
| KV cache at start | 1750 blocks, 28000 tokens, 656.25 MiB, growing to 131072 tokens |
| Experts in host RAM | 57600 MiB |
| N-gram table | 47.7 GiB mapped, page cache, not resident |
| GPU total | 29128 of 32579 MiB used |
The PLE table is mapped with MADV_RANDOM and never read in full. Over 1300 decode
tokens, 25 MiB came off the NVMe; in decode it costs practically nothing. RAM is the
tight part: 56 GiB of pinned experts plus PLE on a 78 GB host leave 9 to 16 GB free
after loading. If RAM is too short to pin, imp skips pinning rather than risk the
process. The price is steep: with a 40 GB container limit decode drops to 4.3 tokens
per second.
Honest limits
- One sequence at a time. Until the PLE context is kept per sequence there is no batched decode. A multi-user server has the throughput of a single stream.
- Hit rate depends on the text. Benchmark mode generates repetitive text and hits 98.8 % in the cache; real text hits 62.7 %. A 120-token request ran at 54 tokens per second, the same build in the benchmark at 102. Benchmark numbers overstate this model’s real speed by a wide margin.
- PCIe is the floor. Almost half of a decode step is gather over PCIe. Neither better eviction nor CPU compute promises more than a few percent.
- The cache budget has a hard ceiling. More slots take the GDN state’s room. 55 % is measured, not derived.
- Not done: the MTP head (1 layer, in the FP8 shard) does not run. Prefix-cache resumes for the PLE context are missing too.
- QSA gains nothing on this model. Measured up to 23k context; 131k is untested.
The core of the approach is unspectacular: all experts on the host, VRAM for a cache that works without the host, and every step checked against real text instead of the benchmark.