how llms remember context

Attention vs Mamba

Two ways for a language model to see its own past — and why one of them stays fast at a million tokens.

Attention layer

memory grows every token
KV cache
0 entries
Each new token looks back at every stored token to decide what matters…
…and adds itself to a cache that never stops growing. More context → more memory → slower decoding.

Mamba layer

fixed-size state, forever
h
running state
same size, every step
Each token updates one small state — a running summary — then is discarded.
Token 1,000,000 costs the same as token 5. Nothing grows. The trade-off: exact old details can fade.

Memory as context grows

all-attention
Gemma 4
Soofi S
cache context → 4K 64K 256K 1M
Soofi S keeps a KV cache in only 6 of 52 layers — the rest run on Mamba's fixed state. Gemma 4 slows the growth with local windows, but every layer is still attention.

Two recipes, layer by layer

Soofi S 30B-A3B · hybrid
23× Mamba · 23× MoE · only 6× attention
Gemma 4 26B-A4B · all-attention
5× local window : 1× global, MoE everywhere
M Mamba  ·  E MoE experts  ·  A/G full attention  ·  l local attention
Attention keeps the whole transcript — perfect recall, growing cost.
Mamba keeps a running summary — flat cost, softer memory.
Soofi S bets on Mamba. Gemma 4 bets on smarter attention.
sources: Soofi S report · Gemma 4 technical report — schematic, not to scale
a ~75s animated explainer · press play