Factorize queries, keys, and values as contextual tensor products — compressing the KV cache by up to ~10×.
$\;Q_t=\sum_{r=1}^{R}\,a^{Q}_{t,r}\otimes b^{Q}_{t,r}\,,\qquad$ likewise for $K_t,\,V_t$ — rank $R \ll d$.
→Longer context at fixed memory; unifies MHA / MQA / GQA as special cases.
→Compatible with any positional encoding (RoPE, ALiBi, additive RPE, GRAPE); the T6 backbone trains better at equal budget.
"Tensor Product Attention Is All You Need."
Pillar I · Tensor Product Attention
The KV cache is the long-context wall
→Queries, keys, and values factor as contextual tensor products of rank R ≪ d.
→Unifies MHA / MQA / GQA as special cases; compatible with any positional encoding, e.g., RoPE, ALiBi, additive RPE, GRAPE.
NeurIPS 2025 Spotlight · 460 GitHub stars
Pillar I · TPA · FlashTPA Decoding
FlashTPA: decode on the factors
→Per-token attention FLOPs scale as $\Theta(M(R_Q R_K D + H R_Q R_K + H R_V E))$: with small ranks TPA saves prefill and decode compute, not just KV-cache memory.
→Benchmarked against FlashMHA, FlashGQA, FlashMQA, and FlashMLA: faster decoding at long sequence lengths (up to 512K tokens), with Triton decode and prefill kernels in the repo.
Algorithms 2 & 3 in the paper · arXiv:2501.06425 · CUDA kernel in progress
Pillar I · FlashSampling · arXiv:2603.15854
FlashSampling: sample without the logits
→Exact, not approximate: tilewise Gumbel-max keeps one maximizer per row and per vocabulary tile; a small reduction over tiles finishes the categorical sample.
→Fused into the LM-head matmul: avoids logits materialization in HBM. Kernels reach 2.23× over compiled PyTorch multinomial; vLLM TPOT falls up to 10.2% on Qwen3-1.7B (B200).
Each mechanism controls a different part of the computation.
Pillar I · GRAPE · ICLR 2026
Positions as group representations
Positions act via one-parameter subgroups, giving the General Relative Law: $G(t-s)=G(s)^{-1}G(t)$, so scores depend only on the offset $t-s$.
One frame, exact relative position laws, streaming cacheability · see also OpenAI's GPT-5, xAI's Grok 4, Thinking Machines' Inkling, and Jane Street's blog
Pillar I · DeepLoop · the framework
Loop the blocks, scale the depth
→A looped Transformer stores K physical blocks once and revisits them for R rounds: effective depth N = KR with no new parameters.
→Because every visit reuses the same weights, residual updates align across rounds; that alignment is exactly what the α, β rule accounts for.
Pillar I · DeepLoop · arXiv:2607.13491
Loop-aware depth scaling that shows up in loss
$\mathbf{x}_{i+1}=\mathrm{Norm}(\alpha\,\mathbf{x}_i+f(\mathbf{x}_i)),\qquad \alpha=(2N)^{1/2},\quad \beta=(8N)^{-1/2},\qquad p=1/2$ is the aligned-visit threshold.
Gap widens with loop count; recovers DeepNorm's 1/4 when visits decorrelate; also lifts HRM voted accuracy on ARC-AGI.
Pillar I · Recurrent Looped Transformer
Carry computation across tokens
Encoder: known tokens build causal global KV memory in parallel.
Decoder: the previous final state enters the next token’s update; each layer also keeps its own SWA cache.
$\text{Recurrent path after }t\text{ tokens}=tL_D,\qquad\text{blocks per token}=L_E+L_D$
Parity, 256 bits: 100 ± 0% for 5+3 and 7+1; Transformer 8: 50.07 ± 1.63%.S₅ swaps, 256 operations: 97.30 ± 2.76% for 4+4; Transformer 8: 0.85 ± 0.30%.
RLT · October 5, 2026 ↗ · Mean ± sample SD, n = 3; best in-distribution checkpoint. Training lengths: parity ≤40, swaps ≤32. Equal layer counts; parameter counts and compute differ.
Pillar I · RLT · execution tradeoff
Feedback frequency controls parallelism
Decoder schedule at 4+4
Feedback
Parity / 64
S₅ swaps / 64
RLT-1
Every token
100.00%
100.00%
RLT-2
Every 4 tokens
98.99%
19.60%
RLT-0
None
50.23%
Near chance
01Known tokens in a chunk run together. Four-token chunks give 2.27× faster CPU training steps on the mod-5 tasks at 4+4.
02State tracking depends on the feedback interval. Chunking preserves parity much better than permutation tracking. Generated tokens still arrive sequentially.
RLT · October 5, 2026 ↗ · Accuracy: three-seed means. Timing: seed 42, four CPU threads, FP32. Replay rebuilds current-policy state and caches; the recorded sampler probabilities remain fixed.
Pillar I · fixed-state models
A state that stays fixed, and keeps learning
→HLA: higher-order token interactions at linear time; the state does more per byte.
→Falcon: fast weights as episodic memory, so learning continues after pretraining ends.
Pillar I · Higher-order Linear Attention · Zhang, Qin, Gu
HLA: higher-order attention, linear time
→Higher-order interactions from compact prefix sufficient statistics: closed-form streaming identities and a strictly causal masked variant with two extra summaries.
→Chunk-parallel training via associative scans reproduces the serial recurrence exactly; extends to third order and higher.
→Second-order tensor attention $Q(K^{\top}K)Q^{\top}$ touches the context only through second moments, so two streaming summaries are sufficient statistics.
→An associative semidirect-product operator turns the recurrence into a parallel scan; unnormalized and normalized variants share the same state.
Strict causality via a masked variant with two extra summaries · extends to third order and higher
Pillar I · Falcon · ByteDance Seed Technical Report; arXiv:2608.27763
Falcon: fast weights as continual learning
→One lens: attention, recurrent fast-weight memories, and selective SSMs are write rules over a bounded memory.
→Compressing a growing context into a bounded state makes the write rule an online continual-learning rule, connecting architecture design to learning theory.
The continual-learning perspective behind the fixed-state family
Pillar I · Falcon · six write rules
The Falcon family
$S_t=(1-\eta_t\lambda_t)\,S_{t-1}+\eta_t\,x_t r_t^{\top},\qquad (x_t,y_t)=(\phi(k_{t-1}),\,v_t)$ under a read-after-write convention.
→Read-after-write alignment: the prefix-aligned pair $(\phi(k_{t-1}), v_t)$ trains the fast memory on what was actually available at prediction time.
→Consistent semantics: $\beta_t$ plasticity gain, $\lambda_t$ ridge shrinkage, $\eta_t$ induced step size; smoothness-matched or energy-normalized.
→The write is literally one gradient-descent step on an instantaneous ridge objective; the residual $r_t=y_t-S_{t-1}^{\top}x_t$ is measured against the previous state's prediction, unlike standard DeltaNet.
→The loss is $L_t$-smooth with $L_t=\|x_t\|_2^2+\lambda_t$, so $\eta_t$ is smoothness-matched; with $\lambda_t=\varepsilon=0$ it reduces exactly to classical NLMS.
Pillar I · Falcon · experiments
Aligned writes extrapolate
→Language modeling holds: 124M–130M models on FineWeb-Edu at a matched ~50B-token budget stay competitive with Transformer, Mamba-2, DeltaNet, and Gated DeltaNet baselines.
→Length extrapolation improves: the shifted, normalized writes win where storage and carry propagation dominate, 87.2 vs 65.8 mean accuracy out of distribution.
A controlled diagnostic that isolates the memory-writing behavior of the recurrent state
→A rank-1 Householder-style update on the identity shortcut: erasing along k and writing v are coupled, with the gate as a synchronous dynamic step size.
→The network controls the spectrum of its layer transition operator, modeling non-monotonic dynamics while keeping gated-residual training stability.
Derive the objective, then scale agentic RL to long horizons.
Pillar II · a unifying view
REINFORCE is all you need
$\nabla_\theta J=\mathbb{E}_{x\sim\mu}\big[\textstyle\sum_t A_t\,\nabla_\theta\log\pi_\theta(x_t\mid x_{<t})\big]$: pretraining is $\mu=$ corpus, $A_t\equiv 1$; RL is $\mu=\pi_\theta$, $A_t$ from reward.
→Pretraining is off-policy REINFORCE with A ≡ 1: the corpus is the behavior policy, and every next token is a positively rewarded action.
→Scaling REINFORCE across the spectrum, off-policy (corpus, replayed rollouts, IS-corrected) to on-policy (verifier-rewarded rollouts), unifies the training stack.
One objective family from pretraining to agentic RL
Pillar II · FlashREINFORCE
Single-rollout asynchronous REINFORCE
Signed feedbackOne fresh batchOne rollout per prompt. Center rewards across independent prompts, then take one update and discard the batch.No learned critic
Behavior correctionRatio + trajectory gateUse recorded sampler probabilities for token importance ratios; screen complete trajectories by a divergence statistic.No ratio clipping
Length controlSample-mean optimizationAverage tokens within each trajectory, then average trajectories, so long failures do not dominate by length.One outer weight per trajectory
01Regress the optimality condition. Profile out the intercept instead of learning a prompt-dependent normalizer.
02One complete rollout per prompt. Terminal returns support critic-free updates, including stochastic tool outputs.
03No multiplicative importance weights. Sampler probabilities enter through log-ratios.
KLPO · October 5, 2026 ↗ · Yifan Zhang · On KL-Regularized Policy Optimization S is the improvement signal; β is the KL coefficient; Z normalizes the conditional policy.
01M ≥ 1 auxiliary tokens per prefix. Draw from q independently of the complete rollout; no extra completions.
02Center scores under the sampler. Independent MC-KL recovers the Full-KL token gradient in expectation.
03Fresh draws preserve the stated unbiasedness. Reusing fixed historical-sampler records defines an empirical surrogate.
KLPO · October 5, 2026 ↗ · Formula: one-response loss gradient; probabilities condition on the stored prefix. Released: theory, CPU-verified losses and Molt integration. Paper-scale GPU results are not reported.
Alignment · ICML 2025
General Preference Model
Beyond Bradley–Terry: a skew-symmetric operator captures cyclic, intransitive human preference.
MoltAn agentic-first, PyTorch-native RL framework: Ray, vLLM, and NVIDIA AutoModel. About 9.2K lines of RL code that scale to 1T-class MoE.NVIDIA Technical Report; arXiv:2607.21653
RL Kernel MismatchSame checkpoint, different policy: rollout and learner kernels define different effective policies. Four remedies, from importance ratios to exact rejection sampling.Tech Report · 08/2026
Scaling agentic RL takes both the infrastructure and the science.
Pillar II · Molt · NVIDIA Technical Report; arXiv:2607.21653
Molt: three boxes, one async loop
→About 9.2K lines of RL code; MoE-native to 1T-class, think DeepSeek-V3 at EP 256.
→reinforce · rloo · grpo · gae · on-policy distillation; router replay (R3) and IS correction for async rollout.
Pillar II · RL Kernel Mismatch · 08/2026
Same checkpoint, different policy
→Chunkwise, scan, and recurrent kernels are algebraically equal but different finite-precision operators; DeltaNet's corrective write feeds the error into future updates.
→Four remedies: actor-emitted importance ratios, one canonical execution rule, higher-precision state, and exact modified rejection sampling over a fast proposal.
Pillar III
AutoResearch
Agents that automate research itself.
Reasoning
Reasoning as structure, not length
Structured inferenceCumulative Reasoning & Diagram of Thought — reasoning as a DAG, not a linear chain.Proposer–verifier loops that build and check intermediate results.
ControlMeta Prompting — structured prompting as a control layer for LLMs.Composable, reusable reasoning scaffolds for AI systems.
Better structure beats longer chains.
Pillar III · structured reasoning
A chain forgets; a DAG accumulates
→A proposer suggests steps; a verifier checks them; only verified results join the context.
→Cumulative Reasoning (TMLR, 704 citations) and Diagram of Thought formalize the DAG view.
Agentic systems
Agents that act on the world
FlagshipMathCode — a terminal coding agent that formalizes natural-language math into Lean 4 and proves it.
→Lanser-CLI — RL from compiler & language-server feedback (cf. Claude Code v2.0.74).
→Web World Models — controllable, open-ended worlds for language agents: rules in web code, context from LLMs.
Agents that browse, compute, and prove.
Pillar III · MathCode
From natural language to verified proof
→A terminal coding agent where the compiler is the referee; 732 GitHub stars, the most-starred project.
→CriticLean (ACL 2026) adds critic-guided RL to the formalization step itself.
Pillar III · environments
Rewards with executable ground truth
Executable checks make success reproducible; the specification and evaluator still need to match the intended task.
Pillar III · closing the loop
Model and harness coevolve
→We build our own harness; token-exact traces survive even opaque context compaction.
→This loop is what turns research agents into AutoResearch.
Pillar III · Agora · collective AutoResearch
Git records the research process
01Append-only contributions. Code, hypotheses, results and verifications form a dependency graph that workers can check out and reproduce.
02Evidence guides attention. Reuse and verification propagate scores; analyze() exposes frontier results, contested claims and underexplored branches.
Agora · arXiv:2609.18094 ↗ · NeurIPS 2026 AutoMLR Workshop Oral · NVIDIA Technical Report Workers choose experiments from shared state; there is no central planner assigning tasks.
Pillar III · Agora · sustained research run
Collective search leaves reproducible evidence
141 donors, 32 architecture families. Transfer into a 119.6M-parameter attention–SSM hybrid using donor weights and forward passes.
3.3923 → 1.899 bpb. The development evaluator scores 200 FineWeb-Edu texts; pretraining and fine-tuning are prohibited.