SWE-2 - Tokenstead
- MoE premier 2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime.
- The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.
- Serving stack: MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers.
Unverified
- MoE premier 2.8T total params, 104B active per token (MoE) - the Kimi K3 base with Cognition’s post-training on top, and the first time Cognition has scaled RL into the multi-trillion-parameter regime.
- The base had already been RL-heavy for agentic coding; Cognition’s pass added another 5 to 6 points on most benchmarks.
- Serving stack: MoE inference on NVFP4 and FP8 kernels with quantization-aware training; FP8 carries K, Q, V, and score computations in the MLA layers.
Sources: Tokenstead