About
(1)
Research
(2)
Join
(3)
/blog
8.03.2026
(Some of) The Models, They Just Don't Want to Learn
We are launching One Layer Deeper, a competition for co-designing architectures, objectives, and optimizers to learn deeper serial computation and extrapolate beyond training depth.
6.08.2026
Parallax: Parameterized Local Linear Attention
Parallax upgrades softmax attention to a local linear fit via a single learned covariance correction.
6.02.2026
Wall Attention: Length Generalization With Diagonal Gates
We generalize diagonal gating from linear RNNs to softmax attention, introducing Wall Attention - a data-dependent positional encoding that achieves strong length generalization and beats RoPE across the board.
5.05.2026
Aurora: A Leverage-Aware Optimizer for Rectangular Matrices
Muon kills neurons in tall MLP matrices. Aurora fixes it with a joint orthogonality–leverage constraint, resulting in an incredibly powerful optimizer.
4.28.2026
Nitrobrew: Fast, Lossless Distillation for Free
8.10.2025
A Telescope, Pointed at the Stars
7.25.2025
MoMoE: Memory-optimized Mixture of Experts
06.25.2025
Sparsity is Cool
3.17.2025
Activault: Scalable, Efficient, and Fast Model Activation Storage
12.15.2024
Sieve: SAEs Beat Baselines on a Real-World Task (A Code Generation Case Study)
11.12.2024
The Rate Distortion Dance of Sparse Autoencoders