news

May 15, 2026 In long contexts, RoPE provably loses both its locality bias and consistent token relevance, while increasing its base trades off distinguishing positions against distinguishing tokens—see our new preprint RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provably.
May 13, 2026 Continuously consolidating useful experiences into LLM-written memory can eventually make agents worse than using no memory; preserving raw episodes and gating consolidation is more reliable—see our new preprint Useful Memories Become Faulty When Continuously Updated by LLMs.
Feb 17, 2026 Stronger SFT can hurt downstream RL; our PEAR reweighting makes SFT checkpoints better starters for RL and improves post-RL performance—new preprint Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning.
Feb 17, 2026 Surprisingly, plain SGD matches (or beats) AdamW for RL in LLMs while updating <0.02% of parameters—see our new preprint Do We Need Adam?.