Oct 22, 2025/26 mins readNanoChat·Technical Deep Dives·Part 5/13Modern Transformer Architecture: RoPE, QK Norm, and Design Choices
Oct 20, 2025/26 mins readNanoChat·Technical Deep Dives·Part 4/13KV Caching Deep-Dive: Memory-Efficient Transformer Inference
Oct 15, 2025/21 mins readNanoChat·Technical Deep Dives·Part 2/13The Muon Optimizer Explained: Why Orthogonal Gradients Work
Sep 10, 2025/12 mins readTiny Language Models·Foundations & Architecture·Part 2/10Mathematical Foundations of Model Compression: Theory Behind Tiny LLMs