Oct 20, 2025/26 mins readNanoChat·Technical Deep Dives·Part 4/13KV Caching Deep-Dive: Memory-Efficient Transformer Inference
Oct 17, 2025/20 mins readNanoChat·Technical Deep Dives·Part 3/13Distributed Muon: Custom Gradient Synchronization for Memory-Efficient Training
Oct 15, 2025/21 mins readNanoChat·Technical Deep Dives·Part 2/13The Muon Optimizer Explained: Why Orthogonal Gradients Work
Oct 04, 2025/17 mins readTiny Language Models·Deployment & Production·Part 9/10Edge Device Deployment: Running Tiny LLMs on Raspberry Pi, Mobile, and IoT
Sep 27, 2025/15 mins readTiny Language Models·Training & Optimization·Part 7/10Quantization-Aware Training: INT8/INT4 Models That Maintain Quality
Sep 19, 2025/26 mins readTiny Language Models·Foundations & Architecture·Part 4/10Efficient Attention Mechanisms for Tiny Language Models