Aug 12, 2026/10 mins readDistilled Engineering·Part 2/8Dark Knowledge: What a Teacher's Wrong Answers Are Actually Worth
Jul 28, 2026/19 mins readKimi K3·Engineering Guide·Part 2/6Inside Kimi K3: How KDA, AttnRes, and 896 Experts Work
Nov 05, 2025/23 mins readNanoChat·Technical Deep Dives·Part 7/13Loss Landscape & Scaling Laws: Understanding Training Dynamics
Oct 31, 2025/21 mins readNanoChat·Technical Deep Dives·Part 6/13Training Data Pipeline: Streaming Tokenization at Scale
Oct 22, 2025/26 mins readNanoChat·Technical Deep Dives·Part 5/13Modern Transformer Architecture: RoPE, QK Norm, and Design Choices
Oct 20, 2025/26 mins readNanoChat·Technical Deep Dives·Part 4/13KV Caching Deep-Dive: Memory-Efficient Transformer Inference
Oct 17, 2025/20 mins readNanoChat·Technical Deep Dives·Part 3/13Distributed Muon: Custom Gradient Synchronization for Memory-Efficient Training
Oct 15, 2025/21 mins readNanoChat·Technical Deep Dives·Part 2/13The Muon Optimizer Explained: Why Orthogonal Gradients Work
Oct 13, 2025/17 mins readTiny Language Models·Deployment & Production·Part 10/10Tiny LLM Deployment Patterns: Architecture Blueprints from Published Benchmarks
Oct 04, 2025/17 mins readTiny Language Models·Deployment & Production·Part 9/10Edge Device Deployment: Running Tiny LLMs on Raspberry Pi, Mobile, and IoT
Oct 01, 2025/18 mins readTiny Language Models·Training & Optimization·Part 8/10Fine-Tuning Tiny Models: LoRA, QLoRA, and Domain Adaptation Strategies
Sep 27, 2025/15 mins readTiny Language Models·Training & Optimization·Part 7/10Quantization-Aware Training: INT8/INT4 Models That Maintain Quality
Sep 23, 2025/20 mins readTiny Language Models·Training & Optimization·Part 6/10Knowledge Distillation: How to Train a 1.5B Model That Matches Your 7B
Sep 20, 2025/23 mins readTiny Language Models·Foundations & Architecture·Part 5/10Tiny LLM Architecture Comparison: TinyLlama vs Phi-2 vs Gemma vs MobileLLM
Sep 19, 2025/26 mins readTiny Language Models·Foundations & Architecture·Part 4/10Efficient Attention Mechanisms for Tiny Language Models
Sep 15, 2025/26 mins readTiny Language Models·Foundations & Architecture·Part 3/10Model Compression: 14GB to 450MB While Keeping 90% Quality
Sep 10, 2025/12 mins readTiny Language Models·Foundations & Architecture·Part 2/10Mathematical Foundations of Model Compression: Theory Behind Tiny LLMs
Sep 05, 2025/34 mins readTiny Language Models·Foundations & Architecture·Part 1/10Tiny Language Models: How 1.3B Parameters Can Beat 7B on Reasoning