6 LLM architecture lessons the tutorials leave out
Six things I ran into building a GPT-2 clone from scratch in PyTorch: the scaling flaw RsLoRA fixes, why RoPE replaced additive positional embeddings, why weight tying stopped paying off, and what KV-caching costs you in VRAM.