Cutting vector search costs by 80% with MRL and int8
Vector infrastructure gets expensive fast: 100 million 1024-dimensional float32 vectors need over 1.2TB of RAM. Pairing Matryoshka truncation with int8 quantization cuts storage by roughly 80%, and this post covers where the savings stop and recall starts to break.