Visual-Language-Action models and the case for unification
VLA models put vision, language and motion in one latent space instead of three separate pipelines. Notes on why that unification works, and how action tokenization, diffusion heads and flow matching differ once the model has to output real movement.