Zero-shot classification instead of regex for messy data
Regex and K-means both fall apart on short free-text data where the same idea gets phrased three different ways. Zero-shot classification with a local model like Gemma 2 handles the paraphrasing, and this is the Ollama pipeline I ran over several thousand security annotations.