Build an LLM evaluation system that catches hallucinations
A scoring layer that measures attribution and specificity, so you can catch confident hallucinations instead of trusting responses that just sound right. Includes the Python logic and a regression suite for CI.