Agentic AI gives you a dot, not an ROC curve
A health-tech client had 99.8% accuracy on a disease that shows up in 0.2% of patients. Their agent said no to everyone. Here is how I pulled a continuous risk score out of a binary agent so the AUC comparison against their XGBoost model meant something.