Neuro-symbolic fraud detection with differentiable rules
Weighted BCE starves when fraud is 0.2% of your data. I walk through adding a differentiable rule loss in PyTorch, what it does to ROC-AUC, and why both models need thresholds tuned the same way.