First training run on a new churn dataset: 99.8% accuracy on the held-out test set. What's the most likely explanation?