Why can a model perform extremely well offline and badly in production?
Offline results estimate behavior under a historical dataset and evaluation protocol. Production adds a different population, delayed labels, feedback loops, system failures, and business costs.
Verify leakage and split quality, compare feature and prediction distributions by cohort, reproduce online preprocessing, and measure the decision outcome—not only the model score.