Assessing Models: Validation, Bias and Ethics
Correct evaluation determines whether a model actually works, while ethics determine if it can be applied without harm or discrimination.
Validation protocols include cross-validation, stratification, and time splits. Quality metrics are chosen based on the task and business goals; there's a trade-off between precision and recall.
Calibration of Probabilities: Platt Scaling, Isotonic Regression
The ‘accuracy’ metric can be misleading on imbalanced classes. Incorrect splits lead to inflated quality assessments. There's a lack of a formal process for challenging decisions in high-risk domains.
Evaluation isn't just one number; it’s a set of practices that ensure the model is honest, correct, and legitimate.
High Test Metrics Don’t Guarantee Model Responsibility
1) Choose appropriate validation and stratification protocols. 2) Measure multiple metrics with different trade-offs. 3) Verify fairness for subgroups. 4) Document limitations and risks, implement a challenge process.
In credit scoring, in addition to AUC, the difference in rejection rates between demographic groups is measured, probabilities are calibrated, and decisions are made considering ethical policies and regulatory requirements.
Frequently asked questions
What is time-series validation used for?
Time-series validation involves splitting the data into different periods to assess how well a model predicts future events based on past trends.
How do Platt scaling and isotonic regression calibrate probabilities?
Platt scaling and isotonic regression are techniques used to adjust predicted probabilities, particularly when dealing with imbalanced datasets, to better reflect the true likelihood of events.
How should fairness metrics be evaluated and documented?
Fairness metrics like Demographic Parity and Equalized Odds need to be carefully evaluated, and the results thoroughly documented to identify potential biases in the model's predictions.
Should stress tests and edge cases be included during evaluation?
Yes, including stress tests and examining extreme or unusual scenarios helps to assess the robustness of a model and identify potential vulnerabilities that might not be apparent in standard testing.
▶ Try it live
Everything above runs in your browser — open Hash Function Avalanche Visualizer and change the parameters while it is running. Nothing is installed, nothing is uploaded, the whole model lives in one tab.