AI evaluation
Choose inputs, comparisons and metrics before trusting a score.
Choose a useful starting point.
Two scores are comparable only when the task, input split, normalization, aggregation and failure handling are stated. Begin by freezing those choices and a simple baseline. The evaluation-harness brief asks whether two tools agree on the same saved predictions; it does not claim either tool has been run here.
For agent evaluations, preserve the environment, action trace, budget and unsuccessful attempts alongside the scoring rule. The linked protocols help you plan this record. They do not make a benchmark suitable for every agent or establish that a qualified runtime supports its dependencies.
Questions to work through.
Put uncertainty around a small agent-success benchmark
How misleading can a nominal 95% success-rate interval be when only twenty independent tasks are observed?
Open the brief RESEARCH BRIEF · 6 SOURCESExplain a score difference before blaming the model
Do evaluation harnesses agree on the same frozen predictions once normalization, aggregation and failure handling are made explicit?
Open the brief RESEARCH BRIEF · 6 SOURCESMake an agent benchmark attempt auditable
Can another operator reconstruct one bounded agent-evaluation attempt from its task, environment, actions, failures and scoring rule?
Open the briefExplore the sources.
All 191 resources →Iris
Flower measurements for a compact, interpretable multiclass baseline and leakage audit.
3.1. Cross-validation: evaluating estimator performance
A guide to model-evaluation splits, including grouped and time-series data, nested model selection and permutation-based assessment.
1.13. Feature selection
A guide to variance filters, univariate tests, recursive elimination, model-based selection and selection inside pipelines.
1.16. Probability calibration
A probability-calibration guide covering reliability curves and sigmoid, isotonic and temperature-scaling approaches.
2.7. Novelty and Outlier Detection
A guide distinguishing outlier detection from novelty detection and comparing assumptions of several anomaly-detection methods.
2D elastodynamic metamaterials
Pixelated metamaterial unit-cell designs and band-gap locations and widths for investigating simulation surrogate reliability.
3.4. Metrics and scoring: quantifying the quality of predictions
A broad scoring reference that distinguishes classification, regression, ranking and clustering metrics and their scorer interfaces.
3W dataset
Oil-production process signals for studying event detection across operating scenarios.
5.1. Partial Dependence and Individual Conditional Expectation plots
A guide to partial-dependence and individual-conditional-expectation curves, including definitions, computation and correlated-input caveats.
5.2. Permutation feature importance
A model-inspection guide explaining permutation importance and its limitations when predictors are strongly correlated.
[Re] Badder Seeds: Reproducing the Evaluation of Lexical Methods for Bias Measurement
A replication examining lexical methods for measuring bias and sensitivity to seed choices.
[Re] BiRT: Bio-inspired Replay in Vision Transformers for Continual Learning
A replication of replay-based continual learning with vision transformers, relevant to evaluating retention across sequential tasks.