Statistics & uncertainty
Understand small samples and the assumptions behind an apparent improvement.
Choose a useful starting point.
A success rate from twenty tasks contains less information than the same rate from a large study. The binomial brief gives a fixed independent-trial model for comparing interval rules. Correlated tasks, repeated prompts and unequal task difficulty require a different model; an interval formula alone does not settle those choices.
Trying many variants also changes how an apparent improvement should be interpreted. Declare the family of comparisons and distinguish the chance of any false positive from the false-discovery rate. The resources below connect those questions to source methods and unexecuted plans.
Questions to work through.
Put uncertainty around a small agent-success benchmark
How misleading can a nominal 95% success-rate interval be when only twenty independent tasks are observed?
Open the brief RESEARCH BRIEF · 4 SOURCESCheck the false-positive cost of trying many variants
How does testing twenty null variants change the chance of declaring at least one apparent improvement?
Open the brief RESEARCH BRIEF · 7 SOURCESDesign a generative-model calibration check
Can inference recover declared generating quantities without conflating posterior predictive fit and parameter calibration?
Open the briefExplore the sources.
All 195 resources →Iris
Flower measurements for a compact, interpretable multiclass baseline and leakage audit.
Exact binomial coverage of Wald and Wilson intervals
Enumerate all binomial outcomes at n=20 and p=0.05, then compare exact interval-coverage calculations with a fixed-seed simulation.
SciPy
Numerical routines for integration, differential equations, optimization, transforms, and statistical analysis.
2D elastodynamic metamaterials
Pixelated metamaterial unit-cell designs and band-gap locations and widths for investigating simulation surrogate reliability.
3W dataset
Oil-production process signals for studying event detection across operating scenarios.
[Re] Easy Bayesian Transfer Learning with Informative Priors
A replication concerning Bayesian transfer learning with informative priors.
[Re] Fast-Activating Voltage- and Calcium-Dependent Potassium (BK) Conductance Promotes Bursting in Pituitary Cells: A Dynamic Clamp Study
A replication of a pituitary-cell conductance model, with uncertainty quantification identified in source keywords.
[Re] Measures for investigating the contextual modulation of information transmission
A replication concerning measures of contextual modulation in information transmission.
[Re] Modeling Insect Phenology Using Ordinal Regression and Continuation Ratio Models
An R replication of insect phenology modeling with ordinal regression and continuation-ratio models.
[Re] Replication Study of DECAF: Generating Fair Synthetic Data Using Causally-Aware Generative Networks
A replication of causally aware synthetic-data generation aimed at fairness evaluation.
[Re] Resampling methods for evaluating classification accuracy of wildlife habitat models
A replication of resampling methods for evaluating wildlife-habitat classification models.
[Re] Spatial constraints underlying the retinal mosaics of two types of horizontal cells in cat and macaque
An R replication of spatial constraints in retinal-cell mosaics.