Put uncertainty around a small agent-success benchmark
How misleading can a nominal 95% success-rate interval be when only twenty independent tasks are observed?
Make this plan your own ↓Read, edit and export without an account. This is preparation; no run or result is claimed.
What to compare
Keep n=20,p=0.05 and all 21 success counts fixed for the worked case; compare coverage with the nominal level. Do not infer coverage for correlated benchmark tasks from this binomial model.
Useful outputs
- Exact count-enumeration table for the fixed binomial model, with both interval methods.
- A separate Monte Carlo table with its fixed seed, replicate count and uncertainty interpretation.
- A transfer note stating when repeated agents/tasks violate the independent-trials model.
Inputs and prerequisites
- Original pilot jobs require the existing qualified isolated standard-library runtime.
- SciPy extensions require a separately qualified dependency environment.
What this work would not establish
- This is a worked model of benchmark uncertainty, not an evaluation of a named AI model.
- No sampler-calibration verdict or completed scientific check is implied.
Start with these sources.
Exact binomial coverage of Wald and Wilson intervals
Enumerate all binomial outcomes at n=20 and p=0.05, then compare exact interval-coverage calculations with a fixed-seed simulation.
Conditional moments of a bootstrap mean generator
Resample a fixed twenty-value empirical distribution and compare bootstrap moments with their conditional analytical values.
An exact small-sample permutation reference
Enumerate every balanced assignment for two groups of four observations and compare the exact two-sided randomization p-value with a sampled estimate.
bootstrap
The bootstrap reference explains percentile, basic and bias-corrected accelerated confidence intervals, paired resampling and computation controls.
permutation_test
A reference for permutation tests that distinguishes exchangeability structures, exact enumeration and randomized null sampling.
Make the question your own.
Edit freely. Export a copy before reloading or leaving. Your text stays in this tab until you choose to copy, download or save it privately.