Exact binomial coverage of Wald and Wilson intervals
Enumerate all binomial outcomes at n=20 and p=0.05, then compare exact interval-coverage calculations with a fixed-seed simulation.
Proposed original example. Source release and any execution are separate steps.
Do nominal 95% intervals achieve that coverage in this prespecified small-sample setting?
Weight coverage over all n+1 outcomes by their binomial probabilities; separately report Monte Carlo coverage and its finite-sample limitation.
What you could produce
- A scoped observations JSON record and a comparison with the stated reference.
Before you use it
- Python 3.13 standard library
Limits to keep in view
- No research code was executed by the preparation tool.
- Author output cannot issue an independent scientific-verification result.
Source and permission context
Original local preparation by the Executable Science seed collection; upstream API references remain separately attributed.
Rights need review. Review the scope and upstream conditions before reuse.
Still unresolved
- Proposed local-draft licenses: original code MIT, explanations CC-BY-4.0, synthetic numeric data CC0-1.0; publication/disclosure approval remains separate.
Put this resource to work.
Put uncertainty around a small agent-success benchmark
How misleading can a nominal 95% success-rate interval be when only twenty independent tasks are observed?
Open the brief RESEARCH BRIEF · 6 SOURCESExplain a score difference before blaming the model
Do evaluation harnesses agree on the same frozen predictions once normalization, aggregation and failure handling are made explicit?
Open the brief