AgentBench
Multi-environment agent benchmark repository for investigating task adapters and evaluation consistency.
External resource. No execution or independent verification is claimed here.
Can a single reviewed environment separate malformed agent actions, infrastructure errors and genuine task failures?
Prespecify finite trajectories and expected environment states; independently check the adapter's state transitions and score calculation.
What you could produce
- A version-pinned protocol stating inputs, rights, expected behavior, tolerances and resource limits before execution.
- A retained per-case result and failure ledger with an independently controlled comparison, if qualified execution is later authorized.
Before you use it
- Python and environment-specific services, containers or models; each environment needs isolated qualification and a separate budget.
- A separately qualified isolated runtime with a reviewed, pinned dependency and input closure.
Limits to keep in view
- No source program, example, build hook, package, dataset, model or generated research code has been executed or downloaded as a payload.
- The documented self-contained Python 3.13, 90-second pilot does not establish support for this package, its compiled dependencies, GPUs, services or agent sandboxes.
- Installation success, scientific outcomes, runtime compatibility, latency, memory use, API costs and security properties are unmeasured.
Source and permission context
Preserve the upstream project name, repository, exact source commit and applicable contributor/notices; resolve upstream citation guidance for any later formal use.
License evidence recorded. Review the scope and upstream conditions before reuse.
code · Apache-2.0
Apache License version 2.0 is identified in the pinned root license text. Observation is limited to LICENSE at commit d1e4a10db08c87075c78972e48ecc182be03e2d5; this is not blanket artifact clearance.
Inspect the license evidence ↗Still unresolved
- Only the cited license and README texts were observed; file exceptions, dependency closure, vendored components and submodules are not comprehensively audited.
- Dataset files, task prompts, generated outputs, model weights, tokenizer assets and hosted APIs are not cleared by a root code license.
- README/documentation reuse rights and version-specific citation guidance remain separately unresolved; only links and original descriptions are retained.
- Individual environments, benchmark content, external services and linked models are not presumed covered by the repository's root code terms.