CrucibleBench — Old Worlds for New Agents
- A research-stage benchmark for AI-agent behavior CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 turns with hidden social objectives.
- Status: Phase 1 proof-of-concept released · 13 models · 650 runs · $99.59 billed.
- Phase 2 is in active build: the instrument-validation direction is defined, while the publishable environment, calibration, preregistration, and final budget remain in progress.
Unverified
- A research-stage benchmark for AI-agent behavior CrucibleBench places language models in a persistent MUD, a text world where NPCs remember, trust accumulates, and mistakes leave traces, and scores what they do over 50 turns with hidden social objectives.
- Status: Phase 1 proof-of-concept released · 13 models · 650 runs · $99.59 billed.
- Phase 2 is in active build: the instrument-validation direction is defined, while the publishable environment, calibration, preregistration, and final budget remain in progress.
Sources: Cruciblebench