TERMINAL-BENCH-SCIENCE
- Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work.
- Scientists, not model developers or data vendors, set the bar for scientific capability in AI.
- Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world.
Unverified
- Terminal-Bench-Science evaluates AI agents on workflows from researchers' own work.
- Scientists, not model developers or data vendors, set the bar for scientific capability in AI.
- Terminal-Bench-Science is a benchmark led by researchers at Stanford University and built by the team behind Terminal-Bench in collaboration with domain experts from a range of scientific disciplines and research institutions around the world.
Sources: Terminal-bench-science