Benchmarking Open-Ended Inference Optimization by AI Agents
Measuring how well CLI agents like Claude Code or Codex CLI…