Model comparisons, agent setups, cost, and modalities, measured on real tasks. Read the methodology and results, or rerun a study on your own data.
Loop and Braintrust MCP + Codex each ran five workflows on identical 200-trace projects. Each tool was scored against a handwritten rubric measuring how well it surfaced insights about the traces and cited supporting evidence so traceability could be verified.
22 condition-neutral checks
39% faster
Four models answered 1,329 questions about specific events from LiveNewsBench. Each model ran without web search and with limited and wide You.com Search API access. GPT-5.6 Terra and Claude Sonnet 5 also used their built-in search tools.
Each enforcement variant was scored 2 ways. First by whether they adhered to that behavior, and second by the number of tests passed. We found that the agent could have leaky behavior that wasn't surfaced by output-only scoring.
30 tasksdeterministic detector
test_passed on the same 30 tasks
Each skill packages best practices from leading AI teams and researchers into a plain SKILL.md file your coding agent can read. Use them to get started running evals on your data.
Ask it to do what the skill describes, like validating a scorer against human labels or finding hidden failure modes.