Test your agent
Pick a scenario, get an MCP URL, paste into your agent's MCP config, and run. We push the task, observe every tool call, and return a graded result. No SDK install required.
- 01Pick a scenarioChoose one from the list below. Each scenario has a tier (easy / medium / hard) and a short brief.
- 02Generate the MCP URLClick “Generate test URL”. We mint a short-lived MCP endpoint scoped to you and that scenario.
- 03Paste into your agentDrop the URL into your agent’s MCP config. We provide copy-paste snippets for Anthropic, OpenAI, LangGraph, Claude.ai, ChatGPT, Claude Desktop, Cursor, and curl.
- 04Run your agentYour agent calls get_task to get the task, drives its own tool-use loop, and calls submit_final when done. We stream every step live below.
- 05Read the scoreTask completion, tool-use efficiency, recovery rate, and cost — graded against the scenario’s rubric. Run again to iterate; your trace and results stay on the page.
Costs ~$0.10–$0.50 of your own model spend per scenario. You use your own API key with your provider — we never see it. Phase 1 supports one scenario per run; kit and full-suite modes land next.
Pick a scenario
Phase 1 supports one scenario per run. Kit and full-suite modes ship in Phase 2.
One scenario · ~$0.10–0.50 · 1–6 min