Test your agent

Pick a scenario, get an MCP URL, paste into your agent's MCP config, and run. We push the task, observe every tool call, and return a graded result. No SDK install required.

  1. 01
    Pick a scenario
    Choose one from the list below. Each scenario has a tier (easy / medium / hard) and a short brief.
  2. 02
    Generate the MCP URL
    Click “Generate test URL”. We mint a short-lived MCP endpoint scoped to you and that scenario.
  3. 03
    Paste into your agent
    Drop the URL into your agent’s MCP config. We provide copy-paste snippets for Anthropic, OpenAI, LangGraph, Claude.ai, ChatGPT, Claude Desktop, Cursor, and curl.
  4. 04
    Run your agent
    Your agent calls get_task to get the task, drives its own tool-use loop, and calls submit_final when done. We stream every step live below.
  5. 05
    Read the score
    Task completion, tool-use efficiency, recovery rate, and cost — graded against the scenario’s rubric. Run again to iterate; your trace and results stay on the page.

Costs ~$0.10–$0.50 of your own model spend per scenario. You use your own API key with your provider — we never see it. Phase 1 supports one scenario per run; kit and full-suite modes land next.

Pick a scenario

Phase 1 supports one scenario per run. Kit and full-suite modes ship in Phase 2.

One scenario · ~$0.10–0.50 · 1–6 min