Point a real agent at a real target. Debug your MCP, REST, A2A, WebMCP and Embedded Checkout Protocol integrations, benchmark models like Claude, ChatGPT and Gemini, and share session replays with your team.
Two runtimes decide where a session executes. Every surface drives both — the same session definition everywhere.
Store builders and agent builders — opposite sides of the same handshake, one harness.
Schema quality, identity enforcement and funnel completion — measured at runtime, against the store you actually run.
Every session runs through our router. Pick a model, pick a target, and watch it shop.
| Model | Provider | Class |
|---|---|---|
| Claude Opus 5 | Anthropic | Frontier |
| Claude Sonnet 5 | Anthropic | Frontier |
| Claude Opus 4.8 | Anthropic | Frontier |
| Claude Sonnet 4.6 | Anthropic | Frontier |
| GPT-5.5 | OpenAI | Frontier |
| GPT-6 Luna | OpenAI | Fast |
| GPT-4o | OpenAI | Mid-tier |
| o4-mini | OpenAI | Reasoning |
| Gemini 3.1 Pro | Frontier | |
| Gemini 2.5 Pro | Mid-tier | |
| Gemini 3.8 Flash | Fast | |
| Gemini 3.7 Flash | Fast | |
| Gemini 3.6 Flash | Fast |
| Model | Provider | Class |
|---|---|---|
| Grok 4.5 | xAI | Frontier |
| Grok 4.3 | xAI | Frontier |
| Kimi K3 | Moonshot AI | Open-weight |
| Qwen 3.7 Max | Alibaba | Open-weight |
| Qwen 3.8 Flash | Alibaba | Fast |
| DeepSeek R1 | DeepSeek | Open-weight |
| DeepSeek V3.2 | DeepSeek | Open-weight |
| DeepSeek V4 Pro | DeepSeek | Open-weight |
| DeepSeek V4.1 Flash | DeepSeek | Open-weight |
| DeepSeek V4 Flash | DeepSeek | Open-weight |
| Muse Spark 1.3 | Meta | Frontier |
| Llama 4 Maverick | Meta | Open-weight |
| Llama 3.3 70B | Meta | Open-weight |
Debug what worked, share what didn't. Embeddable session replays your whole team can review.
Real products pulled from real targets through their own MCP endpoints, right now — not a screenshot.
One team for your whole org. Shared sessions, private profiles, headless API access and usage tracking — built for storefront engineers and agent teams working in AI commerce.
Define multi-turn shopping sequences, run them across models and transports, and get pass/fail reports with funnel completion and token analysis.
A collection is a saved target × model × sequence. Run it from the API, from CI, or let it run on a schedule in the Cloud runtime — and get told when a verdict changes.
$ ucp run --collection checkout-flow → target thebodyshop.com · claude-sonnet-4.5 ✓ search_catalog 775ms ✓ get_product_details 263ms ✓ update_cart 583ms ✓ checkout_url reached 13.4s verdict: funnel 4/4 · 94,779 tokens
A Chrome side panel that runs a session on the tab you are already looking at. Your storefront on the left, what the agent actually received on the right.
When your product page says one thing and get_product_details says another, the two panes disagree on screen. Every tool call, argument, response and timing, beside the page that produced it.
It works on any target, not just yours — open a store that publishes a UCP profile and read its payloads in real time while you browse.

One harness for both halves of an agentic checkout.