FetchSandbox
A deterministic verification engine for agents. Proves a fix: fails on the old code, passes on new.
我該用這個嗎
品質與安全性
根據工具定義與協定合規性的自動化分析。
上下文成本
這是每次將伺服器的工具載入模型上下文時所消耗的約略 token 數量。數量越高,可用於其他工作的注意力就越少。
安裝
一鍵安裝
將以下內容加入你的 `claude_desktop_config.json` 檔案:
{
"mcpServers": {
"mcp": {
"command": "npx",
"args": [
"fetchsandbox-mcp"
]
}
}
}可執行的套件
0.7.0stdio遠端端點
https://fetchsandbox.com/mcp/v1streamable-http它能做什麼
工具清單
工具(7)
🟢coach(intent, session_id, user_response, context)
Conversational integration coach for FetchSandbox. Server-side orchestrator that walks the user through adding an API integration (payments / email / auth / etc.) — intake the goal, elicit domain-aware discovery questions from the spec's brain.yaml, route to the right workflow, prove the contract via FetchSandbox, surface compliance notes. Call this BEFORE any other FetchSandbox tool when the user has an open-ended 'help me add X', 'integrate X', 'test my X integration' ask. BEHAVIOR — strict, do exactly this each turn: (1) Say `message_for_user` to the user (verbatim or lightly paraphrased to fit your voice — but don't add new content). (2) If `next_action=wait_for_user` AND `options` is present + non-empty: USE YOUR CLIENT'S NATIVE QUESTION-PICKER TOOL (in Cursor / Claude Code this is `AskUserQuestion`) to render the lettered picker with the `question` text, the `options[].label` as rows, `default_option` as default, and an 'Other...' freeform row when `allow_freeform=true`. When the user picks or types, call coach again with `{session_id, user_response: <picked value or freeform text>}`. (3) If `next_action=wait_for_user` AND no `options`: just wait for free text. (4) If `next_action=call_tool`: invoke `tool_call.tool` with `tool_call.args`, then call coach again with the result in `context`. (5) If `next_action=done`: end the session. For a hosted app-builder request to build or test an integration, keep the user's simple goal intact and let the server route all named providers into one validation session. When FetchSandbox returns twin bindings, treat those as temporary TEST credentials: configure them server-side in the app-builder project, never ask the user for live Paddle/Resend keys, and never claim they were applied until the builder reports the change and the verifier observes app traffic. Infer app configuration details from the project when available; ask the user only for a fact the builder cannot discover. The deterministic verifier, not this conversation, owns the verdict. The state machine is server-side — DON'T try to predict the next step or skip ahead; let the server drive.
輸入結構描述
{
"type": "object",
"properties": {
"intent": {
"type": "string",
"description": "User's free-form integration ask, on the FIRST call only. Pass through verbatim — the server's intent router benefits from the full phrasing."
},
"session_id": {
"type": "string",
"description": "Returned by a previous coach call. Required on every call after the first."
},
"user_response": {
"type": "string",
"description": "The user's reply to the previous coach turn's question. Required when the previous turn returned `next_action: wait_for_user`."
},
"context": {
"type": "object",
"description": "Optional context the LLM brings to the turn — e.g. a summary of the user's repo (if you ran an introspect step), or the result of a previously-instructed `tool_call`."
}
}
}🟢guide(intent, hints)
ROUTE a symptom to the provider behaviour that explains it. Use this when the user names a provider or a domain (payments, email, auth, SMS, subscriptions) and you need to know what that provider ACTUALLY does — not what its docs say, and not what can be inferred from reading the integration code. Reading the code tells you what your app does with a field. It cannot tell you what the field MEANS at the provider — whether a line item is a seat count, whether a 200 body carries ok:false, whether an event can arrive out of order. That is the class of bug this routes. Returns {spec, workflow, scenario, confidence, reasoning} plus `next_actions`: a typed list of what to call next, with arguments pre-filled. Follow it rather than improvising the next step.
輸入結構描述
{
"type": "object",
"properties": {
"intent": {
"type": "string",
"description": "The developer's free-form prompt as they typed it. Don't pre-process or shorten — the router benefits from the full phrasing (capture timing, geo cues, failure mode language)."
},
"hints": {
"type": "object",
"description": "Optional caller-supplied overrides. Each field short-circuits the corresponding detection step.",
"properties": {
"spec": {
"type": "string",
"description": "Force a specific spec slug."
},
"workflow": {
"type": "string",
"description": "Force a specific workflow_id."
},
"scenario": {
"type": "string",
"description": "Force a specific failure scenario name."
}
}
}
},
"required": [
"intent"
]
}⚪quickrun(spec_slug, workflow_name, scenario)
Run a curated proof workflow against a KNOWN, bundled spec (stripe, clerk, descope, resend, twilio, and 50+ others) in ONE call — it spins up the sandbox by slug, so you do NOT need import_spec or a sandbox_id first. This is the normal path when the user is testing an integration with a well-known provider: call `guide` on their prompt, then call `quickrun` with the returned spec + workflow. To reproduce a failure, pass the `scenario` from guide's matched_bug_pattern.reproduce_with.scenario (e.g. webhook_retries, payment_declined). Returns sandbox_id + flow_run_id — pass BOTH to verify_behavior to prove the fix — plus a receipt URL. Use run_workflow instead ONLY when you already hold a sandbox_id from import_spec of a custom/private spec.
輸入結構描述
{
"type": "object",
"properties": {
"spec_slug": {
"type": "string",
"description": "The bundled spec slug from guide (e.g. 'stripe'). Lowercase."
},
"workflow_name": {
"type": "string",
"description": "The workflow id from guide (e.g. 'accept_payment')."
},
"scenario": {
"type": "string",
"description": "OPTIONAL failure scenario to reproduce (e.g. webhook_retries, payment_declined). Take it from guide's matched_bug_pattern.reproduce_with.scenario. Omit for the happy path."
}
},
"required": [
"spec_slug",
"workflow_name"
],
"additionalProperties": false
}🟡validate_integration(providers, app_base_url, session_id, arm, probe, ...)
Start a validation session for the user's integration. quickrun and run_workflow execute OUR curated workflow against a twin — green there says the provider behaves as documented, and says nothing about whether their integration wired to it correctly. This instead hands back a FRESH twin per provider and asks the user to point their app at it, then reads external provider traffic. Direct agent calls also appear there; caller identity is not established. Use it when someone wants their integration verified before launch, or when you have just run a happy path and need to say what it did and did not prove. Call once with `providers` to start, then again with the returned `session_id` after their app has run a flow. The matrix is returned before any traffic: pass `arm` with a probe ID before exercising the app. Send the returned request_headers on every provider request in that attempt, then pass `probe` and `run_id` to evaluate. Use `cancel: true` with `run_id` to abort or retry failed cleanup. Request logs alone do not prove application state or production readiness. For a supported payment-to-receipt integration, follow the returned next_tool_call immediately on the SAME session; do not substitute guide, quickrun, list_workflows, or reference probes for application verification. Use the widest suite the server recommends. receipt_recovery_v1 includes R1-R4 and adds checks R5-R8 for accepted-send response loss, concurrent redelivery, source/email transient failures and delayed redelivery. Pass suite plus receipt_config with the app's exact webhook URL, test recipients and observable purchase marker. FetchSandbox freezes all four receipt rules for both same-customer and different-customer purchases. Configure the actual app with the issued twins, signing secret and request headers; then call execute:true with its run_id. FetchSandbox drives signed events and independently observes accepted email records. Keep every check and the observation windows in your report. If you cannot configure the app, report that blocker; reference tests cannot replace app execution. Each validation session can run only one receipt suite. After its terminal result, start a new session with providers ['paddle','resend'] for fresh twins before retrying; a completed session cannot be reused. A verified receipt suite covers only its declared rules and windows, never all production behavior. A missing or untriggered recovery fault stays unmeasured. This application suite requires sign-in.
輸入結構描述
{
"type": "object",
"properties": {
"providers": {
"type": "array",
"items": {
"type": "string"
},
"description": "Provider slugs the app integrates, e.g. ['stripe'] or ['paddle','resend']. Required to start a session."
},
"app_base_url": {
"type": "string",
"description": "Where their app is deployed, if known. Recorded so a later step can deliver webhooks to it."
},
"session_id": {
"type": "string",
"description": "Returned by the first call. Pass it back after their app has run a flow to see what it did."
},
"arm": {
"type": "string",
"description": "Probe ID from the matrix to arm before running the app. Requires session_id."
},
"probe": {
"type": "string",
"description": "Probe ID from the matrix to evaluate after running the app. Requires session_id and run_id."
},
"run_id": {
"type": "string",
"description": "Attempt ID returned by arm. Required with probe or cancel; prevents stale calls from consuming a newer attempt."
},
"cancel": {
"type": "boolean",
"description": "Abort the attempt or retry failed cleanup. Requires session_id and run_id; cannot combine with arm or probe."
},
"suite": {
"type": "string",
"enum": [
"receipt_delivery_v1",
"receipt_recovery_v1"
],
"description": "Configure the declared application receipt suite on this session."
},
"receipt_config": {
"type": "object",
"properties": {
"webhook_url": {
"type": "string",
"description": "Exact normal application webhook URL, including path."
},
"recipient_a": {
"type": "string",
"description": "Expected test recipient for purchase A, fixed before execution."
},
"recipient_b": {
"type": "string",
"description": "Different test recipient for the different-customer case."
},
"purchase_marker": {
"type": "object",
"properties": {
"field": {
"type": "string",
"enum": [
"html",
"text",
"subject",
"tags"
]
},
"tag_name": {
"type": "string"
}
},
"required": [
"field"
],
"additionalProperties": false,
"description": "Where the normal receipt includes its Paddle transaction ID; same-record recipient/marker matching is required."
},
"observation_seconds": {
"type": "number",
"minimum": 0.1,
"maximum": 5,
"description": "Observation window after each checkpoint; default 2 seconds. Verdict covers only measured windows."
},
"app_version": {
"type": "string",
"description": "Declared application build identifier; not code attestation."
},
"source_lookup_required": {
"type": "boolean",
"description": "Whether the app reads the source provider during its normal webhook handler."
},
"purchase_context": {
"type": "object",
"description": "Application routing custom_data, e.g. test workspace ID. Never credentials."
}
},
"required": [
"webhook_url",
"recipient_a",
"recipient_b",
"purchase_marker",
"app_version",
"source_lookup_required"
],
"additionalProperties": false
},
"execute": {
"type": "boolean",
"description": "Drive the configured app and evaluate the frozen receipt suite. Requires session_id/run_id; configuration acknowledgment alone is not proof."
}
},
"additionalProperties": false
}🟢list_workflows(spec_id, spec_slug)
List the named, runnable workflows for a previously-imported spec. Workflows are realistic multi-step API journeys (e.g. 'create customer → attach payment method → create subscription'). Use this after import_spec for exploration ("what can I do?", "show me the flows") OR before run_all_workflows when the user wants a SCOPED validation: list, filter by user intent ("checkout", "webhooks"), then pass the matching ids as `workflow_names` to run_all_workflows. Returns the workflows AND the failure scenarios this spec can simulate — each with a plain description and the exact call to run it. Read the scenarios out to the user and ask which matter; do not guess scenario names.
輸入結構描述
{
"type": "object",
"properties": {
"spec_id": {
"type": "string",
"description": "The spec_id returned by import_spec."
},
"spec_slug": {
"type": "string",
"description": "A human name like 'resend' or 'stripe'. Use this when you do not have a spec_id — you do not need to look one up first."
}
},
"additionalProperties": false
}🟢run_workflow(sandbox_id, workflow_name, scenario)
Execute ONE specific workflow by name and return its step-by-step trace PLUS a `share_url` — a public, replayable proof URL that renders the full timeline (every request, response, webhook event) for this run. The share_url is the canonical 'here's what happened' artifact: surface it verbatim in any reply that needs evidence (PR comments, Slack threads, blog posts, X replies). Do NOT substitute a docs URL or any other link as the proof — the share_url is the only valid receipt. Use ONLY when the user explicitly names a single workflow to run (e.g., "run accept_payment", "just check the refund workflow"). For ANY validation-style request — "validate stripe", "check coverage", "run all workflows", "test this integration", or even "validate stripe checkout" (multiple workflows match "checkout") — use `run_all_workflows` instead. The batch tool collapses N approvals to 1 and supports a workflow_names filter for scope. Calling this in a loop is an anti-pattern.
輸入結構描述
{
"type": "object",
"properties": {
"sandbox_id": {
"type": "string",
"description": "The sandbox_id returned by import_spec."
},
"workflow_name": {
"type": "string",
"description": "Workflow id or name from list_workflows. Case-insensitive; dashes and underscores are interchangeable."
},
"scenario": {
"type": "string",
"description": "OPTIONAL failure scenario to exercise (e.g. payment_declined, insufficient_funds, fraud_hold). Toggles the sandbox engine's scenario for the duration of the run, then restores. Use this for 'test with declined card' / 'simulate failure X' intents. Omit for the happy path."
}
},
"required": [
"sandbox_id",
"workflow_name"
],
"additionalProperties": false
}🟢verify_behavior(bug_pattern_id, prompt, sandbox_id, flow_run_id)
Prove a known fix survives a bug — the 'prove' half of reproduce→prove. The backend spawns a buggy AND a fixed reference handler and fires the bug_pattern's probes at both, returning the side-by-side diff (e.g. the buggy handler double-charges on a duplicate webhook, the fixed handler dedupes). Call this AFTER run_workflow reproduces a failure, when the matched bug_pattern has a simulation block, to show the fix actually holds — not just that the failure reproduced. Pass sandbox_id + flow_run_id from the run so the diff is saved onto that run's receipt URL. bug_pattern_id comes from guide's matched_bug_pattern. The buggy/fixed handlers are FetchSandbox reference implementations, NOT the user's code. A confirmed reference result does not verify the user's app or show that it is ready to ship. Keep the returned provider and evidence scope with any reported result.
輸入結構描述
{
"type": "object",
"properties": {
"bug_pattern_id": {
"type": "string",
"description": "The bug_pattern to prove (e.g. webhook_duplicate_side_effect). Comes from guide's matched_bug_pattern.id or the spec's brain."
},
"prompt": {
"type": "string",
"description": "OPTIONAL. The user's own description of the symptom. For patterns that can originate in the handler OR the provider, this classifies which side to simulate. Omit to run both."
},
"sandbox_id": {
"type": "string",
"description": "OPTIONAL. The sandbox from the run. Pass with flow_run_id to save the diff onto that run's receipt. This binds the pattern to that sandbox's provider; provide it when verifying a provider integration."
},
"flow_run_id": {
"type": "string",
"description": "OPTIONAL. The flow_run_id returned by run_workflow. Pass with sandbox_id so the receipt URL renders the diff alongside the steps."
}
},
"required": [
"bug_pattern_id"
],
"additionalProperties": false
}建議的提示詞
find_bugsfind_bugslist_specslist_specsfind_bugsset_scenario社群
證據