Word Is Bond
Test the voice agents you run: scored transcripts, pass/fail verdicts, latency and WER metrics.
使うべきか
品質と安全性
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"wordis-bond": {
"url": "https://api.wordis-bond.com/mcp"
}
}
}リモートエンドポイント
https://api.wordis-bond.com/mcpstreamable-httpできること
ツール一覧
ツール(16)
🟢run_demo(goal, persona, language, expected, bargeIn)
Run the hosted demo voice agent (a dental front desk) end-to-end and get a real, fully-scored result in about a minute — no target of your own needed. Returns the score (0–100), pass/fail verdict, per-turn metrics, the transcript, and a shareable public report URL. Zero carrier cost. Optional inputs override the scenario.
入力スキーマ
{
"type": "object",
"properties": {
"goal": {
"type": "string",
"description": "What the synthetic caller should try to accomplish."
},
"persona": {
"type": "object",
"properties": {
"description": {
"type": "string"
}
},
"description": "Override the synthetic caller persona."
},
"language": {
"type": "string",
"description": "BCP-47 language tag, default \"en\"."
},
"expected": {
"type": "array",
"items": {
"type": "string"
},
"description": "Expected agent lines to pin for word-error-rate."
},
"bargeIn": {
"type": "boolean",
"description": "Inject one caller-initiated barge-in."
}
},
"additionalProperties": false
}🟢run_test(transcript, turns, scenarioGoal, assertions, targetAgent, ...)
Run a test against a voice agent you control. Two modes: (1) score a captured transcript offline — pass "transcript" (or "turns") plus "scenarioGoal"; the judge returns a scored run synchronously. (2) run a live synthetic call — pass "targetAgent" and "goal". A live "direct" (SIP/WebRTC) target returns a tokenized media WebSocket URL for your agent-side harness to dial; a "pstn" target places a real carrier call (pro+ plans) to a number you have verified — an unverified destination returns 403 NUMBER_NOT_VERIFIED and no call is placed. Poll get_run / GET /api/tests/{id} for the terminal scored state of a live run.
入力スキーマ
{
"type": "object",
"properties": {
"transcript": {
"type": "string",
"description": "Captured conversation to score (offline mode)."
},
"turns": {
"type": "array",
"description": "Per-turn capture; formatted into a transcript when \"transcript\" is omitted.",
"items": {
"type": "object",
"properties": {
"role": {
"type": "string",
"enum": [
"agent",
"caller"
]
},
"text": {
"type": "string"
}
}
}
},
"scenarioGoal": {
"type": "string",
"description": "What the conversation was meant to accomplish (offline mode)."
},
"assertions": {
"type": "array",
"items": {
"type": "string"
},
"description": "Extra pass/fail checks for the judge."
},
"targetAgent": {
"type": "object",
"description": "The voice agent under test — a system you run. transport \"direct\" (SIP/WebRTC) has near-zero cost and works on every plan; \"pstn\" places a real carrier call (pro+ plans) and only to a number you have verified. The hosted demo target is { \"transport\": \"direct\", \"peerId\": \"wb-demo-dental\" }.",
"properties": {
"transport": {
"type": "string",
"enum": [
"pstn",
"webrtc",
"direct",
"sip"
]
},
"toNumber": {
"type": "string",
"description": "E.164 number to dial (required for a pstn target). It must be a number this account has proven it controls: verify it first with verify_number_start then verify_number_confirm, otherwise the run is rejected with 403 NUMBER_NOT_VERIFIED and no call is placed."
},
"peerId": {
"type": "string",
"description": "Hosted-agent id, e.g. \"wb-demo-dental\"."
}
},
"required": [
"transport"
]
},
"goal": {
"type": "string",
"description": "What the synthetic caller should accomplish (live mode)."
},
"persona": {
"type": "object",
"properties": {
"name": {
"type": "string"
},
"description": {
"type": "string"
}
},
"description": "The synthetic caller persona (live mode)."
},
"suiteId": {
"type": "string",
"description": "Link this run to a suite (optional)."
},
"language": {
"type": "string",
"description": "BCP-47 language tag, default \"en\"."
},
"transport": {
"type": "string",
"enum": [
"pstn",
"webrtc"
],
"description": "Force a transport (live mode)."
},
"concurrency": {
"type": "integer",
"minimum": 1,
"maximum": 25,
"description": "Number of parallel live calls."
},
"expected": {
"type": "array",
"items": {
"type": "string"
}
},
"bargeIn": {
"type": "boolean"
},
"record": {
"type": "boolean",
"description": "Record this call to YOUR OWN storage (BYOS, default false). Requires the byosRecording capability (pro/enterprise) and an enabled recording target; the customer-owned pointer comes back in the run’s recording_url. wordis-bond keeps only the pointer, never the audio."
},
"recordTargetId": {
"type": "string",
"description": "Record to THIS target (else the most-recent enabled one)."
}
},
"additionalProperties": false
}🟢get_trends(days)
Read account-wide testing trends over a window: pass-rate, average score, and run-to-run regressions per suite, plus overall totals. Use this to spot behaviour drift in the voice agents you test.
入力スキーマ
{
"type": "object",
"properties": {
"days": {
"type": "integer",
"minimum": 1,
"maximum": 365,
"description": "Look-back window in days (default 90)."
}
},
"additionalProperties": false
}🟡list_suites
List your reusable test suites (each is a set of scenarios/personas pinned to a target voice agent). Returns their ids, names, targets, and schedules — use a suite id with run_test or get_trends.
入力スキーマ
{
"type": "object",
"properties": {},
"additionalProperties": false
}🟡create_suite(name, targetAgent, language, scenarios, schedule)
Create a reusable test suite: a named set of scenarios/personas pinned to a target voice agent you run. Optionally give it a schedule ("weekly" | "daily" | "hourly", plan-gated) so it runs automatically and flags drift.
入力スキーマ
{
"type": "object",
"properties": {
"name": {
"type": "string"
},
"targetAgent": {
"type": "object",
"description": "The voice agent under test — a system you run. transport \"direct\" (SIP/WebRTC) has near-zero cost and works on every plan; \"pstn\" places a real carrier call (pro+ plans) and only to a number you have verified. The hosted demo target is { \"transport\": \"direct\", \"peerId\": \"wb-demo-dental\" }.",
"properties": {
"transport": {
"type": "string",
"enum": [
"pstn",
"webrtc",
"direct",
"sip"
]
},
"toNumber": {
"type": "string",
"description": "E.164 number to dial (required for a pstn target). It must be a number this account has proven it controls: verify it first with verify_number_start then verify_number_confirm, otherwise the run is rejected with 403 NUMBER_NOT_VERIFIED and no call is placed."
},
"peerId": {
"type": "string",
"description": "Hosted-agent id, e.g. \"wb-demo-dental\"."
}
},
"required": [
"transport"
]
},
"language": {
"type": "string",
"description": "BCP-47 language tag, default \"en\"."
},
"scenarios": {
"type": "array",
"description": "Persona + goal pairs the suite exercises.",
"items": {
"type": "object",
"properties": {
"persona": {
"description": "A string, or { name, description }."
},
"goal": {
"type": "string"
},
"expected": {
"type": "array",
"items": {
"type": "string"
}
}
}
}
},
"schedule": {
"type": "string",
"description": "Cadence for scheduled runs, e.g. \"weekly\" | \"daily\" | \"hourly\"."
}
},
"required": [
"name",
"targetAgent"
],
"additionalProperties": false
}⚪verify_number_start(number)
Prove your account controls a phone number, which it must do before any PSTN test call to it. Word Is Bond places one short call to the number, speaks a 6-digit code twice, and hangs up. Confirm the code within ten minutes with verify_number_confirm. Calling an already-verified number places no call. This is the only action that dials an unverified number and it is tightly capped (3 calls per number and 5 distinct numbers per account, per day).
入力スキーマ
{
"type": "object",
"properties": {
"number": {
"type": "string",
"description": "The number to verify, in E.164 (e.g. \"+15045204977\")."
}
},
"required": [
"number"
],
"additionalProperties": false
}⚪verify_number_confirm(number, code, label)
Finish verifying a phone number by supplying the 6-digit code spoken on the verification call. On success the number becomes a permitted PSTN test destination for your account. The code is single-use, expires after ten minutes, and locks after five incorrect attempts.
入力スキーマ
{
"type": "object",
"properties": {
"number": {
"type": "string",
"description": "The number being verified, in E.164."
},
"code": {
"type": "string",
"description": "The 6 digits spoken on the call."
},
"label": {
"type": "string",
"description": "Optional label stored alongside the verified number."
}
},
"required": [
"number",
"code"
],
"additionalProperties": false
}🔴list_verified_numbers
List the phone numbers your account has proven it controls. Only these numbers (and Word Is Bond DIDs) may be used as a "pstn" targetAgent.toNumber. Revoke one with DELETE /api/numbers/{id}.
入力スキーマ
{
"type": "object",
"properties": {},
"additionalProperties": false
}🟡register_recording_target(name, callbackUrl, track)
Register where BYOS call recordings go, so you can pass "record": true to run_test and have that call’s audio teed to YOUR OWN storage. wordis-bond keeps only a pointer (the run’s recording_url), never the audio. "callbackUrl" is a public https endpoint that returns a presigned PUT URL per recording (so wordis-bond never holds your cloud credentials). Pro/enterprise capability, bundled free — you pay your own storage; starter → 402.
入力スキーマ
{
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "A label for this target."
},
"callbackUrl": {
"type": "string",
"description": "Public https endpoint that mints a presigned PUT URL per recording."
},
"track": {
"type": "string",
"enum": [
"inbound",
"outbound",
"both"
],
"description": "Which side to capture (default inbound = the agent)."
}
},
"required": [
"name",
"callbackUrl"
],
"additionalProperties": false
}⚪import_flow(platform, config, save, name)
Import a phone-system flow or voice-agent config from a platform (e.g. Vapi) into the canonical, diffable Flow IR — the first step of putting your phone system under version control. Reports the fields the IR abstracts away. Pass save:true to persist it as a versioned flow. The same IR exports back out, so it doubles as a migration surface between platforms.
入力スキーマ
{
"type": "object",
"properties": {
"platform": {
"type": "string",
"enum": [
"vapi"
],
"description": "The source platform."
},
"config": {
"type": "object",
"description": "The platform's native flow/agent object."
},
"save": {
"type": "boolean",
"description": "Persist the imported IR as a new versioned flow."
},
"name": {
"type": "string",
"description": "Name for the saved flow (defaults to the config name)."
}
},
"required": [
"platform",
"config"
],
"additionalProperties": false
}🟡diff_flow(flowId, from, to)
Show the structural diff between two versions of a version-controlled flow: which nodes and edges were added, removed, or changed. This is the code review for your phone system — see exactly what a change did before you ship it.
入力スキーマ
{
"type": "object",
"properties": {
"flowId": {
"type": "string",
"description": "The flow id."
},
"from": {
"type": "integer",
"minimum": 1,
"description": "Base version number."
},
"to": {
"type": "integer",
"minimum": 1,
"description": "Target version number (defaults to the current version)."
}
},
"required": [
"flowId",
"from"
],
"additionalProperties": false
}🔴test_flow(flowId)
Run the regression gate on a flow now: compile the current version into a synthetic-caller test, run it against the flow's target, and compare the result to the previous version's baseline. A behavior change that regressed (a pass turning into a fail, or a score drop past the threshold) is caught and blocks the change — continuous integration for your phone-system logic.
入力スキーマ
{
"type": "object",
"properties": {
"flowId": {
"type": "string",
"description": "The flow id to gate."
}
},
"required": [
"flowId"
],
"additionalProperties": false
}🟢list_monitors
List your production monitors and their current health (unknown | healthy | degraded | critical). A monitor watches ONE live production line: you stream it completed calls, and it scores each with the same versioned judge that scores your tests, tracks a rolling baseline, and alerts when quality drifts. Requires a pro or enterprise plan.
入力スキーマ
{
"type": "object",
"properties": {},
"additionalProperties": false
}🟡create_monitor(name, scenarioGoal, assertions, rubric, language, ...)
Create a production monitor. "scenarioGoal" is what GOOD looks like for this agent — the judge scores every ingested call against it, exactly as a test scenario goal works. Optional "assertions" are plain-English checks the agent must satisfy. Read the ingest secret afterwards from GET /api/monitor/{id}/secret. Requires a pro or enterprise plan.
入力スキーマ
{
"type": "object",
"properties": {
"name": {
"type": "string",
"description": "Human name, e.g. \"Support line — main agent\"."
},
"scenarioGoal": {
"type": "string",
"description": "What the agent is supposed to accomplish on every call."
},
"assertions": {
"type": "array",
"items": {
"type": "string"
},
"description": "Plain-English checks the agent must satisfy."
},
"rubric": {
"type": "string",
"description": "Extra free-text rubric appended to the judge instructions."
},
"language": {
"type": "string",
"description": "BCP-47 language tag, default \"en\"."
},
"sampleRate": {
"type": "number",
"description": "Fraction 0..1 of ingested calls to score. Default 1 (score every call)."
},
"alertWebhookUrl": {
"type": "string",
"description": "Optional public https URL to receive signed drift alerts."
}
},
"required": [
"name",
"scenarioGoal"
],
"additionalProperties": false
}🟡ingest_call(monitorId, transcript, externalId, platform, durationSec, ...)
Send one completed PRODUCTION call to a monitor to be scored. Returns 202 immediately; scoring runs in the background and the monitor health updates. The transcript is scored in flight and never stored — only the scorecard and safe metadata are kept. Pass "externalId" (your own call id) so a re-delivered call scores exactly once.
入力スキーマ
{
"type": "object",
"properties": {
"monitorId": {
"type": "string",
"description": "The monitor to ingest into."
},
"transcript": {
"type": "string",
"description": "Formatted \"AGENT: … / CALLER: …\" transcript of the finished call."
},
"externalId": {
"type": "string",
"description": "Your own call id — makes the ingest idempotent."
},
"platform": {
"type": "string",
"description": "Where the call ran, e.g. \"retell\", \"vapi\", \"telnyx\"."
},
"durationSec": {
"type": "number",
"description": "Call duration in seconds."
},
"occurredAt": {
"type": "string",
"description": "ISO-8601 timestamp of when the call happened."
}
},
"required": [
"monitorId",
"transcript"
],
"additionalProperties": false
}🟢get_monitor_health(monitorId)
Read a monitor's live rolling health and drift. Health is unknown | healthy | degraded | critical. Drift is isolated by judge version: when our judge changes, the boundary is reported as judgeVersionChanged and NEVER as an agent regression — a score delta across that boundary says nothing about your agent.
入力スキーマ
{
"type": "object",
"properties": {
"monitorId": {
"type": "string",
"description": "The monitor to read."
}
},
"required": [
"monitorId"
],
"additionalProperties": false
}コミュニティ
エビデンス