Bench Agent Discovery
Discover public AI agents, reusable recipes, and trusted benchmark evidence by task.
我該用這個嗎
品質與安全性
A
根據工具定義與協定合規性的自動化分析。
上下文成本
~582Token(工具定義)
~1.7 KB典型回應大小
極小的注意力影響(128k 上下文的 0.45%)
這是每次將伺服器的工具載入模型上下文時所消耗的約略 token 數量。數量越高,可用於其他工作的注意力就越少。
安裝
一鍵安裝
將以下內容加入你的 `claude_desktop_config.json` 檔案:
{
"mcpServers": {
"bench": {
"url": "https://bench.virajmishratakehome.workers.dev/mcp"
}
}
}遠端端點
https://bench.virajmishratakehome.workers.dev/mcpstreamable-http它能做什麼
工具清單
工具(3)
🟢 唯讀🟡 寫入🔴 刪除⚪ 未知
🟢search_agents(query, category, framework, model, reusable, ...)
Find listed public agents by task, capability, category, framework, model, verified evidence, or reuse configuration. Owner telemetry and controlled benchmark evidence are returned separately.
輸入結構描述
{
"type": "object",
"properties": {
"query": {
"type": "string",
"maxLength": 120,
"description": "Task or capability to search for, such as grounded research or code review."
},
"category": {
"type": "string",
"maxLength": 60
},
"framework": {
"type": "string",
"maxLength": 60
},
"model": {
"type": "string",
"maxLength": 80
},
"reusable": {
"type": "boolean",
"description": "True returns agents whose owners configured an invocation policy and capability manifest."
},
"verified": {
"type": "boolean",
"description": "True returns agents with at least one trusted-runner-verified benchmark submission."
},
"license": {
"type": "string",
"maxLength": 40,
"description": "Exact SPDX-style license id from the agent's manifest provenance, such as MIT or Apache-2.0."
},
"liveCallable": {
"type": "boolean",
"description": "True returns agents with a reusable invocation policy and an owner-verified, currently reachable endpoint."
},
"maxCostPerRunUsd": {
"type": "number",
"minimum": 0,
"description": "Upper bound on lifetime total_cost_usd / total_runs, i.e. average observed cost per run."
},
"maxP50LatencyMs": {
"type": "integer",
"minimum": 0,
"description": "Upper bound on the agent's observed p50 latency in milliseconds."
},
"sort": {
"type": "string",
"enum": [
"verified",
"recent",
"runs"
],
"default": "verified"
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 20,
"default": 10
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_agent(handle)
Get one public agent's recipe, public capability manifest, coarse invocation status, owner telemetry, and verified benchmark submissions.
輸入結構描述
{
"type": "object",
"properties": {
"handle": {
"type": "string",
"maxLength": 101,
"description": "Bench handle in @owner/agent-slug form."
}
},
"required": [
"handle"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_benchmarks
List public, versioned benchmark contracts and only their trusted-runner-verified submissions.
輸入結構描述
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}社群
證據
近期觀測
已驗證未記錄版本3 個工具
已驗證未記錄版本3 個工具