Bench Agent Discovery
Discover public AI agents, reusable recipes, and trusted benchmark evidence by task.
使うべきか
品質と安全性
A
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
~582トークン数(ツール定義)
~1.7 KB一般的なレスポンスサイズ
注意への影響は最小限(128k コンテキストの 0.45%)
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"bench": {
"url": "https://bench.virajmishratakehome.workers.dev/mcp"
}
}
}リモートエンドポイント
https://bench.virajmishratakehome.workers.dev/mcpstreamable-httpできること
ツール一覧
ツール(3)
🟢 読み取り専用🟡 書き込み🔴 削除⚪ 不明
🟢search_agents(query, category, framework, model, reusable, ...)
Find listed public agents by task, capability, category, framework, model, verified evidence, or reuse configuration. Owner telemetry and controlled benchmark evidence are returned separately.
入力スキーマ
{
"type": "object",
"properties": {
"query": {
"type": "string",
"maxLength": 120,
"description": "Task or capability to search for, such as grounded research or code review."
},
"category": {
"type": "string",
"maxLength": 60
},
"framework": {
"type": "string",
"maxLength": 60
},
"model": {
"type": "string",
"maxLength": 80
},
"reusable": {
"type": "boolean",
"description": "True returns agents whose owners configured an invocation policy and capability manifest."
},
"verified": {
"type": "boolean",
"description": "True returns agents with at least one trusted-runner-verified benchmark submission."
},
"license": {
"type": "string",
"maxLength": 40,
"description": "Exact SPDX-style license id from the agent's manifest provenance, such as MIT or Apache-2.0."
},
"liveCallable": {
"type": "boolean",
"description": "True returns agents with a reusable invocation policy and an owner-verified, currently reachable endpoint."
},
"maxCostPerRunUsd": {
"type": "number",
"minimum": 0,
"description": "Upper bound on lifetime total_cost_usd / total_runs, i.e. average observed cost per run."
},
"maxP50LatencyMs": {
"type": "integer",
"minimum": 0,
"description": "Upper bound on the agent's observed p50 latency in milliseconds."
},
"sort": {
"type": "string",
"enum": [
"verified",
"recent",
"runs"
],
"default": "verified"
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 20,
"default": 10
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_agent(handle)
Get one public agent's recipe, public capability manifest, coarse invocation status, owner telemetry, and verified benchmark submissions.
入力スキーマ
{
"type": "object",
"properties": {
"handle": {
"type": "string",
"maxLength": 101,
"description": "Bench handle in @owner/agent-slug form."
}
},
"required": [
"handle"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_benchmarks
List public, versioned benchmark contracts and only their trusted-runner-verified submissions.
入力スキーマ
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}コミュニティ
エビデンス
最近の観測
検証済みバージョンは記録されていませんツール 3 件
検証済みバージョンは記録されていませんツール 3 件
検証済みバージョンは記録されていませんツール 3 件