Bench Agent Discovery

Discover public AI agents, reusable recipes, and trusted benchmark evidence by task.

我該用這個嗎

品質與安全性

A
說明品質
93%
結構描述完整度
71%
命名品質
100%
汙染風險
100%
權限相符程度
100%
協定合規性
100%

根據工具定義與協定合規性的自動化分析。

上下文成本

~582Token(工具定義)
~1.7 KB典型回應大小
極小的注意力影響(128k 上下文的 0.45%)

這是每次將伺服器的工具載入模型上下文時所消耗的約略 token 數量。數量越高,可用於其他工作的注意力就越少。

安裝

一鍵安裝

將以下內容加入你的 `claude_desktop_config.json` 檔案:

{
  "mcpServers": {
    "bench": {
      "url": "https://bench.virajmishratakehome.workers.dev/mcp"
    }
  }
}

遠端端點

https://bench.virajmishratakehome.workers.dev/mcpstreamable-http

它能做什麼

工具清單

工具(3)

🟢 唯讀🟡 寫入🔴 刪除⚪ 未知
🟢search_agents(query, category, framework, model, reusable, ...)

Find listed public agents by task, capability, category, framework, model, verified evidence, or reuse configuration. Owner telemetry and controlled benchmark evidence are returned separately.

輸入結構描述

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "maxLength": 120,
      "description": "Task or capability to search for, such as grounded research or code review."
    },
    "category": {
      "type": "string",
      "maxLength": 60
    },
    "framework": {
      "type": "string",
      "maxLength": 60
    },
    "model": {
      "type": "string",
      "maxLength": 80
    },
    "reusable": {
      "type": "boolean",
      "description": "True returns agents whose owners configured an invocation policy and capability manifest."
    },
    "verified": {
      "type": "boolean",
      "description": "True returns agents with at least one trusted-runner-verified benchmark submission."
    },
    "license": {
      "type": "string",
      "maxLength": 40,
      "description": "Exact SPDX-style license id from the agent's manifest provenance, such as MIT or Apache-2.0."
    },
    "liveCallable": {
      "type": "boolean",
      "description": "True returns agents with a reusable invocation policy and an owner-verified, currently reachable endpoint."
    },
    "maxCostPerRunUsd": {
      "type": "number",
      "minimum": 0,
      "description": "Upper bound on lifetime total_cost_usd / total_runs, i.e. average observed cost per run."
    },
    "maxP50LatencyMs": {
      "type": "integer",
      "minimum": 0,
      "description": "Upper bound on the agent's observed p50 latency in milliseconds."
    },
    "sort": {
      "type": "string",
      "enum": [
        "verified",
        "recent",
        "runs"
      ],
      "default": "verified"
    },
    "limit": {
      "type": "integer",
      "minimum": 1,
      "maximum": 20,
      "default": 10
    }
  },
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢get_agent(handle)

Get one public agent's recipe, public capability manifest, coarse invocation status, owner telemetry, and verified benchmark submissions.

輸入結構描述

{
  "type": "object",
  "properties": {
    "handle": {
      "type": "string",
      "maxLength": 101,
      "description": "Bench handle in @owner/agent-slug form."
    }
  },
  "required": [
    "handle"
  ],
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢list_benchmarks

List public, versioned benchmark contracts and only their trusted-runner-verified submissions.

輸入結構描述

{
  "type": "object",
  "properties": {},
  "$schema": "http://json-schema.org/draft-07/schema#"
}

社群

為此伺服器評分

證據

近期觀測

已驗證未記錄版本3 個工具
已驗證未記錄版本3 個工具