Bench Agent Discovery

Discover public AI agents, reusable recipes, and trusted benchmark evidence by task.

我该使用它吗

质量与安全性

A
描述质量
93%
模式完整度
71%
命名质量
100%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~582token 数(工具定义)
~1.7 KB典型响应大小
对注意力的影响极小(占 128k 上下文窗口的 0.45%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "bench": {
      "url": "https://bench.virajmishratakehome.workers.dev/mcp"
    }
  }
}

远程端点

https://bench.virajmishratakehome.workers.dev/mcpstreamable-http

它能做什么

工具清单

工具(3)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢search_agents(query, category, framework, model, reusable, ...)

Find listed public agents by task, capability, category, framework, model, verified evidence, or reuse configuration. Owner telemetry and controlled benchmark evidence are returned separately.

输入模式

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "maxLength": 120,
      "description": "Task or capability to search for, such as grounded research or code review."
    },
    "category": {
      "type": "string",
      "maxLength": 60
    },
    "framework": {
      "type": "string",
      "maxLength": 60
    },
    "model": {
      "type": "string",
      "maxLength": 80
    },
    "reusable": {
      "type": "boolean",
      "description": "True returns agents whose owners configured an invocation policy and capability manifest."
    },
    "verified": {
      "type": "boolean",
      "description": "True returns agents with at least one trusted-runner-verified benchmark submission."
    },
    "license": {
      "type": "string",
      "maxLength": 40,
      "description": "Exact SPDX-style license id from the agent's manifest provenance, such as MIT or Apache-2.0."
    },
    "liveCallable": {
      "type": "boolean",
      "description": "True returns agents with a reusable invocation policy and an owner-verified, currently reachable endpoint."
    },
    "maxCostPerRunUsd": {
      "type": "number",
      "minimum": 0,
      "description": "Upper bound on lifetime total_cost_usd / total_runs, i.e. average observed cost per run."
    },
    "maxP50LatencyMs": {
      "type": "integer",
      "minimum": 0,
      "description": "Upper bound on the agent's observed p50 latency in milliseconds."
    },
    "sort": {
      "type": "string",
      "enum": [
        "verified",
        "recent",
        "runs"
      ],
      "default": "verified"
    },
    "limit": {
      "type": "integer",
      "minimum": 1,
      "maximum": 20,
      "default": 10
    }
  },
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢get_agent(handle)

Get one public agent's recipe, public capability manifest, coarse invocation status, owner telemetry, and verified benchmark submissions.

输入模式

{
  "type": "object",
  "properties": {
    "handle": {
      "type": "string",
      "maxLength": 101,
      "description": "Bench handle in @owner/agent-slug form."
    }
  },
  "required": [
    "handle"
  ],
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢list_benchmarks

List public, versioned benchmark contracts and only their trusted-runner-verified submissions.

输入模式

{
  "type": "object",
  "properties": {},
  "$schema": "http://json-schema.org/draft-07/schema#"
}

社区

评价此服务器

证据

最近观测

已验证未记录版本3 个工具
已验证未记录版本3 个工具
已验证未记录版本3 个工具