Bench Agent Discovery

Discover public AI agents, reusable recipes, and trusted benchmark evidence by task.

사용해야 할까요

품질 및 안전성

A
설명 품질
93%
스키마 완전성
71%
이름 품질
100%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~582토큰 (도구 정의)
~1.7 KB일반적인 응답 크기
최소한의 주의 영향 (128k 컨텍스트의 0.45%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "bench": {
      "url": "https://bench.virajmishratakehome.workers.dev/mcp"
    }
  }
}

원격 엔드포인트

https://bench.virajmishratakehome.workers.dev/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (3)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢search_agents(query, category, framework, model, reusable, ...)

Find listed public agents by task, capability, category, framework, model, verified evidence, or reuse configuration. Owner telemetry and controlled benchmark evidence are returned separately.

입력 스키마

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "maxLength": 120,
      "description": "Task or capability to search for, such as grounded research or code review."
    },
    "category": {
      "type": "string",
      "maxLength": 60
    },
    "framework": {
      "type": "string",
      "maxLength": 60
    },
    "model": {
      "type": "string",
      "maxLength": 80
    },
    "reusable": {
      "type": "boolean",
      "description": "True returns agents whose owners configured an invocation policy and capability manifest."
    },
    "verified": {
      "type": "boolean",
      "description": "True returns agents with at least one trusted-runner-verified benchmark submission."
    },
    "license": {
      "type": "string",
      "maxLength": 40,
      "description": "Exact SPDX-style license id from the agent's manifest provenance, such as MIT or Apache-2.0."
    },
    "liveCallable": {
      "type": "boolean",
      "description": "True returns agents with a reusable invocation policy and an owner-verified, currently reachable endpoint."
    },
    "maxCostPerRunUsd": {
      "type": "number",
      "minimum": 0,
      "description": "Upper bound on lifetime total_cost_usd / total_runs, i.e. average observed cost per run."
    },
    "maxP50LatencyMs": {
      "type": "integer",
      "minimum": 0,
      "description": "Upper bound on the agent's observed p50 latency in milliseconds."
    },
    "sort": {
      "type": "string",
      "enum": [
        "verified",
        "recent",
        "runs"
      ],
      "default": "verified"
    },
    "limit": {
      "type": "integer",
      "minimum": 1,
      "maximum": 20,
      "default": 10
    }
  },
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢get_agent(handle)

Get one public agent's recipe, public capability manifest, coarse invocation status, owner telemetry, and verified benchmark submissions.

입력 스키마

{
  "type": "object",
  "properties": {
    "handle": {
      "type": "string",
      "maxLength": 101,
      "description": "Bench handle in @owner/agent-slug form."
    }
  },
  "required": [
    "handle"
  ],
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢list_benchmarks

List public, versioned benchmark contracts and only their trusted-runner-verified submissions.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "http://json-schema.org/draft-07/schema#"
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 3개
검증됨버전이 기록되지 않음도구 3개
검증됨버전이 기록되지 않음도구 3개