operant-mcp
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
사용해야 할까요
품질 및 안전성
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"operant-mcp": {
"command": "npx",
"args": [
"saagar-operant-mcp"
]
}
}
}실행 가능한 패키지
0.1.0stdio원격 엔드포인트
https://operant-mcp.saagar210.workers.dev/mcpstreamable-http할 수 있는 일
도구 목록
도구 (5)
🟢get_results
Return retained model calculation profiles plus generated_at, claim_status, claims_at_risk, and the bounded evidence_binding. These rows are not durable named-model performance claims. Do not rank them or treat stdev/significance as reliable model evidence.
입력 스키마
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢compare_models(model_a, model_b)
Place two retained calculation profiles side by side by display_name substring. Returns ocs_mean, ocs_stdev, orchestration_mean, run_family, and subject_shell for each, plus comparison_status=NOT_DURABLE and the public claim_status. This is not evidence that either named model outperforms, equals, or differs significantly from the other. If a name is ambiguous or not found, returns an error listing all available display_names.
입력 스키마
{
"type": "object",
"properties": {
"model_a": {
"type": "string",
"minLength": 1,
"description": "Display name (or substring) of the first model"
},
"model_b": {
"type": "string",
"minLength": 1,
"description": "Display name (or substring) of the second model"
}
},
"required": [
"model_a",
"model_b"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_methodology
Return the benchmark design: the 4 axes, the OCS formula (Youden's J), the 5 decision labels (PROCEED, PROCEED_SANCTIONED, REFUSE, ESCALATE, REROUTE), scoring blocks, and a concise what-it-measures / what-it-doesn't summary. Sourced from the baked case-file metadata.
입력 스키마
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_cases(axis)
Return case metadata (no full task prompts): pair_id/id, axis, tier, grounding, and side indicators (malign/benign for axes 1/2/4; null for axis 3). Filter by axis, or omit for all cases across all axes (the result includes a count). Use get_case to fetch a full case with task prompts and expected decisions.
입력 스키마
{
"type": "object",
"properties": {
"axis": {
"type": "string",
"enum": [
"refusal-calibration",
"sanctioned-path",
"orchestration",
"escalation-reroute"
],
"description": "Axis to filter by: refusal-calibration | sanctioned-path | orchestration | escalation-reroute"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_case(pair_id, axis)
Return the full case for a given pair_id (axes 1/2/4) or id (axis 3): malign and benign task prompts, expected decisions, grounding rationale, and bypass patterns. Axis 3 cases are single (unmatched) and use an 'id' field instead of 'pair_id'. Use list_cases to browse available ids.
입력 스키마
{
"type": "object",
"properties": {
"pair_id": {
"type": "string",
"minLength": 1,
"description": "The pair_id (axes 1/2/4) or id (axis 3) to retrieve"
},
"axis": {
"type": "string",
"enum": [
"refusal-calibration",
"sanctioned-path",
"orchestration",
"escalation-reroute"
],
"description": "The axis this case belongs to"
}
},
"required": [
"pair_id",
"axis"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}커뮤니티
증거