operant-mcp
Read-only MCP server for the OPERANT AI operating-agent calibration benchmark.
¿Debería usar esto?
Calidad y seguridad
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"operant-mcp": {
"command": "npx",
"args": [
"saagar-operant-mcp"
]
}
}
}Paquetes ejecutables
0.1.0stdioPuntos de conexión remotos
https://operant-mcp.saagar210.workers.dev/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (5)
🟢get_results
Return retained model calculation profiles plus generated_at, claim_status, claims_at_risk, and the bounded evidence_binding. These rows are not durable named-model performance claims. Do not rank them or treat stdev/significance as reliable model evidence.
Esquema de entrada
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢compare_models(model_a, model_b)
Place two retained calculation profiles side by side by display_name substring. Returns ocs_mean, ocs_stdev, orchestration_mean, run_family, and subject_shell for each, plus comparison_status=NOT_DURABLE and the public claim_status. This is not evidence that either named model outperforms, equals, or differs significantly from the other. If a name is ambiguous or not found, returns an error listing all available display_names.
Esquema de entrada
{
"type": "object",
"properties": {
"model_a": {
"type": "string",
"minLength": 1,
"description": "Display name (or substring) of the first model"
},
"model_b": {
"type": "string",
"minLength": 1,
"description": "Display name (or substring) of the second model"
}
},
"required": [
"model_a",
"model_b"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_methodology
Return the benchmark design: the 4 axes, the OCS formula (Youden's J), the 5 decision labels (PROCEED, PROCEED_SANCTIONED, REFUSE, ESCALATE, REROUTE), scoring blocks, and a concise what-it-measures / what-it-doesn't summary. Sourced from the baked case-file metadata.
Esquema de entrada
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_cases(axis)
Return case metadata (no full task prompts): pair_id/id, axis, tier, grounding, and side indicators (malign/benign for axes 1/2/4; null for axis 3). Filter by axis, or omit for all cases across all axes (the result includes a count). Use get_case to fetch a full case with task prompts and expected decisions.
Esquema de entrada
{
"type": "object",
"properties": {
"axis": {
"type": "string",
"enum": [
"refusal-calibration",
"sanctioned-path",
"orchestration",
"escalation-reroute"
],
"description": "Axis to filter by: refusal-calibration | sanctioned-path | orchestration | escalation-reroute"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_case(pair_id, axis)
Return the full case for a given pair_id (axes 1/2/4) or id (axis 3): malign and benign task prompts, expected decisions, grounding rationale, and bypass patterns. Axis 3 cases are single (unmatched) and use an 'id' field instead of 'pair_id'. Use list_cases to browse available ids.
Esquema de entrada
{
"type": "object",
"properties": {
"pair_id": {
"type": "string",
"minLength": 1,
"description": "The pair_id (axes 1/2/4) or id (axis 3) to retrieve"
},
"axis": {
"type": "string",
"enum": [
"refusal-calibration",
"sanctioned-path",
"orchestration",
"escalation-reroute"
],
"description": "The axis this case belongs to"
}
},
"required": [
"pair_id",
"axis"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}Comunidad
Evidencia