Whetstone

Verifier-grounded AI promotion gates, disposable report cards, and signed PASS/HOLD/BLOCK receipts.

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
82%
Calidad de los nombres
89%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~3,212Tokens (definiciones de herramientas)
~2.3 KBTamaño de respuesta típico
Impacto significativo en la atención (2.51% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "tools": {
      "url": "https://whetstone.cyberelf.link/mcp"
    }
  }
}

Puntos de conexión remotos

https://whetstone.cyberelf.link/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (14)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟢inspect_promotion(baseline, baseline_name, candidate, candidate_name, domains, ...)

Quarantine declared exposure, compare paired baseline/candidate outcomes on the clean remainder, and issue a promotion receipt. Bring your own exam rows, exposure records, and per-item results. Full example: GET /api/examples key 'inspector'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "baseline": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the baseline system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "baseline_name": {
      "default": "baseline",
      "type": "string"
    },
    "candidate": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the candidate system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "candidate_name": {
      "default": "candidate",
      "type": "string"
    },
    "domains": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "Optional item_id -> domain label mapping.",
      "type": "object"
    },
    "enable_behavioral_fingerprint": {
      "default": true,
      "type": "boolean"
    },
    "enable_text_similarity": {
      "default": true,
      "type": "boolean"
    },
    "exam": {
      "description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "exposure": {
      "description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    },
    "fingerprint_max_n": {
      "default": 4,
      "maximum": 5,
      "minimum": 3,
      "type": "integer"
    },
    "policy": {
      "additionalProperties": false,
      "description": "Explicit promotion policy. Omitted fields use the documented defaults.",
      "properties": {
        "confidence_alpha": {
          "default": 0.05,
          "exclusiveMinimum": 0,
          "maximum": 1,
          "type": "number"
        },
        "max_regressions": {
          "default": 0,
          "minimum": 0,
          "type": "integer"
        },
        "min_gains": {
          "default": 1,
          "minimum": 0,
          "type": "integer"
        },
        "require_retained_probe": {
          "default": false,
          "type": "boolean"
        }
      },
      "type": "object"
    },
    "retained_probe": {
      "additionalProperties": false,
      "description": "Optional retained-capability result checked alongside the paired cohort.",
      "properties": {
        "base_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "candidate_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "items": {
          "minimum": 0,
          "type": "integer"
        }
      },
      "required": [
        "base_verified",
        "candidate_verified",
        "items"
      ],
      "type": "object"
    },
    "similarity_threshold": {
      "default": 0.6,
      "maximum": 1,
      "minimum": 0.5,
      "type": "number"
    }
  },
  "required": [
    "exam",
    "baseline",
    "candidate"
  ],
  "additionalProperties": false,
  "description": "Audit exposure, prove a complete clean cohort, then gate baseline versus candidate."
}
🟢audit_leakage(enable_behavioral_fingerprint, enable_text_similarity, exam, exposure, fingerprint_max_n, ...)

Exact declared-exposure audit over your exam rows: row identity, behavioral fingerprints for graph-DSL expressions, text-similarity review flags, and a clean exam export. Full example: GET /api/examples key 'leakage'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "enable_behavioral_fingerprint": {
      "default": true,
      "type": "boolean"
    },
    "enable_text_similarity": {
      "default": true,
      "type": "boolean"
    },
    "exam": {
      "description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "exposure": {
      "description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    },
    "fingerprint_max_n": {
      "default": 4,
      "maximum": 5,
      "minimum": 3,
      "type": "integer"
    },
    "similarity_threshold": {
      "default": 0.6,
      "maximum": 1,
      "minimum": 0.5,
      "type": "number"
    }
  },
  "required": [
    "exam"
  ],
  "additionalProperties": false,
  "description": "Audit declared exposure against an exam and export the clean remainder."
}
🟢promotion_gate(baseline, baseline_name, candidate, candidate_name, domains, ...)

PASS, HOLD, or BLOCK from paired per-item results: gains, regressions, exact McNemar p-value, per-domain breakdown. Full example: GET /api/examples key 'gate'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "baseline": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the baseline system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "baseline_name": {
      "default": "baseline",
      "type": "string"
    },
    "candidate": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the candidate system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "candidate_name": {
      "default": "candidate",
      "type": "string"
    },
    "domains": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "Optional item_id -> domain label mapping.",
      "type": "object"
    },
    "policy": {
      "additionalProperties": false,
      "description": "Explicit promotion policy. Omitted fields use the documented defaults.",
      "properties": {
        "confidence_alpha": {
          "default": 0.05,
          "exclusiveMinimum": 0,
          "maximum": 1,
          "type": "number"
        },
        "max_regressions": {
          "default": 0,
          "minimum": 0,
          "type": "integer"
        },
        "min_gains": {
          "default": 1,
          "minimum": 0,
          "type": "integer"
        },
        "require_retained_probe": {
          "default": false,
          "type": "boolean"
        }
      },
      "type": "object"
    },
    "retained_probe": {
      "additionalProperties": false,
      "description": "Optional retained-capability result checked alongside the paired cohort.",
      "properties": {
        "base_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "candidate_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "items": {
          "minimum": 0,
          "type": "integer"
        }
      },
      "required": [
        "base_verified",
        "candidate_verified",
        "items"
      ],
      "type": "object"
    }
  },
  "required": [
    "baseline",
    "candidate"
  ],
  "additionalProperties": false,
  "description": "Compare identical baseline and candidate item cohorts under an explicit policy."
}
🟢bank_health(history, items)

Item-lifecycle diagnostics over your grading history: discriminators, saturated and flaky items, frontier gaps. Full example: GET /api/examples key 'health'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "history": {
      "description": "Observed item/system outcomes.",
      "items": {
        "additionalProperties": true,
        "properties": {
          "domain": {
            "type": "string"
          },
          "item_id": {
            "minLength": 1,
            "type": "string"
          },
          "passed": {
            "type": "boolean"
          },
          "system": {
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "item_id",
          "system",
          "passed"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "items": {
      "description": "Optional item definitions.",
      "items": {
        "additionalProperties": true,
        "properties": {
          "domain": {
            "type": "string"
          },
          "item_id": {
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "item_id"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    }
  },
  "required": [
    "history"
  ],
  "additionalProperties": false,
  "description": "Diagnose item lifecycle health from one or more grading-history rows."
}
🟡safe_patch(document, operations, reason)

Apply a section-scoped Markdown patch under conservation checks (untouched sections stay byte-identical; protected tokens preserved). Full example: GET /api/examples key 'safepatch'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "document": {
      "description": "Complete Markdown document to patch.",
      "maxLength": 200000,
      "minLength": 1,
      "type": "string"
    },
    "operations": {
      "items": {
        "additionalProperties": false,
        "properties": {
          "allow_token_changes": {
            "description": "Protected literal tokens that this operation may intentionally change.",
            "items": {
              "type": "string"
            },
            "maxItems": 100,
            "type": "array"
          },
          "find": {
            "minLength": 1,
            "type": "string"
          },
          "replace": {
            "type": "string"
          },
          "target_heading": {
            "description": "Markdown heading text without the leading # characters.",
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "target_heading",
          "find",
          "replace"
        ],
        "type": "object"
      },
      "maxItems": 50,
      "minItems": 1,
      "type": "array"
    },
    "reason": {
      "type": "string"
    }
  },
  "required": [
    "document",
    "operations"
  ],
  "additionalProperties": false,
  "description": "Apply deterministic, section-scoped Markdown replacements under conservation checks."
}
🟢counterexample_hunt(expression, ns, restarts, seed, steps)

Bounded simulated-annealing search for a graph counterexample inside a DSL predicate class, with an exact certificate when found. CPU-bounded and strictly rate-limited. Full example: GET /api/examples key 'counterexample'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "expression": {
      "description": "Graph predicate in the Whetstone DSL, for example: is_connected and is_triangle_free and not is_bipartite",
      "maxLength": 500,
      "minLength": 1,
      "type": "string"
    },
    "ns": {
      "default": [
        8,
        9,
        10,
        11
      ],
      "description": "Graph sizes searched.",
      "items": {
        "maximum": 12,
        "minimum": 4,
        "type": "integer"
      },
      "maxItems": 5,
      "minItems": 1,
      "type": "array"
    },
    "restarts": {
      "default": 4,
      "maximum": 6,
      "minimum": 1,
      "type": "integer"
    },
    "seed": {
      "default": 0,
      "type": "integer"
    },
    "steps": {
      "default": 800,
      "maximum": 1500,
      "minimum": 50,
      "type": "integer"
    }
  },
  "required": [
    "expression"
  ],
  "additionalProperties": false,
  "description": "Run a bounded graph search against one Whetstone predicate expression."
}
🟡memory_relevance(context_entities, current_step, memories, objective, objective_entities, ...)

Compare query-free salience against objective-conditioned relevance for a set of memories under a token budget. Full example: GET /api/examples key 'memory'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "context_entities": {
      "description": "Entities already active in context.",
      "items": {
        "type": "string"
      },
      "maxItems": 100,
      "type": "array"
    },
    "current_step": {
      "minimum": 0,
      "type": "integer"
    },
    "memories": {
      "description": "Memories to rank.",
      "items": {
        "additionalProperties": false,
        "properties": {
          "age": {
            "default": 0,
            "minimum": 0,
            "type": "integer"
          },
          "confidence": {
            "default": 0.8,
            "maximum": 1,
            "minimum": 0,
            "type": "number"
          },
          "content": {
            "minLength": 1,
            "type": "string"
          },
          "entities": {
            "description": "Entities explicitly present in this memory.",
            "items": {
              "type": "string"
            },
            "maxItems": 100,
            "type": "array"
          },
          "kind": {
            "default": "episodic",
            "type": "string"
          },
          "source": {
            "default": "uploaded",
            "type": "string"
          },
          "use_count": {
            "default": 0,
            "minimum": 0,
            "type": "integer"
          }
        },
        "required": [
          "content"
        ],
        "type": "object"
      },
      "maxItems": 1000,
      "minItems": 1,
      "type": "array"
    },
    "objective": {
      "minLength": 1,
      "type": "string"
    },
    "objective_entities": {
      "description": "Optional explicit entities when the objective text is not self-describing.",
      "items": {
        "type": "string"
      },
      "maxItems": 100,
      "type": "array"
    },
    "question_kind": {
      "default": "generic",
      "type": "string"
    },
    "token_budget": {
      "default": 90,
      "maximum": 10000,
      "minimum": 1,
      "type": "integer"
    }
  },
  "required": [
    "objective",
    "memories"
  ],
  "additionalProperties": false,
  "description": "Rank caller-supplied memories against a concrete objective under a token budget."
}
🟢replay_trace(events, notes)

Turn reasoning-emulator control events into checkpoints, rewinds, notes, and a timeline. Full example: GET /api/examples key 'replay'.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "events": {
      "description": "Ordered reasoning-emulator events.",
      "items": {
        "additionalProperties": false,
        "properties": {
          "detail": {
            "type": "string"
          },
          "kind": {
            "description": "Event class such as control, verifier, model, or observation.",
            "type": "string"
          },
          "source": {
            "default": "native",
            "type": "string"
          },
          "step": {
            "minimum": 0,
            "type": "integer"
          }
        },
        "required": [
          "kind",
          "detail"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "notes": {
      "description": "Optional analyst notes.",
      "items": {
        "type": "string"
      },
      "maxItems": 5000,
      "type": "array"
    }
  },
  "required": [
    "events"
  ],
  "additionalProperties": false,
  "description": "Reconstruct checkpoints, rewinds, branches, and verifier outcomes from control events."
}
🟡report_card_start(challenge)

TIER 1: start a disposable report-card session. Returns exam items (graph-repair prompts minted from the repository's public frontier) for THIS agent to answer. Answer every item, then call report_card_submit exactly once. Sessions are one-shot, expire in 15 minutes, and are strictly rate-limited. This demonstrates the promotion-gate mechanism on disposable items; it is not a private-bank credential.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "challenge": {
      "description": "Caller nonce bound into the signed receipt for replay detection.",
      "maxLength": 128,
      "minLength": 8,
      "type": "string"
    }
  },
  "additionalProperties": false
}
🟡report_card_submit(answers, session_id)

TIER 1: submit answers for a report-card session and receive the graded report (per-item verdicts, per-domain totals, SHA-256 commitments). Grading is by checker spec: verified strict refinements are reported separately, and promotion grade requires at least 5% clean-support retention. No answer key exists. The session is destroyed by this call.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "answers": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "item_id -> answer (a DSL predicate, or the JSON reply the prompt asked for)",
      "type": "object"
    },
    "session_id": {
      "type": "string"
    }
  },
  "required": [
    "session_id",
    "answers"
  ],
  "additionalProperties": false
}
🟡open_bench_start(challenge)

TIER 2: start a one-shot Open Promotion Bench session. Returns six fresh virtual-repository scope-integrity tasks. Run a baseline and candidate independently on the same cohort, then submit both answer maps with open_bench_submit. This is an open, procedural, self-attested track rather than a private-bank credential.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "challenge": {
      "description": "Caller nonce bound into the signed receipt for replay detection.",
      "maxLength": 128,
      "minLength": 8,
      "type": "string"
    }
  },
  "additionalProperties": false
}
🟡open_bench_submit(attestation, baseline_answers, baseline_manifest, candidate_answers, candidate_manifest, ...)

TIER 2: grade paired baseline and candidate patches, count gains/regressions/ties, and issue PASS/HOLD/BLOCK. Set publish=true plus attestation=true to append only the safe manifests and sanitized receipt to the public board; tasks and answers are never persisted.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "attestation": {
      "type": "boolean"
    },
    "baseline_answers": {
      "type": "object"
    },
    "baseline_manifest": {
      "additionalProperties": false,
      "properties": {
        "harness": {
          "type": "string"
        },
        "model": {
          "type": "string"
        },
        "name": {
          "type": "string"
        },
        "version": {
          "type": "string"
        }
      },
      "required": [
        "name"
      ],
      "type": "object"
    },
    "candidate_answers": {
      "type": "object"
    },
    "candidate_manifest": {
      "additionalProperties": false,
      "properties": {
        "harness": {
          "type": "string"
        },
        "model": {
          "type": "string"
        },
        "name": {
          "type": "string"
        },
        "version": {
          "type": "string"
        }
      },
      "required": [
        "name"
      ],
      "type": "object"
    },
    "publish": {
      "type": "boolean"
    },
    "session_id": {
      "type": "string"
    }
  },
  "required": [
    "session_id",
    "baseline_manifest",
    "candidate_manifest",
    "baseline_answers",
    "candidate_answers"
  ],
  "additionalProperties": false
}
🟢open_bench_leaderboard

TIER 2: list the self-attested public Open Promotion Bench receipts. Entries contain manifests, verdicts, item-level transitions, and commitments but never task contents or submitted answers.

Esquema de entrada

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
⚪about_whetstone

What this service is: the tool catalog, the tier boundaries, and where the source lives.

Esquema de entrada

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada14 herramientas
verificadoversión no registrada14 herramientas