Whetstone
Verifier-grounded AI promotion gates, disposable report cards, and signed PASS/HOLD/BLOCK receipts.
使うべきか
品質と安全性
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"tools": {
"url": "https://whetstone.cyberelf.link/mcp"
}
}
}リモートエンドポイント
https://whetstone.cyberelf.link/mcpstreamable-httpできること
ツール一覧
ツール(14)
🟢inspect_promotion(baseline, baseline_name, candidate, candidate_name, domains, ...)
Quarantine declared exposure, compare paired baseline/candidate outcomes on the clean remainder, and issue a promotion receipt. Bring your own exam rows, exposure records, and per-item results. Full example: GET /api/examples key 'inspector'.
入力スキーマ
{
"type": "object",
"properties": {
"baseline": {
"additionalProperties": {
"type": "boolean"
},
"description": "item_id -> boolean pass/fail result for the baseline system.",
"maxProperties": 5000,
"minProperties": 1,
"type": "object"
},
"baseline_name": {
"default": "baseline",
"type": "string"
},
"candidate": {
"additionalProperties": {
"type": "boolean"
},
"description": "item_id -> boolean pass/fail result for the candidate system.",
"maxProperties": 5000,
"minProperties": 1,
"type": "object"
},
"candidate_name": {
"default": "candidate",
"type": "string"
},
"domains": {
"additionalProperties": {
"type": "string"
},
"description": "Optional item_id -> domain label mapping.",
"type": "object"
},
"enable_behavioral_fingerprint": {
"default": true,
"type": "boolean"
},
"enable_text_similarity": {
"default": true,
"type": "boolean"
},
"exam": {
"description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
"items": {
"additionalProperties": true,
"properties": {},
"type": "object"
},
"maxItems": 5000,
"minItems": 1,
"type": "array"
},
"exposure": {
"description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
"items": {
"additionalProperties": true,
"properties": {},
"type": "object"
},
"maxItems": 5000,
"minItems": 0,
"type": "array"
},
"fingerprint_max_n": {
"default": 4,
"maximum": 5,
"minimum": 3,
"type": "integer"
},
"policy": {
"additionalProperties": false,
"description": "Explicit promotion policy. Omitted fields use the documented defaults.",
"properties": {
"confidence_alpha": {
"default": 0.05,
"exclusiveMinimum": 0,
"maximum": 1,
"type": "number"
},
"max_regressions": {
"default": 0,
"minimum": 0,
"type": "integer"
},
"min_gains": {
"default": 1,
"minimum": 0,
"type": "integer"
},
"require_retained_probe": {
"default": false,
"type": "boolean"
}
},
"type": "object"
},
"retained_probe": {
"additionalProperties": false,
"description": "Optional retained-capability result checked alongside the paired cohort.",
"properties": {
"base_verified": {
"minimum": 0,
"type": "integer"
},
"candidate_verified": {
"minimum": 0,
"type": "integer"
},
"items": {
"minimum": 0,
"type": "integer"
}
},
"required": [
"base_verified",
"candidate_verified",
"items"
],
"type": "object"
},
"similarity_threshold": {
"default": 0.6,
"maximum": 1,
"minimum": 0.5,
"type": "number"
}
},
"required": [
"exam",
"baseline",
"candidate"
],
"additionalProperties": false,
"description": "Audit exposure, prove a complete clean cohort, then gate baseline versus candidate."
}🟢audit_leakage(enable_behavioral_fingerprint, enable_text_similarity, exam, exposure, fingerprint_max_n, ...)
Exact declared-exposure audit over your exam rows: row identity, behavioral fingerprints for graph-DSL expressions, text-similarity review flags, and a clean exam export. Full example: GET /api/examples key 'leakage'.
入力スキーマ
{
"type": "object",
"properties": {
"enable_behavioral_fingerprint": {
"default": true,
"type": "boolean"
},
"enable_text_similarity": {
"default": true,
"type": "boolean"
},
"exam": {
"description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
"items": {
"additionalProperties": true,
"properties": {},
"type": "object"
},
"maxItems": 5000,
"minItems": 1,
"type": "array"
},
"exposure": {
"description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
"items": {
"additionalProperties": true,
"properties": {},
"type": "object"
},
"maxItems": 5000,
"minItems": 0,
"type": "array"
},
"fingerprint_max_n": {
"default": 4,
"maximum": 5,
"minimum": 3,
"type": "integer"
},
"similarity_threshold": {
"default": 0.6,
"maximum": 1,
"minimum": 0.5,
"type": "number"
}
},
"required": [
"exam"
],
"additionalProperties": false,
"description": "Audit declared exposure against an exam and export the clean remainder."
}🟢promotion_gate(baseline, baseline_name, candidate, candidate_name, domains, ...)
PASS, HOLD, or BLOCK from paired per-item results: gains, regressions, exact McNemar p-value, per-domain breakdown. Full example: GET /api/examples key 'gate'.
入力スキーマ
{
"type": "object",
"properties": {
"baseline": {
"additionalProperties": {
"type": "boolean"
},
"description": "item_id -> boolean pass/fail result for the baseline system.",
"maxProperties": 5000,
"minProperties": 1,
"type": "object"
},
"baseline_name": {
"default": "baseline",
"type": "string"
},
"candidate": {
"additionalProperties": {
"type": "boolean"
},
"description": "item_id -> boolean pass/fail result for the candidate system.",
"maxProperties": 5000,
"minProperties": 1,
"type": "object"
},
"candidate_name": {
"default": "candidate",
"type": "string"
},
"domains": {
"additionalProperties": {
"type": "string"
},
"description": "Optional item_id -> domain label mapping.",
"type": "object"
},
"policy": {
"additionalProperties": false,
"description": "Explicit promotion policy. Omitted fields use the documented defaults.",
"properties": {
"confidence_alpha": {
"default": 0.05,
"exclusiveMinimum": 0,
"maximum": 1,
"type": "number"
},
"max_regressions": {
"default": 0,
"minimum": 0,
"type": "integer"
},
"min_gains": {
"default": 1,
"minimum": 0,
"type": "integer"
},
"require_retained_probe": {
"default": false,
"type": "boolean"
}
},
"type": "object"
},
"retained_probe": {
"additionalProperties": false,
"description": "Optional retained-capability result checked alongside the paired cohort.",
"properties": {
"base_verified": {
"minimum": 0,
"type": "integer"
},
"candidate_verified": {
"minimum": 0,
"type": "integer"
},
"items": {
"minimum": 0,
"type": "integer"
}
},
"required": [
"base_verified",
"candidate_verified",
"items"
],
"type": "object"
}
},
"required": [
"baseline",
"candidate"
],
"additionalProperties": false,
"description": "Compare identical baseline and candidate item cohorts under an explicit policy."
}🟢bank_health(history, items)
Item-lifecycle diagnostics over your grading history: discriminators, saturated and flaky items, frontier gaps. Full example: GET /api/examples key 'health'.
入力スキーマ
{
"type": "object",
"properties": {
"history": {
"description": "Observed item/system outcomes.",
"items": {
"additionalProperties": true,
"properties": {
"domain": {
"type": "string"
},
"item_id": {
"minLength": 1,
"type": "string"
},
"passed": {
"type": "boolean"
},
"system": {
"minLength": 1,
"type": "string"
}
},
"required": [
"item_id",
"system",
"passed"
],
"type": "object"
},
"maxItems": 5000,
"minItems": 1,
"type": "array"
},
"items": {
"description": "Optional item definitions.",
"items": {
"additionalProperties": true,
"properties": {
"domain": {
"type": "string"
},
"item_id": {
"minLength": 1,
"type": "string"
}
},
"required": [
"item_id"
],
"type": "object"
},
"maxItems": 5000,
"minItems": 0,
"type": "array"
}
},
"required": [
"history"
],
"additionalProperties": false,
"description": "Diagnose item lifecycle health from one or more grading-history rows."
}🟡safe_patch(document, operations, reason)
Apply a section-scoped Markdown patch under conservation checks (untouched sections stay byte-identical; protected tokens preserved). Full example: GET /api/examples key 'safepatch'.
入力スキーマ
{
"type": "object",
"properties": {
"document": {
"description": "Complete Markdown document to patch.",
"maxLength": 200000,
"minLength": 1,
"type": "string"
},
"operations": {
"items": {
"additionalProperties": false,
"properties": {
"allow_token_changes": {
"description": "Protected literal tokens that this operation may intentionally change.",
"items": {
"type": "string"
},
"maxItems": 100,
"type": "array"
},
"find": {
"minLength": 1,
"type": "string"
},
"replace": {
"type": "string"
},
"target_heading": {
"description": "Markdown heading text without the leading # characters.",
"minLength": 1,
"type": "string"
}
},
"required": [
"target_heading",
"find",
"replace"
],
"type": "object"
},
"maxItems": 50,
"minItems": 1,
"type": "array"
},
"reason": {
"type": "string"
}
},
"required": [
"document",
"operations"
],
"additionalProperties": false,
"description": "Apply deterministic, section-scoped Markdown replacements under conservation checks."
}🟢counterexample_hunt(expression, ns, restarts, seed, steps)
Bounded simulated-annealing search for a graph counterexample inside a DSL predicate class, with an exact certificate when found. CPU-bounded and strictly rate-limited. Full example: GET /api/examples key 'counterexample'.
入力スキーマ
{
"type": "object",
"properties": {
"expression": {
"description": "Graph predicate in the Whetstone DSL, for example: is_connected and is_triangle_free and not is_bipartite",
"maxLength": 500,
"minLength": 1,
"type": "string"
},
"ns": {
"default": [
8,
9,
10,
11
],
"description": "Graph sizes searched.",
"items": {
"maximum": 12,
"minimum": 4,
"type": "integer"
},
"maxItems": 5,
"minItems": 1,
"type": "array"
},
"restarts": {
"default": 4,
"maximum": 6,
"minimum": 1,
"type": "integer"
},
"seed": {
"default": 0,
"type": "integer"
},
"steps": {
"default": 800,
"maximum": 1500,
"minimum": 50,
"type": "integer"
}
},
"required": [
"expression"
],
"additionalProperties": false,
"description": "Run a bounded graph search against one Whetstone predicate expression."
}🟡memory_relevance(context_entities, current_step, memories, objective, objective_entities, ...)
Compare query-free salience against objective-conditioned relevance for a set of memories under a token budget. Full example: GET /api/examples key 'memory'.
入力スキーマ
{
"type": "object",
"properties": {
"context_entities": {
"description": "Entities already active in context.",
"items": {
"type": "string"
},
"maxItems": 100,
"type": "array"
},
"current_step": {
"minimum": 0,
"type": "integer"
},
"memories": {
"description": "Memories to rank.",
"items": {
"additionalProperties": false,
"properties": {
"age": {
"default": 0,
"minimum": 0,
"type": "integer"
},
"confidence": {
"default": 0.8,
"maximum": 1,
"minimum": 0,
"type": "number"
},
"content": {
"minLength": 1,
"type": "string"
},
"entities": {
"description": "Entities explicitly present in this memory.",
"items": {
"type": "string"
},
"maxItems": 100,
"type": "array"
},
"kind": {
"default": "episodic",
"type": "string"
},
"source": {
"default": "uploaded",
"type": "string"
},
"use_count": {
"default": 0,
"minimum": 0,
"type": "integer"
}
},
"required": [
"content"
],
"type": "object"
},
"maxItems": 1000,
"minItems": 1,
"type": "array"
},
"objective": {
"minLength": 1,
"type": "string"
},
"objective_entities": {
"description": "Optional explicit entities when the objective text is not self-describing.",
"items": {
"type": "string"
},
"maxItems": 100,
"type": "array"
},
"question_kind": {
"default": "generic",
"type": "string"
},
"token_budget": {
"default": 90,
"maximum": 10000,
"minimum": 1,
"type": "integer"
}
},
"required": [
"objective",
"memories"
],
"additionalProperties": false,
"description": "Rank caller-supplied memories against a concrete objective under a token budget."
}🟢replay_trace(events, notes)
Turn reasoning-emulator control events into checkpoints, rewinds, notes, and a timeline. Full example: GET /api/examples key 'replay'.
入力スキーマ
{
"type": "object",
"properties": {
"events": {
"description": "Ordered reasoning-emulator events.",
"items": {
"additionalProperties": false,
"properties": {
"detail": {
"type": "string"
},
"kind": {
"description": "Event class such as control, verifier, model, or observation.",
"type": "string"
},
"source": {
"default": "native",
"type": "string"
},
"step": {
"minimum": 0,
"type": "integer"
}
},
"required": [
"kind",
"detail"
],
"type": "object"
},
"maxItems": 5000,
"minItems": 1,
"type": "array"
},
"notes": {
"description": "Optional analyst notes.",
"items": {
"type": "string"
},
"maxItems": 5000,
"type": "array"
}
},
"required": [
"events"
],
"additionalProperties": false,
"description": "Reconstruct checkpoints, rewinds, branches, and verifier outcomes from control events."
}🟡report_card_start(challenge)
TIER 1: start a disposable report-card session. Returns exam items (graph-repair prompts minted from the repository's public frontier) for THIS agent to answer. Answer every item, then call report_card_submit exactly once. Sessions are one-shot, expire in 15 minutes, and are strictly rate-limited. This demonstrates the promotion-gate mechanism on disposable items; it is not a private-bank credential.
入力スキーマ
{
"type": "object",
"properties": {
"challenge": {
"description": "Caller nonce bound into the signed receipt for replay detection.",
"maxLength": 128,
"minLength": 8,
"type": "string"
}
},
"additionalProperties": false
}🟡report_card_submit(answers, session_id)
TIER 1: submit answers for a report-card session and receive the graded report (per-item verdicts, per-domain totals, SHA-256 commitments). Grading is by checker spec: verified strict refinements are reported separately, and promotion grade requires at least 5% clean-support retention. No answer key exists. The session is destroyed by this call.
入力スキーマ
{
"type": "object",
"properties": {
"answers": {
"additionalProperties": {
"type": "string"
},
"description": "item_id -> answer (a DSL predicate, or the JSON reply the prompt asked for)",
"type": "object"
},
"session_id": {
"type": "string"
}
},
"required": [
"session_id",
"answers"
],
"additionalProperties": false
}🟡open_bench_start(challenge)
TIER 2: start a one-shot Open Promotion Bench session. Returns six fresh virtual-repository scope-integrity tasks. Run a baseline and candidate independently on the same cohort, then submit both answer maps with open_bench_submit. This is an open, procedural, self-attested track rather than a private-bank credential.
入力スキーマ
{
"type": "object",
"properties": {
"challenge": {
"description": "Caller nonce bound into the signed receipt for replay detection.",
"maxLength": 128,
"minLength": 8,
"type": "string"
}
},
"additionalProperties": false
}🟡open_bench_submit(attestation, baseline_answers, baseline_manifest, candidate_answers, candidate_manifest, ...)
TIER 2: grade paired baseline and candidate patches, count gains/regressions/ties, and issue PASS/HOLD/BLOCK. Set publish=true plus attestation=true to append only the safe manifests and sanitized receipt to the public board; tasks and answers are never persisted.
入力スキーマ
{
"type": "object",
"properties": {
"attestation": {
"type": "boolean"
},
"baseline_answers": {
"type": "object"
},
"baseline_manifest": {
"additionalProperties": false,
"properties": {
"harness": {
"type": "string"
},
"model": {
"type": "string"
},
"name": {
"type": "string"
},
"version": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"candidate_answers": {
"type": "object"
},
"candidate_manifest": {
"additionalProperties": false,
"properties": {
"harness": {
"type": "string"
},
"model": {
"type": "string"
},
"name": {
"type": "string"
},
"version": {
"type": "string"
}
},
"required": [
"name"
],
"type": "object"
},
"publish": {
"type": "boolean"
},
"session_id": {
"type": "string"
}
},
"required": [
"session_id",
"baseline_manifest",
"candidate_manifest",
"baseline_answers",
"candidate_answers"
],
"additionalProperties": false
}🟢open_bench_leaderboard
TIER 2: list the self-attested public Open Promotion Bench receipts. Entries contain manifests, verdicts, item-level transitions, and commitments but never task contents or submitted answers.
入力スキーマ
{
"type": "object",
"properties": {},
"additionalProperties": false
}⚪about_whetstone
What this service is: the tool catalog, the tier boundaries, and where the source lives.
入力スキーマ
{
"type": "object",
"properties": {},
"additionalProperties": false
}コミュニティ
エビデンス