Whetstone

Verifier-grounded AI promotion gates, disposable report cards, and signed PASS/HOLD/BLOCK receipts.

我该使用它吗

质量与安全性

A
描述质量
100%
模式完整度
82%
命名质量
89%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~3,212token 数(工具定义)
~2.3 KB典型响应大小
对注意力有显著影响(占 128k 上下文窗口的 2.51%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "tools": {
      "url": "https://whetstone.cyberelf.link/mcp"
    }
  }
}

远程端点

https://whetstone.cyberelf.link/mcpstreamable-http

它能做什么

工具清单

工具(14)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢inspect_promotion(baseline, baseline_name, candidate, candidate_name, domains, ...)

Quarantine declared exposure, compare paired baseline/candidate outcomes on the clean remainder, and issue a promotion receipt. Bring your own exam rows, exposure records, and per-item results. Full example: GET /api/examples key 'inspector'.

输入模式

{
  "type": "object",
  "properties": {
    "baseline": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the baseline system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "baseline_name": {
      "default": "baseline",
      "type": "string"
    },
    "candidate": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the candidate system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "candidate_name": {
      "default": "candidate",
      "type": "string"
    },
    "domains": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "Optional item_id -> domain label mapping.",
      "type": "object"
    },
    "enable_behavioral_fingerprint": {
      "default": true,
      "type": "boolean"
    },
    "enable_text_similarity": {
      "default": true,
      "type": "boolean"
    },
    "exam": {
      "description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "exposure": {
      "description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    },
    "fingerprint_max_n": {
      "default": 4,
      "maximum": 5,
      "minimum": 3,
      "type": "integer"
    },
    "policy": {
      "additionalProperties": false,
      "description": "Explicit promotion policy. Omitted fields use the documented defaults.",
      "properties": {
        "confidence_alpha": {
          "default": 0.05,
          "exclusiveMinimum": 0,
          "maximum": 1,
          "type": "number"
        },
        "max_regressions": {
          "default": 0,
          "minimum": 0,
          "type": "integer"
        },
        "min_gains": {
          "default": 1,
          "minimum": 0,
          "type": "integer"
        },
        "require_retained_probe": {
          "default": false,
          "type": "boolean"
        }
      },
      "type": "object"
    },
    "retained_probe": {
      "additionalProperties": false,
      "description": "Optional retained-capability result checked alongside the paired cohort.",
      "properties": {
        "base_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "candidate_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "items": {
          "minimum": 0,
          "type": "integer"
        }
      },
      "required": [
        "base_verified",
        "candidate_verified",
        "items"
      ],
      "type": "object"
    },
    "similarity_threshold": {
      "default": 0.6,
      "maximum": 1,
      "minimum": 0.5,
      "type": "number"
    }
  },
  "required": [
    "exam",
    "baseline",
    "candidate"
  ],
  "additionalProperties": false,
  "description": "Audit exposure, prove a complete clean cohort, then gate baseline versus candidate."
}
🟢audit_leakage(enable_behavioral_fingerprint, enable_text_similarity, exam, exposure, fingerprint_max_n, ...)

Exact declared-exposure audit over your exam rows: row identity, behavioral fingerprints for graph-DSL expressions, text-similarity review flags, and a clean exam export. Full example: GET /api/examples key 'leakage'.

输入模式

{
  "type": "object",
  "properties": {
    "enable_behavioral_fingerprint": {
      "default": true,
      "type": "boolean"
    },
    "enable_text_similarity": {
      "default": true,
      "type": "boolean"
    },
    "exam": {
      "description": "Exam rows. Each row needs item_id (or id) plus prompt/content/input/task/question/expression.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "exposure": {
      "description": "Declared exposure rows carrying identity/content fields and an optional source or path.",
      "items": {
        "additionalProperties": true,
        "properties": {},
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    },
    "fingerprint_max_n": {
      "default": 4,
      "maximum": 5,
      "minimum": 3,
      "type": "integer"
    },
    "similarity_threshold": {
      "default": 0.6,
      "maximum": 1,
      "minimum": 0.5,
      "type": "number"
    }
  },
  "required": [
    "exam"
  ],
  "additionalProperties": false,
  "description": "Audit declared exposure against an exam and export the clean remainder."
}
🟢promotion_gate(baseline, baseline_name, candidate, candidate_name, domains, ...)

PASS, HOLD, or BLOCK from paired per-item results: gains, regressions, exact McNemar p-value, per-domain breakdown. Full example: GET /api/examples key 'gate'.

输入模式

{
  "type": "object",
  "properties": {
    "baseline": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the baseline system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "baseline_name": {
      "default": "baseline",
      "type": "string"
    },
    "candidate": {
      "additionalProperties": {
        "type": "boolean"
      },
      "description": "item_id -> boolean pass/fail result for the candidate system.",
      "maxProperties": 5000,
      "minProperties": 1,
      "type": "object"
    },
    "candidate_name": {
      "default": "candidate",
      "type": "string"
    },
    "domains": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "Optional item_id -> domain label mapping.",
      "type": "object"
    },
    "policy": {
      "additionalProperties": false,
      "description": "Explicit promotion policy. Omitted fields use the documented defaults.",
      "properties": {
        "confidence_alpha": {
          "default": 0.05,
          "exclusiveMinimum": 0,
          "maximum": 1,
          "type": "number"
        },
        "max_regressions": {
          "default": 0,
          "minimum": 0,
          "type": "integer"
        },
        "min_gains": {
          "default": 1,
          "minimum": 0,
          "type": "integer"
        },
        "require_retained_probe": {
          "default": false,
          "type": "boolean"
        }
      },
      "type": "object"
    },
    "retained_probe": {
      "additionalProperties": false,
      "description": "Optional retained-capability result checked alongside the paired cohort.",
      "properties": {
        "base_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "candidate_verified": {
          "minimum": 0,
          "type": "integer"
        },
        "items": {
          "minimum": 0,
          "type": "integer"
        }
      },
      "required": [
        "base_verified",
        "candidate_verified",
        "items"
      ],
      "type": "object"
    }
  },
  "required": [
    "baseline",
    "candidate"
  ],
  "additionalProperties": false,
  "description": "Compare identical baseline and candidate item cohorts under an explicit policy."
}
🟢bank_health(history, items)

Item-lifecycle diagnostics over your grading history: discriminators, saturated and flaky items, frontier gaps. Full example: GET /api/examples key 'health'.

输入模式

{
  "type": "object",
  "properties": {
    "history": {
      "description": "Observed item/system outcomes.",
      "items": {
        "additionalProperties": true,
        "properties": {
          "domain": {
            "type": "string"
          },
          "item_id": {
            "minLength": 1,
            "type": "string"
          },
          "passed": {
            "type": "boolean"
          },
          "system": {
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "item_id",
          "system",
          "passed"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "items": {
      "description": "Optional item definitions.",
      "items": {
        "additionalProperties": true,
        "properties": {
          "domain": {
            "type": "string"
          },
          "item_id": {
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "item_id"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 0,
      "type": "array"
    }
  },
  "required": [
    "history"
  ],
  "additionalProperties": false,
  "description": "Diagnose item lifecycle health from one or more grading-history rows."
}
🟡safe_patch(document, operations, reason)

Apply a section-scoped Markdown patch under conservation checks (untouched sections stay byte-identical; protected tokens preserved). Full example: GET /api/examples key 'safepatch'.

输入模式

{
  "type": "object",
  "properties": {
    "document": {
      "description": "Complete Markdown document to patch.",
      "maxLength": 200000,
      "minLength": 1,
      "type": "string"
    },
    "operations": {
      "items": {
        "additionalProperties": false,
        "properties": {
          "allow_token_changes": {
            "description": "Protected literal tokens that this operation may intentionally change.",
            "items": {
              "type": "string"
            },
            "maxItems": 100,
            "type": "array"
          },
          "find": {
            "minLength": 1,
            "type": "string"
          },
          "replace": {
            "type": "string"
          },
          "target_heading": {
            "description": "Markdown heading text without the leading # characters.",
            "minLength": 1,
            "type": "string"
          }
        },
        "required": [
          "target_heading",
          "find",
          "replace"
        ],
        "type": "object"
      },
      "maxItems": 50,
      "minItems": 1,
      "type": "array"
    },
    "reason": {
      "type": "string"
    }
  },
  "required": [
    "document",
    "operations"
  ],
  "additionalProperties": false,
  "description": "Apply deterministic, section-scoped Markdown replacements under conservation checks."
}
🟢counterexample_hunt(expression, ns, restarts, seed, steps)

Bounded simulated-annealing search for a graph counterexample inside a DSL predicate class, with an exact certificate when found. CPU-bounded and strictly rate-limited. Full example: GET /api/examples key 'counterexample'.

输入模式

{
  "type": "object",
  "properties": {
    "expression": {
      "description": "Graph predicate in the Whetstone DSL, for example: is_connected and is_triangle_free and not is_bipartite",
      "maxLength": 500,
      "minLength": 1,
      "type": "string"
    },
    "ns": {
      "default": [
        8,
        9,
        10,
        11
      ],
      "description": "Graph sizes searched.",
      "items": {
        "maximum": 12,
        "minimum": 4,
        "type": "integer"
      },
      "maxItems": 5,
      "minItems": 1,
      "type": "array"
    },
    "restarts": {
      "default": 4,
      "maximum": 6,
      "minimum": 1,
      "type": "integer"
    },
    "seed": {
      "default": 0,
      "type": "integer"
    },
    "steps": {
      "default": 800,
      "maximum": 1500,
      "minimum": 50,
      "type": "integer"
    }
  },
  "required": [
    "expression"
  ],
  "additionalProperties": false,
  "description": "Run a bounded graph search against one Whetstone predicate expression."
}
🟡memory_relevance(context_entities, current_step, memories, objective, objective_entities, ...)

Compare query-free salience against objective-conditioned relevance for a set of memories under a token budget. Full example: GET /api/examples key 'memory'.

输入模式

{
  "type": "object",
  "properties": {
    "context_entities": {
      "description": "Entities already active in context.",
      "items": {
        "type": "string"
      },
      "maxItems": 100,
      "type": "array"
    },
    "current_step": {
      "minimum": 0,
      "type": "integer"
    },
    "memories": {
      "description": "Memories to rank.",
      "items": {
        "additionalProperties": false,
        "properties": {
          "age": {
            "default": 0,
            "minimum": 0,
            "type": "integer"
          },
          "confidence": {
            "default": 0.8,
            "maximum": 1,
            "minimum": 0,
            "type": "number"
          },
          "content": {
            "minLength": 1,
            "type": "string"
          },
          "entities": {
            "description": "Entities explicitly present in this memory.",
            "items": {
              "type": "string"
            },
            "maxItems": 100,
            "type": "array"
          },
          "kind": {
            "default": "episodic",
            "type": "string"
          },
          "source": {
            "default": "uploaded",
            "type": "string"
          },
          "use_count": {
            "default": 0,
            "minimum": 0,
            "type": "integer"
          }
        },
        "required": [
          "content"
        ],
        "type": "object"
      },
      "maxItems": 1000,
      "minItems": 1,
      "type": "array"
    },
    "objective": {
      "minLength": 1,
      "type": "string"
    },
    "objective_entities": {
      "description": "Optional explicit entities when the objective text is not self-describing.",
      "items": {
        "type": "string"
      },
      "maxItems": 100,
      "type": "array"
    },
    "question_kind": {
      "default": "generic",
      "type": "string"
    },
    "token_budget": {
      "default": 90,
      "maximum": 10000,
      "minimum": 1,
      "type": "integer"
    }
  },
  "required": [
    "objective",
    "memories"
  ],
  "additionalProperties": false,
  "description": "Rank caller-supplied memories against a concrete objective under a token budget."
}
🟢replay_trace(events, notes)

Turn reasoning-emulator control events into checkpoints, rewinds, notes, and a timeline. Full example: GET /api/examples key 'replay'.

输入模式

{
  "type": "object",
  "properties": {
    "events": {
      "description": "Ordered reasoning-emulator events.",
      "items": {
        "additionalProperties": false,
        "properties": {
          "detail": {
            "type": "string"
          },
          "kind": {
            "description": "Event class such as control, verifier, model, or observation.",
            "type": "string"
          },
          "source": {
            "default": "native",
            "type": "string"
          },
          "step": {
            "minimum": 0,
            "type": "integer"
          }
        },
        "required": [
          "kind",
          "detail"
        ],
        "type": "object"
      },
      "maxItems": 5000,
      "minItems": 1,
      "type": "array"
    },
    "notes": {
      "description": "Optional analyst notes.",
      "items": {
        "type": "string"
      },
      "maxItems": 5000,
      "type": "array"
    }
  },
  "required": [
    "events"
  ],
  "additionalProperties": false,
  "description": "Reconstruct checkpoints, rewinds, branches, and verifier outcomes from control events."
}
🟡report_card_start(challenge)

TIER 1: start a disposable report-card session. Returns exam items (graph-repair prompts minted from the repository's public frontier) for THIS agent to answer. Answer every item, then call report_card_submit exactly once. Sessions are one-shot, expire in 15 minutes, and are strictly rate-limited. This demonstrates the promotion-gate mechanism on disposable items; it is not a private-bank credential.

输入模式

{
  "type": "object",
  "properties": {
    "challenge": {
      "description": "Caller nonce bound into the signed receipt for replay detection.",
      "maxLength": 128,
      "minLength": 8,
      "type": "string"
    }
  },
  "additionalProperties": false
}
🟡report_card_submit(answers, session_id)

TIER 1: submit answers for a report-card session and receive the graded report (per-item verdicts, per-domain totals, SHA-256 commitments). Grading is by checker spec: verified strict refinements are reported separately, and promotion grade requires at least 5% clean-support retention. No answer key exists. The session is destroyed by this call.

输入模式

{
  "type": "object",
  "properties": {
    "answers": {
      "additionalProperties": {
        "type": "string"
      },
      "description": "item_id -> answer (a DSL predicate, or the JSON reply the prompt asked for)",
      "type": "object"
    },
    "session_id": {
      "type": "string"
    }
  },
  "required": [
    "session_id",
    "answers"
  ],
  "additionalProperties": false
}
🟡open_bench_start(challenge)

TIER 2: start a one-shot Open Promotion Bench session. Returns six fresh virtual-repository scope-integrity tasks. Run a baseline and candidate independently on the same cohort, then submit both answer maps with open_bench_submit. This is an open, procedural, self-attested track rather than a private-bank credential.

输入模式

{
  "type": "object",
  "properties": {
    "challenge": {
      "description": "Caller nonce bound into the signed receipt for replay detection.",
      "maxLength": 128,
      "minLength": 8,
      "type": "string"
    }
  },
  "additionalProperties": false
}
🟡open_bench_submit(attestation, baseline_answers, baseline_manifest, candidate_answers, candidate_manifest, ...)

TIER 2: grade paired baseline and candidate patches, count gains/regressions/ties, and issue PASS/HOLD/BLOCK. Set publish=true plus attestation=true to append only the safe manifests and sanitized receipt to the public board; tasks and answers are never persisted.

输入模式

{
  "type": "object",
  "properties": {
    "attestation": {
      "type": "boolean"
    },
    "baseline_answers": {
      "type": "object"
    },
    "baseline_manifest": {
      "additionalProperties": false,
      "properties": {
        "harness": {
          "type": "string"
        },
        "model": {
          "type": "string"
        },
        "name": {
          "type": "string"
        },
        "version": {
          "type": "string"
        }
      },
      "required": [
        "name"
      ],
      "type": "object"
    },
    "candidate_answers": {
      "type": "object"
    },
    "candidate_manifest": {
      "additionalProperties": false,
      "properties": {
        "harness": {
          "type": "string"
        },
        "model": {
          "type": "string"
        },
        "name": {
          "type": "string"
        },
        "version": {
          "type": "string"
        }
      },
      "required": [
        "name"
      ],
      "type": "object"
    },
    "publish": {
      "type": "boolean"
    },
    "session_id": {
      "type": "string"
    }
  },
  "required": [
    "session_id",
    "baseline_manifest",
    "candidate_manifest",
    "baseline_answers",
    "candidate_answers"
  ],
  "additionalProperties": false
}
🟢open_bench_leaderboard

TIER 2: list the self-attested public Open Promotion Bench receipts. Entries contain manifests, verdicts, item-level transitions, and commitments but never task contents or submitted answers.

输入模式

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
⚪about_whetstone

What this service is: the tool catalog, the tier boundaries, and where the source lives.

输入模式

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}

社区

评价此服务器

证据

最近观测

已验证未记录版本14 个工具
已验证未记录版本14 个工具