Reality Graph Verification Tools
Read-only AI coding tools for change verification, release readiness, capacity, and guidance.
我该使用它吗
质量与安全性
发现(6)
- HIGH
- MEDIUM在 validate_task_contract 中
- LOW在 get_verification_report_template 中
- LOW在 calculate_verification_capacity 中
- LOW在 fetch 中
- INFO在 validate_task_contract 中
基于对工具定义和协议合规性的自动分析。
上下文开销
这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。
安装
一键安装
将以下内容添加到你的 `claude_desktop_config.json` 文件中:
{
"mcpServers": {
"verification-tools": {
"url": "https://realitygraph.dev/api/mcp"
}
}
}远程端点
https://realitygraph.dev/api/mcpstreamable-http它能做什么
工具清单
工具(10)
🟢check_verification_debt(team_size, ai_share_percent, prs_per_month, merged_loc_per_week, reviewer_hours_per_week, ...)
Estimate a software team's verification debt from team parameters. Computes the four published metrics (generation-to-verification ratio, review depth, unverified-merge rate, two-week churn) and an annual cost estimate, with the full calculation path, labeled assumptions, thresholds, and sources (GitClear, Sonar, Faros, Veracode). Deterministic arithmetic from published models - no benchmark claims. Only team_size is required; every additional parameter refines the estimate. Set lang='de' for a German report.
输入模式
{
"type": "object",
"properties": {
"team_size": {
"type": "integer",
"minimum": 1,
"maximum": 500,
"description": "Number of developers on the team (required)"
},
"ai_share_percent": {
"description": "Share of merges that are AI-assisted, in percent (default: 60, assumption)",
"type": "number",
"minimum": 0,
"maximum": 100
},
"prs_per_month": {
"description": "Total merged PRs per month (default: team_size x prs_per_engineer_per_month)",
"type": "integer",
"minimum": 1,
"maximum": 100000
},
"merged_loc_per_week": {
"description": "Merged changed lines of code per week (enables the GVR and review-depth metrics)",
"type": "number",
"minimum": 0,
"maximum": 100000000
},
"reviewer_hours_per_week": {
"description": "Reviewer hours actually spent per week (enables the GVR metric)",
"type": "number",
"minimum": 0,
"maximum": 10000
},
"substantive_review_comments_per_week": {
"description": "Substantive review comments per week, excluding bots and nitpicks (enables the review-depth metric)",
"type": "number",
"minimum": 0,
"maximum": 1000000
},
"ai_merges_per_month": {
"description": "AI-assisted merges per month (enables the unverified-merge rate)",
"type": "integer",
"minimum": 0,
"maximum": 100000
},
"ai_merges_with_evidence_per_month": {
"description": "AI-assisted merges per month with recorded validation evidence (enables the unverified-merge rate)",
"type": "integer",
"minimum": 0,
"maximum": 100000
},
"two_week_churn_percent": {
"description": "Share of new lines revised or reverted within 14 days, in percent. A warning signal in the metrics block; it never enters the cost model, because it measures lines and the cost model counts changes",
"type": "number",
"minimum": 0,
"maximum": 100
},
"rework_rate_percent": {
"description": "Share of AI-assisted changes reworked for a defect within 14 days, in percent (default: 2, the illustrative rate from /cost-of-verification-debt - replace it with your own reason-coded rate)",
"type": "number",
"minimum": 0,
"maximum": 100
},
"prs_per_engineer_per_month": {
"description": "Merged PRs per engineer per month (default: 20, derived from the published worked report on /measure-verification-debt)",
"type": "number",
"minimum": 0.1,
"maximum": 500
},
"hourly_rate_eur": {
"description": "Loaded cost per engineer hour in EUR (default: 75, assumption)",
"type": "number",
"minimum": 1,
"maximum": 1000
},
"hours_per_reworked_change": {
"description": "Average hours per reworked change (default: 6, assumption)",
"type": "number",
"minimum": 0.1,
"maximum": 100
},
"review_reconstruction_hours_per_pr": {
"description": "Average reviewer hours spent reconstructing intent per AI-assisted PR (default: 0.5, assumption)",
"type": "number",
"minimum": 0,
"maximum": 20
},
"incident_allowance_eur_per_year": {
"description": "Annual incident allowance in EUR (default: 0; add one only when you have a locally defined incident class, frequency and expected-loss method)",
"type": "number",
"minimum": 0,
"maximum": 10000000
},
"lang": {
"description": "Report language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"team_size"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢lint_task_spec(task, lang)
Check whether a free-text work order for an AI coding agent is verifiable BEFORE handing it over. Heuristic, deterministic lint of the task's form against the four building blocks of a checkable task (goal, boundaries, acceptance criteria, validation plan) plus rule checks (vague adjectives without numbers, unnamed unhappy paths, missing file anchors). Returns a status table with evidence, the concrete questions that close each gap, and a fill-in skeleton. It checks form, not content — no LLM, nothing stored. Set lang='de' for a German report.
输入模式
{
"type": "object",
"properties": {
"task": {
"type": "string",
"minLength": 10,
"maxLength": 8000,
"description": "The work order / task text you intend to give an AI coding agent (English or German)"
},
"lang": {
"description": "Report language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"task"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_task_contract_template(format, lang)
Returns Reality Graph's free fill-in template (v0) for a verifiable task contract: goal, non-goals, boundaries (may change / must not change / forbidden), 3-7 yes/no acceptance criteria, validation plan, expected evidence, assumptions, open questions — with a filled example and fill-in guidance. Write the contract before an AI agent runs; verify the result against it after. format='json' returns a machine-fillable JSON structure; default is a compact markdown skeleton. Set lang='de' for German. Static content, nothing stored.
输入模式
{
"type": "object",
"properties": {
"format": {
"description": "Template format (default: markdown)",
"type": "string",
"enum": [
"markdown",
"json"
]
},
"lang": {
"description": "Language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢validate_task_contract(contract, lang)
Deterministically validates a FILLED task contract (the JSON structure from get_task_contract_template): completeness of goal/non-goals/boundaries, decidability of each acceptance criterion (vague words, missing measurable markers), automated checks in the validation plan, expected evidence, and leftover placeholders. Returns a verdict (PASS / PASS WITH WARNINGS / FAIL), four dimension scores, and a concrete fix per finding. Validates form and completeness, not correctness. No LLM, nothing stored. lang='de' for German.
输入模式
{
"type": "object",
"properties": {
"contract": {
"type": "string",
"minLength": 20,
"maxLength": 16000,
"description": "The filled task contract as a JSON string (structure from get_task_contract_template, format='json')"
},
"lang": {
"description": "Report language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"contract"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_verification_report_template(format, lang)
Returns the free fill-in template (v0) for a verification report — the artifact you write right after an AI-assisted run: task recap, files changed AND files confirmed untouched, validation results per acceptance criterion (not authored by the generating model), what was skipped, limitations, and the explicit decision. format='json' for a machine-fillable structure; default is a compact markdown file. Static content, nothing stored. lang='de' for German.
输入模式
{
"type": "object",
"properties": {
"format": {
"description": "Template format (default: markdown)",
"type": "string",
"enum": [
"markdown",
"json"
]
},
"lang": {
"description": "Language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢search(query, lang)
Full-text search over the Reality Graph knowledge base on AI coding verification: 40+ glossary definitions, 700+ FAQ answers, sourced statistics, and article summaries on verification debt, AI code review, spec-vs-implementation checking, EU compliance (EU AI Act, GDPR, NIS2), and AI coding governance — in English and German. Returns matching documents with title, URL, and snippet. Use fetch to read a result.
输入模式
{
"type": "object",
"properties": {
"query": {
"type": "string",
"minLength": 2,
"maxLength": 300,
"description": "Search query (English or German)"
},
"lang": {
"description": "Restrict results to one language (default: both)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"query"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}输出模式
{
"type": "object",
"properties": {
"results": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"title": {
"type": "string"
},
"url": {
"type": "string"
},
"snippet": {
"type": "string"
}
},
"required": [
"id",
"title",
"url",
"snippet"
],
"additionalProperties": false
}
}
},
"required": [
"results"
],
"$schema": "http://json-schema.org/draft-07/schema#",
"additionalProperties": false
}🟢fetch(id)
Fetch a document from the Reality Graph knowledge base by id (as returned by search, e.g. '/verification-debt') or by full realitygraph.dev URL. Returns the document's summary, definitions, key facts, FAQ, and sources as text, plus the canonical URL.
输入模式
{
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1,
"maxLength": 300,
"description": "Document id from search results, or a realitygraph.dev URL"
}
},
"required": [
"id"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}输出模式
{
"type": "object",
"properties": {
"id": {
"type": "string"
},
"title": {
"type": "string"
},
"text": {
"type": "string"
},
"url": {
"type": "string"
},
"metadata": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {
"type": "string"
}
}
},
"required": [
"id",
"title",
"text",
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#",
"additionalProperties": false
}🟢plan_change_verification(change_summary, change_types, blast_radius, rollback, lang)
Turn explicit change characteristics into a risk tier, required automated checks, manual scenarios, evidence, release blockers, role handoff, and canonical Reality Graph guidance. Use before implementation or review. It does not inspect code and never invents a confidence score.
输入模式
{
"type": "object",
"properties": {
"change_summary": {
"type": "string",
"minLength": 10,
"maxLength": 2000,
"description": "Plain-language summary of the change"
},
"change_types": {
"minItems": 1,
"maxItems": 10,
"type": "array",
"items": {
"type": "string",
"enum": [
"ui",
"api",
"auth",
"database",
"payments",
"personal_data",
"dependency",
"infrastructure",
"public_api",
"compliance"
]
},
"description": "Technical and risk-relevant change types"
},
"blast_radius": {
"type": "string",
"enum": [
"single_component",
"service",
"multi_service",
"customer_data",
"production_wide"
],
"description": "Largest expected impact boundary"
},
"rollback": {
"type": "string",
"enum": [
"automatic",
"documented",
"manual",
"none",
"unknown"
],
"description": "Current rollback or recovery state"
},
"lang": {
"description": "Response language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"change_summary",
"change_types",
"blast_radius",
"rollback"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢check_release_readiness(change_summary, change_types, blast_radius, rollback, lang, ...)
Return GO, CONDITIONAL, or NO_GO from supplied acceptance-criterion results, check evidence, rollback, monitoring, limitations, and independent review. The verdict is deliberately based only on supplied evidence; this tool does not inspect code, CI, or a deployment.
输入模式
{
"type": "object",
"properties": {
"change_summary": {
"type": "string",
"minLength": 10,
"maxLength": 2000,
"description": "Plain-language summary of the change"
},
"change_types": {
"minItems": 1,
"maxItems": 10,
"type": "array",
"items": {
"type": "string",
"enum": [
"ui",
"api",
"auth",
"database",
"payments",
"personal_data",
"dependency",
"infrastructure",
"public_api",
"compliance"
]
},
"description": "Technical and risk-relevant change types"
},
"blast_radius": {
"type": "string",
"enum": [
"single_component",
"service",
"multi_service",
"customer_data",
"production_wide"
],
"description": "Largest expected impact boundary"
},
"rollback": {
"type": "string",
"enum": [
"automatic",
"documented",
"manual",
"none",
"unknown"
],
"description": "Current rollback or recovery state"
},
"lang": {
"description": "Response language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
},
"acceptance_criteria_passed": {
"type": "integer",
"minimum": 0,
"maximum": 10000
},
"acceptance_criteria_failed": {
"type": "integer",
"minimum": 0,
"maximum": 10000
},
"acceptance_criteria_not_run": {
"type": "integer",
"minimum": 0,
"maximum": 10000
},
"checks": {
"maxItems": 100,
"type": "array",
"items": {
"type": "object",
"properties": {
"kind": {
"type": "string",
"enum": [
"build",
"lint",
"typecheck",
"unit",
"integration",
"e2e",
"accessibility",
"security",
"migration",
"manual"
]
},
"name": {
"type": "string",
"minLength": 1,
"maxLength": 200
},
"status": {
"type": "string",
"enum": [
"pass",
"fail",
"not_run"
]
},
"evidence": {
"type": "string",
"maxLength": 1000
}
},
"required": [
"kind",
"status"
]
}
},
"rollback_ready": {
"type": "boolean"
},
"monitoring_ready": {
"type": "boolean"
},
"known_limitations_recorded": {
"type": "boolean"
},
"independent_review": {
"type": "boolean"
}
},
"required": [
"change_summary",
"change_types",
"blast_radius",
"rollback",
"acceptance_criteria_passed",
"acceptance_criteria_failed",
"acceptance_criteria_not_run",
"checks",
"rollback_ready",
"monitoring_ready",
"known_limitations_recorded",
"independent_review"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢calculate_verification_capacity(ai_changes_per_week, average_review_minutes_per_change, available_reviewer_hours_per_week, evidence_coverage_percent, two_week_churn_percent, ...)
Calculate weekly review demand, utilization, capacity gap, supported change throughput, and changes lacking evidence from measured team inputs. No cost model, benchmark, or hidden industry assumption is applied; the output shows the arithmetic and a concrete balancing action.
输入模式
{
"type": "object",
"properties": {
"ai_changes_per_week": {
"type": "integer",
"minimum": 0,
"maximum": 1000000
},
"average_review_minutes_per_change": {
"type": "number",
"minimum": 0.1,
"maximum": 10000
},
"available_reviewer_hours_per_week": {
"type": "number",
"minimum": 0,
"maximum": 100000
},
"evidence_coverage_percent": {
"type": "number",
"minimum": 0,
"maximum": 100
},
"two_week_churn_percent": {
"type": "number",
"minimum": 0,
"maximum": 100
},
"lang": {
"description": "Response language (default: en)",
"type": "string",
"enum": [
"en",
"de"
]
}
},
"required": [
"ai_changes_per_week",
"average_review_minutes_per_change",
"available_reviewer_hours_per_week",
"evidence_coverage_percent"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}社区
证据