ai-eval
Cloudflare Workers MCP server: ai-eval
我該用這個嗎
品質與安全性
B
發現項目(2)
- LOW在 score_response 中
- LOW在 compare_responses 中
根據工具定義與協定合規性的自動化分析。
上下文成本
~214Token(工具定義)
~567 B典型回應大小
極小的注意力影響(128k 上下文的 0.17%)
這是每次將伺服器的工具載入模型上下文時所消耗的約略 token 數量。數量越高,可用於其他工作的注意力就越少。
安裝
一鍵安裝
將以下內容加入你的 `claude_desktop_config.json` 檔案:
{
"mcpServers": {
"ai-eval": {
"url": "https://api.lazy-mac.com/ai-eval/mcp"
}
}
}遠端端點
https://api.lazy-mac.com/ai-eval/mcpstreamable-http它能做什麼
工具清單
工具(3)
🟢 唯讀🟡 寫入🔴 刪除⚪ 未知
⚪score_response(prompt, response, criteria)
Score an AI response against a prompt using heuristic metrics (length, relevance, structure, completeness)
輸入結構描述
{
"type": "object",
"properties": {
"prompt": {
"type": "string",
"description": "The original prompt/question"
},
"response": {
"type": "string",
"description": "The AI response to evaluate"
},
"criteria": {
"type": "array",
"items": {
"type": "string"
},
"description": "Optional keywords that should appear in response"
}
},
"required": [
"prompt",
"response"
]
}⚪compare_responses(prompt, responses)
Compare and rank multiple AI responses to the same prompt
輸入結構描述
{
"type": "object",
"properties": {
"prompt": {
"type": "string"
},
"responses": {
"type": "array",
"minItems": 2,
"items": {
"type": "string"
}
}
},
"required": [
"prompt",
"responses"
]
}🟢text_metrics(text)
Get text quality metrics: word count, sentence count, estimated tokens, readability grade
輸入結構描述
{
"type": "object",
"properties": {
"text": {
"type": "string"
}
},
"required": [
"text"
]
}社群
證據
近期觀測
已驗證未記錄版本3 個工具
已驗證未記錄版本3 個工具
已驗證未記錄版本3 個工具