torneo
Observatory operated and funded by devlo: real tools on frozen tasks; intervals, cost, limits.
Should I use this
Quality & Safety
Based on automated analysis of tool definitions and protocol compliance.
Context Cost
This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.
Install
One-Click Install
Add this to your `claude_desktop_config.json` file:
{
"mcpServers": {
"torneo": {
"url": "https://torneo.ai/api/mcp"
}
}
}Remote endpoints
https://torneo.ai/api/mcpstreamable-httpWhat it can do
Tool inventory
Tools (4)
🟢list_categories
Lists every category with its status (OK: replicated ranking; LOCAL_VALIDITY: single block, no current rank claim; INDETERMINATE: precision insufficient, no rank), observation date, freshness and source run. Categories prefixed 'fixture-' are synthetic demo data validating the machinery, never evidence about real tools.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢get_results(category)
Answers for one category with the canonical answer rule: never 'the best tool', only the best observed evidence for this task, this context, at this date, with intervals, costs, conflicts and limits. On a STALE, SUPERSEDED or INDETERMINATE result the answer is INSUFFICIENT_EVIDENCE and carries no rank. Identical to `torneo query <category>` on the CLI.
Input Schema
{
"type": "object",
"properties": {
"category": {
"type": "string",
"description": "Category id, e.g. 'transcription' (see list_categories)"
}
},
"required": [
"category"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢get_run(run_id)
Returns the canonical result bundle of one run (schema result.v1): protocol lock hash, provenance, reproduce command, per-participant outcomes with intervals. A PRE-REGISTERED run, whose protocol is frozen and timestamped but which has not been executed, returns state PRE_REGISTERED with measured false, its lock and its frozen files, and no result: nothing has been measured yet.
Input Schema
{
"type": "object",
"properties": {
"run_id": {
"type": "string",
"description": "Run id, e.g. 'TRANSCRIPTION-001' (see list_categories, field run_id)"
}
},
"required": [
"run_id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢explain_limits(category, run_id)
Returns what a category's (or run's) result can and cannot tell you: status and its meaning, fixture flag, freshness, expiry, explicit limits, conflicts, funding, published errata (each chained to the served result hash), the legal preflight verdict per tool including tools not run, and the answer rule every consumer must follow.
Input Schema
{
"type": "object",
"properties": {
"category": {
"description": "Category id",
"type": "string"
},
"run_id": {
"description": "Run id (alternative to category)",
"type": "string"
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Community
Evidence