ALPNAI — Agent Performance Tools
Agent cost, latency and quality analysis, report delivery and order status. Crypto payments paused.
我该使用它吗
质量与安全性
基于对工具定义和协议合规性的自动分析。
上下文开销
这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。
安装
一键安装
将以下内容添加到你的 `claude_desktop_config.json` 文件中:
{
"mcpServers": {
"alpnai": {
"url": "https://alpnai.com/api/mcp"
}
}
}远程端点
https://alpnai.com/api/mcpstreamable-http它能做什么
工具清单
工具(10)
🟢get_catalog
Discover cost, latency and quality analysis APIs, their inputs, outputs, free reproducible example, authentication and current payment availability. No payment or budget debit.
输入模式
{
"type": "object",
"properties": {},
"additionalProperties": false
}🟢get_free_sample
Read a dated primary-source sample about OpenAI. Not real-time or exhaustive.
输入模式
{
"type": "object",
"properties": {},
"additionalProperties": false
}🟢audit_agent_costs(runs, config)
Free comparison of cost per successful task on supplied paired traces. Requires a sandbox agent key. No purchase, storage, external model call or automatic deployment. Success labels are supplied by the client.
输入模式
{
"type": "object",
"properties": {
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"task_id": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units."
},
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"variant": {
"type": "string",
"enum": [
"baseline",
"candidate"
]
},
"cost_usd": {
"type": "number",
"minimum": 0,
"maximum": 10000,
"description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals."
},
"success": {
"type": "boolean"
},
"latency_ms": {
"type": "number",
"minimum": 0,
"maximum": 86400000,
"description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured."
}
},
"required": [
"task_id",
"workflow",
"variant",
"cost_usd",
"success"
],
"additionalProperties": false
},
"minItems": 1,
"maxItems": 1000
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [],
"additionalProperties": false
}
},
"required": [
"runs"
],
"additionalProperties": false,
"title": "ALPNAI recorded agent attempts",
"description": "Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
}输出模式
{
"type": "object",
"properties": {
"schema_version": {
"type": "string",
"const": "1.0.0"
},
"purpose": {
"type": "string",
"const": "cost_per_successful_agent_task_audit"
},
"data_provenance": {
"type": "string",
"const": "user_supplied_not_verified"
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [
"minSamples",
"minSuccessRate",
"maxSuccessRateDrop"
],
"additionalProperties": true
},
"input_attempt_records": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"experimental_bias_control": {
"type": "string",
"const": "unknown"
},
"automatic_deployment_authorized": {
"type": "boolean",
"const": false
},
"methodology": {
"type": "object",
"properties": {
"logical_task_key": {
"type": "array",
"items": {
"type": "string"
},
"minItems": 3,
"maxItems": 3
},
"success": {
"type": "string"
},
"duplicates": {
"type": "string"
},
"latency": {
"type": "string"
},
"quality": {
"type": "string"
},
"uncertainty": {
"type": "string"
},
"missing_costs": {
"type": "string"
}
},
"required": [
"logical_task_key",
"success",
"duplicates",
"latency",
"quality",
"uncertainty",
"missing_costs"
],
"additionalProperties": true
},
"workflows": {
"type": "array",
"items": {
"type": "object",
"properties": {
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"baseline": {
"anyOf": [
{
"type": "object",
"properties": {
"tasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"attempts": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"retry_attempts": {
"type": "integer",
"minimum": 0,
"maximum": 999
},
"successful_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"total_cost_usd": {
"type": "number",
"minimum": 0
},
"cost_per_task_usd": {
"type": "number",
"minimum": 0
},
"cost_per_successful_task_usd": {
"anyOf": [
{
"type": "number",
"minimum": 0
},
{
"type": "null"
}
]
},
"success_rate": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"success_rate_interval_95_wilson": {
"type": "array",
"items": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"minItems": 2,
"maxItems": 2
},
"p95_recorded_attempt_latency_per_task_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"latency_complete": {
"type": "boolean"
}
},
"required": [
"tasks",
"attempts",
"retry_attempts",
"successful_tasks",
"total_cost_usd",
"cost_per_task_usd",
"cost_per_successful_task_usd",
"success_rate",
"success_rate_interval_95_wilson",
"p95_recorded_attempt_latency_per_task_ms",
"latency_complete"
],
"additionalProperties": true
},
{
"type": "null"
}
]
},
"candidate": {
"anyOf": [
{
"type": "object",
"properties": {
"tasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"attempts": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"retry_attempts": {
"type": "integer",
"minimum": 0,
"maximum": 999
},
"successful_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"total_cost_usd": {
"type": "number",
"minimum": 0
},
"cost_per_task_usd": {
"type": "number",
"minimum": 0
},
"cost_per_successful_task_usd": {
"anyOf": [
{
"type": "number",
"minimum": 0
},
{
"type": "null"
}
]
},
"success_rate": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"success_rate_interval_95_wilson": {
"type": "array",
"items": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"minItems": 2,
"maxItems": 2
},
"p95_recorded_attempt_latency_per_task_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"latency_complete": {
"type": "boolean"
}
},
"required": [
"tasks",
"attempts",
"retry_attempts",
"successful_tasks",
"total_cost_usd",
"cost_per_task_usd",
"cost_per_successful_task_usd",
"success_rate",
"success_rate_interval_95_wilson",
"p95_recorded_attempt_latency_per_task_ms",
"latency_complete"
],
"additionalProperties": true
},
{
"type": "null"
}
]
},
"comparison": {
"type": "object",
"properties": {
"matched_task_ids": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"baseline_only_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"candidate_only_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"same_task_set": {
"type": "boolean"
}
},
"required": [
"matched_task_ids",
"baseline_only_tasks",
"candidate_only_tasks",
"same_task_set"
],
"additionalProperties": true
},
"gates": {
"type": "object",
"properties": {
"both_variants": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"minimum_distinct_tasks_per_variant": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"same_task_set": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"observed_success_rate": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"recorded_latency": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"lower_cost_per_successful_task": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"experimental_bias_control": {
"type": "string",
"const": "unknown"
}
},
"required": [
"both_variants",
"minimum_distinct_tasks_per_variant",
"same_task_set",
"observed_success_rate",
"recorded_latency",
"lower_cost_per_successful_task",
"experimental_bias_control"
],
"additionalProperties": true
},
"decision": {
"type": "string",
"enum": [
"missing_comparison",
"collect_more_data",
"quality_regression",
"latency_data_required",
"latency_regression",
"no_economic_advantage",
"candidate_for_controlled_trial"
]
},
"automatic_deployment_authorized": {
"type": "boolean",
"const": false
},
"realized_savings_usd": {
"type": "null"
},
"opportunity_monthly": {
"anyOf": [
{
"type": "object",
"properties": {
"status": {
"type": "string",
"const": "conditional_projection_not_realized"
},
"basis": {
"type": "string",
"const": "same_expected_successful_task_volume"
},
"currency": {
"type": "string",
"const": "USD"
},
"period": {
"type": "string",
"const": "month"
},
"baseline_launched_tasks_per_month": {
"type": "integer",
"minimum": 1,
"maximum": 1000000
},
"expected_successful_tasks_per_month": {
"type": "number",
"minimum": 0
},
"candidate_expected_launched_tasks_per_month": {
"type": "number",
"minimum": 0
},
"baseline_expected_cost_usd": {
"type": "number",
"minimum": 0
},
"candidate_expected_cost_usd": {
"type": "number",
"minimum": 0
},
"potential_cost_difference_usd": {
"type": "number",
"minimum": 0
},
"excludes": {
"type": "array",
"items": {
"type": "string"
},
"minItems": 1
},
"conditions": {
"type": "array",
"items": {
"type": "string"
},
"minItems": 1
}
},
"required": [
"status",
"basis",
"currency",
"period",
"baseline_launched_tasks_per_month",
"expected_successful_tasks_per_month",
"candidate_expected_launched_tasks_per_month",
"baseline_expected_cost_usd",
"candidate_expected_cost_usd",
"potential_cost_difference_usd",
"excludes",
"conditions"
],
"additionalProperties": true
},
{
"type": "null"
}
]
}
},
"required": [
"workflow",
"baseline",
"candidate",
"comparison",
"gates",
"decision",
"automatic_deployment_authorized",
"realized_savings_usd",
"opportunity_monthly"
],
"additionalProperties": true
},
"minItems": 1,
"maxItems": 1000
}
},
"required": [
"schema_version",
"purpose",
"data_provenance",
"config",
"input_attempt_records",
"experimental_bias_control",
"automatic_deployment_authorized",
"methodology",
"workflows"
],
"additionalProperties": true
}🟢analyze_agent_latency(runs, config)
Find slow agent workflows and retry overhead before scaling. Supply recorded attempts with task_id, workflow, variant, cost_usd, success and optional latency_ms. Returns P50/P95 of summed recorded attempt durations per task, retry counts, duration coverage and an optional P95 threshold check for each workflow and variant. Missing durations produce null percentiles, not zero latency. This measures supplied durations, not live end-to-end service latency. Free calculation; requires an active ALPNAI agent key; no payment or automatic deployment.
输入模式
{
"type": "object",
"properties": {
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"task_id": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units."
},
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"variant": {
"type": "string",
"enum": [
"baseline",
"candidate"
]
},
"cost_usd": {
"type": "number",
"minimum": 0,
"maximum": 10000,
"description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals."
},
"success": {
"type": "boolean"
},
"latency_ms": {
"type": "number",
"minimum": 0,
"maximum": 86400000,
"description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured."
}
},
"required": [
"task_id",
"workflow",
"variant",
"cost_usd",
"success"
],
"additionalProperties": false
},
"minItems": 1,
"maxItems": 1000
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [],
"additionalProperties": false
}
},
"required": [
"runs"
],
"additionalProperties": false,
"title": "ALPNAI recorded agent attempts",
"description": "Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
}输出模式
{
"type": "object",
"properties": {
"schema_version": {
"type": "string",
"const": "1.0.0"
},
"measurement": {
"type": "string",
"const": "summed_recorded_attempt_durations_per_task"
},
"percentile_method": {
"type": "string",
"const": "nearest_rank"
},
"groups": {
"type": "array",
"items": {
"type": "object",
"properties": {
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"variant": {
"type": "string",
"enum": [
"baseline",
"candidate"
]
},
"tasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"recorded_attempts": {
"type": "integer",
"minimum": 1,
"maximum": 1000
},
"retry_attempts": {
"type": "integer",
"minimum": 0,
"maximum": 999
},
"duration_coverage": {
"type": "number",
"minimum": 0,
"maximum": 1
},
"p50_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"p95_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"max_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"threshold_ms": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
{
"type": "null"
}
]
},
"within_p95_threshold": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
]
}
},
"required": [
"workflow",
"variant",
"tasks",
"recorded_attempts",
"retry_attempts",
"duration_coverage",
"p50_ms",
"p95_ms",
"max_ms",
"threshold_ms",
"within_p95_threshold"
],
"additionalProperties": true
},
"minItems": 1,
"maxItems": 1000
}
},
"required": [
"schema_version",
"measurement",
"percentile_method",
"groups"
],
"additionalProperties": true
}🟢check_agent_quality(runs, config)
Check whether a candidate agent workflow regresses before replacing the baseline. Supply baseline and candidate attempts on matching task IDs with client-provided success labels and costs. Returns observed success rates, task-set matching, sample-size, success, latency and cost gates, plus a decision such as collect_more_data, quality_regression or candidate_for_controlled_trial. Optional thresholds use config. It evaluates the recorded labels, not the correctness of answers or future performance. Free calculation; requires an active ALPNAI agent key; no payment or automatic deployment.
输入模式
{
"type": "object",
"properties": {
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"task_id": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units."
},
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"variant": {
"type": "string",
"enum": [
"baseline",
"candidate"
]
},
"cost_usd": {
"type": "number",
"minimum": 0,
"maximum": 10000,
"description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals."
},
"success": {
"type": "boolean"
},
"latency_ms": {
"type": "number",
"minimum": 0,
"maximum": 86400000,
"description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured."
}
},
"required": [
"task_id",
"workflow",
"variant",
"cost_usd",
"success"
],
"additionalProperties": false
},
"minItems": 1,
"maxItems": 1000
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [],
"additionalProperties": false
}
},
"required": [
"runs"
],
"additionalProperties": false,
"title": "ALPNAI recorded agent attempts",
"description": "Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
}输出模式
{
"type": "object",
"properties": {
"schema_version": {
"type": "string",
"const": "1.0.0"
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [
"minSamples",
"minSuccessRate",
"maxSuccessRateDrop"
],
"additionalProperties": true
},
"workflows": {
"type": "array",
"items": {
"type": "object",
"properties": {
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"comparison": {
"type": "object",
"properties": {
"matched_task_ids": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"baseline_only_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"candidate_only_tasks": {
"type": "integer",
"minimum": 0,
"maximum": 1000
},
"same_task_set": {
"type": "boolean"
}
},
"required": [
"matched_task_ids",
"baseline_only_tasks",
"candidate_only_tasks",
"same_task_set"
],
"additionalProperties": true
},
"gates": {
"type": "object",
"properties": {
"both_variants": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"minimum_distinct_tasks_per_variant": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"same_task_set": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"observed_success_rate": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"recorded_latency": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"lower_cost_per_successful_task": {
"type": "string",
"enum": [
"pass",
"fail",
"unknown",
"not_requested"
]
},
"experimental_bias_control": {
"type": "string",
"const": "unknown"
}
},
"required": [
"both_variants",
"minimum_distinct_tasks_per_variant",
"same_task_set",
"observed_success_rate",
"recorded_latency",
"lower_cost_per_successful_task",
"experimental_bias_control"
],
"additionalProperties": true
},
"decision": {
"type": "string",
"enum": [
"missing_comparison",
"collect_more_data",
"quality_regression",
"latency_data_required",
"latency_regression",
"no_economic_advantage",
"candidate_for_controlled_trial"
]
},
"baseline_success": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 1
},
{
"type": "null"
}
]
},
"candidate_success": {
"anyOf": [
{
"type": "number",
"minimum": 0,
"maximum": 1
},
{
"type": "null"
}
]
},
"automatic_deployment_authorized": {
"type": "boolean",
"const": false
}
},
"required": [
"workflow",
"comparison",
"gates",
"decision",
"baseline_success",
"candidate_success",
"automatic_deployment_authorized"
],
"additionalProperties": true
},
"minItems": 1,
"maxItems": 1000
}
},
"required": [
"schema_version",
"config",
"workflows"
],
"additionalProperties": true
}🟡save_project_report(request_id, title, input)
Compute and save an aggregate performance report in the fixed Projects destination authorized by the account owner. Uses the existing free or paid report allowance. Requires an active agent key and explicit owner write permission. Reuse request_id with identical content to retry within the same permission grant. No report reading, deletion, subscription or payment authorization. Send measurements without secrets.
输入模式
{
"type": "object",
"properties": {
"request_id": {
"type": "string",
"pattern": "^[0-9a-fA-F]{8}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{4}-[0-9a-fA-F]{12}$"
},
"title": {
"type": "string",
"minLength": 1,
"maxLength": 80
},
"input": {
"type": "object",
"properties": {
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"task_id": {
"type": "string",
"minLength": 1,
"maxLength": 128,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 128 UTF-16 code units."
},
"workflow": {
"type": "string",
"minLength": 1,
"maxLength": 80,
"pattern": "^(?!\\s)(?![\\s\\S]*\\s$)[^\\u0000-\\u001F\\u007F]+$",
"description": "Nonempty, no control characters or surrounding whitespace. The engine additionally enforces at most 80 UTF-16 code units."
},
"variant": {
"type": "string",
"enum": [
"baseline",
"candidate"
]
},
"cost_usd": {
"type": "number",
"minimum": 0,
"maximum": 10000,
"description": "USD cost of this attempt. At most six decimal places; the engine verifies whole micro-USD using floating-point tolerance. No multipleOf keyword is used, to avoid rejecting valid JSON decimals."
},
"success": {
"type": "boolean"
},
"latency_ms": {
"type": "number",
"minimum": 0,
"maximum": 86400000,
"description": "Recorded attempt duration in milliseconds; decimals are accepted. Omit if not measured."
}
},
"required": [
"task_id",
"workflow",
"variant",
"cost_usd",
"success"
],
"additionalProperties": false
},
"minItems": 1,
"maxItems": 1000
},
"config": {
"type": "object",
"properties": {
"minSamples": {
"type": "integer",
"minimum": 2,
"maximum": 500,
"default": 30
},
"minSuccessRate": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.95
},
"maxSuccessRateDrop": {
"type": "number",
"minimum": 0,
"maximum": 1,
"default": 0.02
},
"maxP95LatencyMs": {
"type": "number",
"minimum": 0,
"maximum": 86400000000
},
"monthlyTasks": {
"type": "integer",
"minimum": 1,
"maximum": 1000000,
"description": "Baseline logical tasks launched per month; the engine requires exactly one workflow if supplied."
}
},
"required": [],
"additionalProperties": false
}
},
"required": [
"runs"
],
"additionalProperties": false,
"title": "ALPNAI recorded agent attempts",
"description": "Recorded attempts, not prompts or secrets. Identical rows count as separate attempts. All three analysis tools accept this same input. The runtime validator additionally enforces six-decimal micro-USD precision, the same task_id belonging to one workflow, one workflow when monthlyTasks is given, and UTF-16 length limits. HTTP body limit: 512000 UTF-8 bytes."
}
},
"required": [
"request_id",
"title",
"input"
],
"additionalProperties": false
}🟡get_order(order_id)
Follow one existing commercial order owned by the authenticated agent, including after commerce is paused. May record a finalized blockchain receipt or release an expired, never-submitted reservation. Does not create an order, authorize spending or submit a payment. Pending results provide the same order ID and Retry-After; repeat get_order, never purchase again. Requires only the existing agent key.
输入模式
{
"type": "object",
"properties": {
"order_id": {
"type": "string",
"pattern": "^[-A-Za-z0-9]{8,80}$",
"description": "The original order_id returned by the purchase."
}
},
"required": [
"order_id"
],
"additionalProperties": false
}🔴purchase_snapshot(idempotency_key, since, mode, mandate_id)
Get Snapshot for 0.01 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.
输入模式
{
"type": "object",
"properties": {
"idempotency_key": {
"type": "string",
"pattern": "^[A-Za-z0-9_-]{8,100}$"
},
"since": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"mode": {
"type": "string",
"enum": [
"sandbox",
"live"
],
"default": "sandbox",
"description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow."
},
"mandate_id": {
"type": "string",
"pattern": "^[-a-zA-Z0-9]{8,80}$",
"description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers."
}
},
"required": [
"idempotency_key"
],
"additionalProperties": false
}🔴purchase_changes(idempotency_key, since, mode, mandate_id)
Get Change Set for 0.05 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.
输入模式
{
"type": "object",
"properties": {
"idempotency_key": {
"type": "string",
"pattern": "^[A-Za-z0-9_-]{8,100}$"
},
"since": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"mode": {
"type": "string",
"enum": [
"sandbox",
"live"
],
"default": "sandbox",
"description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow."
},
"mandate_id": {
"type": "string",
"pattern": "^[-a-zA-Z0-9]{8,80}$",
"description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers."
}
},
"required": [
"idempotency_key"
],
"additionalProperties": false
}🔴purchase_evidence(idempotency_key, since, mode, mandate_id)
Get Evidence Pack for 0.25 USDC. Default mode sandbox creates a test receipt and debits only the test budget. Explicit mode live uses x402 only when commerce, owner mandate and billing eligibility are enabled. An unpaid result contains PaymentRequired in structuredContent and content. Retry the same tool and idempotency key with the signed PaymentPayload in params._meta["x402/payment"]. Confirmed settlement is returned in result._meta["x402/payment-response"]. For a pending order, call get_order with the original order_id; do not submit another payment. REST continuation and PAYMENT-SIGNATURE headers remain supported. Never send a private key.
输入模式
{
"type": "object",
"properties": {
"idempotency_key": {
"type": "string",
"pattern": "^[A-Za-z0-9_-]{8,100}$"
},
"since": {
"type": "string",
"pattern": "^\\d{4}-\\d{2}-\\d{2}$"
},
"mode": {
"type": "string",
"enum": [
"sandbox",
"live"
],
"default": "sandbox",
"description": "Sandbox is the default. Live explicitly requests the existing authorized x402 purchase flow."
},
"mandate_id": {
"type": "string",
"pattern": "^[-a-zA-Z0-9]{8,80}$",
"description": "Owner-authorized commercial mandate ID. Alternatively send X-AlpNAI-Mandate in the MCP request headers."
}
},
"required": [
"idempotency_key"
],
"additionalProperties": false
}社区
证据