Model Ruler — AI Cost Calculators
Read-only AI/LLM cost tools with current MCP discovery and attributable Registry transport.
我该使用它吗
质量与安全性
发现(15)
- LOW在 self-host-breakeven-calculator 中
- LOW在 fine-tune-roi-calculator 中
- LOW在 agent-loop-cost-calculator 中
- LOW在 token-counter 中
- LOW在 provider-cost-calculator 中
- LOW在 self-host-breakeven-calculator 中
- LOW在 context-window-planner 中
- LOW在 quantization-calculator 中
- LOW在 fine-tune-roi-calculator 中
- LOW在 eval-cost-calculator 中
基于对工具定义和协议合规性的自动分析。
上下文开销
这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。
安装
一键安装
将以下内容添加到你的 `claude_desktop_config.json` 文件中:
{
"mcpServers": {
"model-ruler": {
"url": "https://modelruler.dev/mcp"
}
}
}远程端点
https://modelruler.dev/mcpstreamable-httphttps://modelruler.dev/mcp?dist=model_everywhere_registry_native_official_registrystreamable-http它能做什么
工具清单
工具(12)
🟢token-counter(text, tokenizer, expected_out_tokens)
Use when a user asks how many tokens a given text will consume, or needs to estimate prompt size before pricing a workload. Given text and tokenizer family, returns low/high token range and byte-level measurements.
输入模式
{
"type": "object",
"properties": {
"text": {
"type": "string",
"description": "Text content to count tokens for"
},
"tokenizer": {
"type": "string",
"enum": [
"cl100k_base",
"o200k_base",
"claude",
"gemini",
"llama3",
"mistral",
"default"
],
"description": "Tokenizer family (default: default)"
},
"expected_out_tokens": {
"type": "number",
"minimum": 0,
"description": "Expected output token budget (optional)"
}
},
"required": [
"text"
],
"examples": [
{
"text": "Summarize the Q3 board deck into five concise bullet points for the exec team."
}
]
}🟢provider-cost-calculator(tokens_in, tokens_out, calls_per_month, provider, model, ...)
Use when a user asks what an LLM workload costs on a specific provider/model, or wants to compare cost across providers. Given tokens per call and call volume, returns monthly cost plus a tier comparison table.
输入模式
{
"type": "object",
"properties": {
"tokens_in": {
"type": "number",
"minimum": 0,
"description": "Input tokens per call"
},
"tokens_out": {
"type": "number",
"minimum": 0,
"description": "Output tokens per call"
},
"calls_per_month": {
"type": "number",
"minimum": 1,
"description": "Monthly call volume"
},
"provider": {
"type": "string",
"description": "Target provider (e.g. anthropic, openai, together)"
},
"model": {
"type": "string",
"description": "Target model (e.g. claude-sonnet-4-6)"
},
"include_comparison": {
"type": "boolean",
"description": "Include tier comparison table (default true)"
}
},
"required": [
"tokens_in",
"tokens_out"
],
"examples": [
{
"tokens_in": 1200,
"tokens_out": 400,
"calls_per_month": 100000
}
]
}🟢self-host-breakeven-calculator(monthly_tokens, api_cost_per_1m_out, gpu_provider, gpu_type, utilization_pct, ...)
Use when a user is deciding between API usage and self-hosted GPU inference at a given volume. Returns breakeven token volume, monthly cost comparison, and go/no-go recommendation.
输入模式
{
"type": "object",
"properties": {
"monthly_tokens": {
"type": "number",
"minimum": 0,
"description": "Monthly output token volume"
},
"api_cost_per_1m_out": {
"type": "number",
"minimum": 0,
"description": "Current API output cost per 1M tokens"
},
"gpu_provider": {
"type": "string",
"description": "GPU provider (e.g. runpod, modal)"
},
"gpu_type": {
"type": "string",
"description": "GPU type (e.g. h100, a100-80gb)"
},
"utilization_pct": {
"type": "number",
"minimum": 1,
"maximum": 100,
"description": "Expected GPU utilization % (default 60)"
},
"operational_overhead_pct": {
"type": "number",
"minimum": 0,
"description": "Ops overhead % on GPU cost (default 40)"
}
},
"required": [
"monthly_tokens"
],
"examples": [
{
"monthly_tokens": 750000000
}
]
}🟢context-window-planner(doc_tokens, overhead_tokens, expected_out_tokens, model_context_window, strategy_hint)
Use when a user needs to know whether a document plus prompt plus output fits within a model's context window, or wants a strategy recommendation (truncate/summarize/rag/chunk).
输入模式
{
"type": "object",
"properties": {
"doc_tokens": {
"type": "number",
"minimum": 0,
"description": "Primary document/content tokens"
},
"overhead_tokens": {
"type": "number",
"minimum": 0,
"description": "System prompt + few-shot + history (default 2000)"
},
"expected_out_tokens": {
"type": "number",
"minimum": 0,
"description": "Reserved output budget (default 1000)"
},
"model_context_window": {
"type": "number",
"minimum": 1,
"description": "Target model context size"
},
"strategy_hint": {
"type": "string",
"enum": [
"truncate",
"summarize",
"rag",
"chunk"
],
"description": "Preferred strategy (optional)"
}
},
"required": [
"doc_tokens",
"model_context_window"
],
"examples": [
{
"doc_tokens": 18000,
"model_context_window": 200000
}
]
}🟢quantization-calculator(params_billions, precision_from, precision_to, kv_cache_tokens, batch_size)
Use when a user is planning to quantize an LLM to fit on smaller hardware. Given parameter count and precision transition, returns VRAM requirement, speedup estimate, and approximate quality delta.
输入模式
{
"type": "object",
"properties": {
"params_billions": {
"type": "number",
"minimum": 0.1,
"description": "Model parameter count in billions"
},
"precision_from": {
"type": "string",
"enum": [
"fp32",
"fp16",
"bf16",
"fp8",
"int8",
"int6",
"int5",
"int4",
"int3",
"int2",
"int1"
],
"description": "Starting precision (default bf16)"
},
"precision_to": {
"type": "string",
"enum": [
"fp32",
"fp16",
"bf16",
"fp8",
"int8",
"int6",
"int5",
"int4",
"int3",
"int2",
"int1"
],
"description": "Target precision (default int4)"
},
"kv_cache_tokens": {
"type": "number",
"minimum": 0,
"description": "Max KV cache tokens (default 8192)"
},
"batch_size": {
"type": "number",
"minimum": 1,
"description": "Serving batch size (default 1)"
}
},
"required": [
"params_billions"
],
"examples": [
{
"params_billions": 8
}
]
}🟢fine-tune-roi-calculator(train_tokens, train_cost_per_1m, base_inference_cost_1m, finetuned_inference_cost_1m, monthly_inference_tokens, ...)
Use when a user is considering fine-tuning vs prompt engineering. Returns training cost, monthly inference savings, months-to-ROI, and breakeven volume.
输入模式
{
"type": "object",
"properties": {
"train_tokens": {
"type": "number",
"minimum": 0,
"description": "Training tokens (dataset × epochs)"
},
"train_cost_per_1m": {
"type": "number",
"minimum": 0,
"description": "Training cost per 1M tokens"
},
"base_inference_cost_1m": {
"type": "number",
"minimum": 0,
"description": "Baseline API output cost per 1M"
},
"finetuned_inference_cost_1m": {
"type": "number",
"minimum": 0,
"description": "Fine-tuned inference cost per 1M"
},
"monthly_inference_tokens": {
"type": "number",
"minimum": 0,
"description": "Expected monthly inference volume (output tokens)"
},
"prompt_reduction_pct": {
"type": "number",
"minimum": 0,
"maximum": 100,
"description": "Prompt size reduction % from eliminating few-shot (default 0)"
}
},
"required": [
"train_tokens",
"monthly_inference_tokens"
],
"examples": [
{
"train_tokens": 50000000,
"monthly_inference_tokens": 300000000
}
]
}🟢eval-cost-calculator(samples, models, trials_per_sample, avg_tokens_in, avg_tokens_out, ...)
Use when a user needs to budget an LLM evaluation run. Given samples/models/trials, returns total cost, per-run cost, and parallel time estimate.
输入模式
{
"type": "object",
"properties": {
"samples": {
"type": "number",
"minimum": 1,
"description": "Number of eval samples"
},
"models": {
"type": "number",
"minimum": 1,
"description": "Candidate models (default 1)"
},
"trials_per_sample": {
"type": "number",
"minimum": 1,
"description": "Repeats per sample (default 1)"
},
"avg_tokens_in": {
"type": "number",
"minimum": 0,
"description": "Avg input tokens per sample (default 2000)"
},
"avg_tokens_out": {
"type": "number",
"minimum": 0,
"description": "Avg output tokens per sample (default 500)"
},
"provider": {
"type": "string",
"description": "Eval provider (default anthropic)"
},
"model": {
"type": "string",
"description": "Eval model (default claude-sonnet-4-6)"
},
"judge_enabled": {
"type": "boolean",
"description": "Enable LLM-as-judge second pass (default false)"
},
"judge_tokens_in": {
"type": "number",
"minimum": 0,
"description": "Judge input tokens (default 1500)"
},
"judge_tokens_out": {
"type": "number",
"minimum": 0,
"description": "Judge output tokens (default 200)"
}
},
"required": [
"samples"
],
"examples": [
{
"samples": 2000
}
]
}🟢observability-cost-calculator(requests_per_day, avg_log_bytes, retention_days, provider)
Use when a user needs to budget LLM observability tooling. Returns monthly cost at given request volume with retention adjustment.
输入模式
{
"type": "object",
"properties": {
"requests_per_day": {
"type": "number",
"minimum": 0,
"description": "Average daily LLM requests"
},
"avg_log_bytes": {
"type": "number",
"minimum": 0,
"description": "Avg payload bytes per traced request (default 4096)"
},
"retention_days": {
"type": "number",
"minimum": 1,
"description": "Retention in days (default 30)"
},
"provider": {
"type": "string",
"enum": [
"langsmith",
"langfuse",
"helicone",
"portkey",
"arize",
"braintrust"
],
"description": "Observability provider"
}
},
"required": [
"requests_per_day",
"provider"
],
"examples": [
{
"requests_per_day": 50000,
"provider": "langsmith"
}
]
}🟢rag-pipeline-cost-calculator(queries_per_day, corpus_tokens, chunk_size_tokens, chunks_retrieved, embedding_provider, ...)
Use when a user needs end-to-end RAG cost estimation (embedding + vector store + generation). Returns monthly cost with breakdown and dominant-component identification.
输入模式
{
"type": "object",
"properties": {
"queries_per_day": {
"type": "number",
"minimum": 0,
"description": "User query volume per day"
},
"corpus_tokens": {
"type": "number",
"minimum": 0,
"description": "Indexed corpus size in tokens"
},
"chunk_size_tokens": {
"type": "number",
"minimum": 64,
"description": "Avg chunk size (default 512)"
},
"chunks_retrieved": {
"type": "number",
"minimum": 1,
"description": "Chunks per query (default 5)"
},
"embedding_provider": {
"type": "string",
"enum": [
"openai-3-small",
"openai-3-large",
"voyage-3",
"voyage-3-lite",
"cohere-embed-v3",
"google-text-embed-004"
],
"description": "Embedding model"
},
"vector_store": {
"type": "string",
"enum": [
"pinecone",
"weaviate",
"qdrant",
"chroma",
"turbopuffer"
],
"description": "Vector store"
},
"generator_provider": {
"type": "string",
"description": "Generator LLM provider (default anthropic)"
},
"generator_model": {
"type": "string",
"description": "Generator model (default claude-sonnet-4-6)"
},
"question_tokens": {
"type": "number",
"minimum": 0,
"description": "Avg question tokens (default 100)"
},
"answer_tokens": {
"type": "number",
"minimum": 0,
"description": "Avg answer tokens (default 400)"
},
"reindex_fraction_per_month": {
"type": "number",
"minimum": 0,
"maximum": 1,
"description": "Fraction of corpus re-embedded per month (default 0.1)"
}
},
"required": [
"queries_per_day",
"corpus_tokens"
],
"examples": [
{
"queries_per_day": 5000,
"corpus_tokens": 25000000
}
]
}🟢automation-cost-calculator(workflow_runs_per_month, billable_steps_per_run, make_modules_per_run)
Use for generic or non-agent recurring workflow platform billing across Zapier task billing, Make credit billing, and n8n execution billing. Use agent-workflow-cost-calculator instead when the workflow is explicitly an AI agent with app actions or MCP tool calls. Use agent-loop-cost-calculator instead when the question is about LLM inference/reasoning cost per successful agent task.
输入模式
{
"type": "object",
"properties": {
"workflow_runs_per_month": {
"type": "number",
"minimum": 0,
"description": "Workflow/scenario runs per month"
},
"billable_steps_per_run": {
"type": "number",
"minimum": 1,
"description": "Billable actions/modules per run (default 5)"
},
"make_modules_per_run": {
"type": "number",
"minimum": 1,
"description": "Make module actions per scenario run; defaults to billable_steps_per_run"
}
},
"required": [
"workflow_runs_per_month"
],
"examples": [
{
"workflow_runs_per_month": 120000
}
]
}🟢agent-workflow-cost-calculator(agent_runs_per_month, app_actions_per_run, mcp_tool_calls_per_run, make_modules_per_run)
Use for automation-platform/iPaaS cost of an AI agent workflow, including app-action fan-out and MCP tool-call accounting, explicitly excluding LLM token spend. Use automation-cost-calculator for generic non-agent workflow billing. Use agent-loop-cost-calculator for LLM inference/reasoning cost per successful multi-step agent task.
输入模式
{
"type": "object",
"properties": {
"agent_runs_per_month": {
"type": "number",
"minimum": 0,
"description": "Agent workflow runs per month"
},
"app_actions_per_run": {
"type": "number",
"minimum": 0,
"description": "Downstream app actions per agent run (default 3)"
},
"mcp_tool_calls_per_run": {
"type": "number",
"minimum": 0,
"description": "MCP tool calls per agent run (default 1)"
},
"make_modules_per_run": {
"type": "number",
"minimum": 1,
"description": "Make modules per run; defaults to app actions + MCP calls"
}
},
"required": [
"agent_runs_per_month"
],
"examples": [
{
"agent_runs_per_month": 20000
}
]
}🟢agent-loop-cost-calculator(tasks_per_month, steps_per_task, tool_calls_per_step, avg_tokens_in_per_turn, avg_tokens_out_per_turn, ...)
Use for LLM inference/reasoning cost per successful multi-step agent task, including failure overhead and context growth across turns. Use agent-workflow-cost-calculator for iPaaS/app-action/MCP orchestration fees. Use automation-cost-calculator for generic Zapier/Make/n8n workflow billing.
输入模式
{
"type": "object",
"properties": {
"tasks_per_month": {
"type": "number",
"minimum": 0,
"description": "Tasks attempted per month"
},
"steps_per_task": {
"type": "number",
"minimum": 1,
"description": "Avg reasoning steps per task (default 5)"
},
"tool_calls_per_step": {
"type": "number",
"minimum": 0,
"description": "Avg tool invocations per step (default 2)"
},
"avg_tokens_in_per_turn": {
"type": "number",
"minimum": 0,
"description": "Avg input tokens per turn (default 3000)"
},
"avg_tokens_out_per_turn": {
"type": "number",
"minimum": 0,
"description": "Avg output tokens per turn (default 400)"
},
"success_rate": {
"type": "number",
"minimum": 0.01,
"maximum": 1,
"description": "Task success rate 0-1 (default 0.7)"
},
"provider": {
"type": "string",
"description": "LLM provider (default anthropic)"
},
"model": {
"type": "string",
"description": "LLM model (default claude-sonnet-4-6)"
},
"context_growth_factor": {
"type": "number",
"minimum": 1,
"description": "Multiplier on input tokens as conversation grows (default 1.4)"
}
},
"required": [
"tasks_per_month"
],
"examples": [
{
"tasks_per_month": 60000
}
]
}社区
证据