deepinfra

Run DeepInfra inference, list models and read account rate limits.

我该使用它吗

质量与安全性

A
描述质量
98%
模式完整度
82%
命名质量
80%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~2,232token 数(工具定义)
~804 B典型响应大小
对注意力有中等影响(占 128k 上下文窗口的 1.74%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "deepinfra": {
      "url": "https://deepinfra.usefulapi.io/mcp"
    }
  }
}

远程端点

https://deepinfra.usefulapi.io/mcpstreamable-http

它能做什么

工具清单

工具(18)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢deepinfra_get_account(checklist)

Get the current DeepInfra account / user details (identity, email, quotas). DeepInfra REST: GET /v1/me.

输入模式

{
  "type": "object",
  "properties": {
    "checklist": {
      "description": "Include the onboarding checklist state in the response.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_rate_limit

Get the account's current rate limits (per-model / per-endpoint request and token limits). DeepInfra REST: GET /v1/me/rate_limit.

输入模式

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_list_api_tokens

List the account's API tokens (metadata only, not secret values). DeepInfra REST: GET /v1/api-tokens.

输入模式

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_usage(from, to)

Get spend / usage for a billing period. DeepInfra REST: GET /payment/usage.

输入模式

{
  "type": "object",
  "properties": {
    "from": {
      "type": "string",
      "minLength": 1,
      "description": "Period start (required). Format 'YYYY.MM', or 'current', or 'current(-N)' for N months back, or a unix timestamp in seconds."
    },
    "to": {
      "description": "Period end (same format as `from`); defaults to the current period.",
      "type": "string"
    }
  },
  "required": [
    "from"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_usage_tokens(from, to)

Get per-model token usage for a period. DeepInfra REST: GET /payment/usage/tokens.

输入模式

{
  "type": "object",
  "properties": {
    "from": {
      "type": "string",
      "minLength": 1,
      "description": "Period start (required). Format 'YYYY.MM', or 'current', or 'current(-N)' for N months back, or a unix timestamp in seconds."
    },
    "to": {
      "description": "Period end (same format as `from`); defaults to the current period.",
      "type": "string"
    }
  },
  "required": [
    "from"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_usage_rent(from, to)

Get GPU rental (dedicated hardware) usage for a time range. DeepInfra REST: GET /payment/usage/rent.

输入模式

{
  "type": "object",
  "properties": {
    "from": {
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991,
      "description": "Range start as a unix timestamp in seconds (required)."
    },
    "to": {
      "description": "Range end as a unix timestamp in seconds; defaults to now.",
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "from"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_list_invoices(limit, starting_after, invoice_type)

List billing invoices for the account. DeepInfra REST: GET /payment/invoices.

输入模式

{
  "type": "object",
  "properties": {
    "limit": {
      "description": "Max number of invoices to return.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "starting_after": {
      "description": "Cursor — return invoices after this invoice id (pagination).",
      "type": "string"
    },
    "invoice_type": {
      "description": "Filter by invoice type.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_list_deployments(status)

List the account's dedicated deployments. DeepInfra REST: GET /deploy/list.

输入模式

{
  "type": "object",
  "properties": {
    "status": {
      "description": "Comma-separated statuses to filter by, e.g. 'running,initializing'.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_deployment(deploy_id)

Get one dedicated deployment by id (config + current status). DeepInfra REST: GET /deploy/{deploy_id}.

输入模式

{
  "type": "object",
  "properties": {
    "deploy_id": {
      "type": "string",
      "minLength": 1,
      "description": "The deployment id (required)."
    }
  },
  "required": [
    "deploy_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_deployment_stats(deploy_id, from, to)

Get time-series stats (throughput / latency / replicas) for a dedicated deployment. DeepInfra REST: GET /deploy/{deploy_id}/stats2.

输入模式

{
  "type": "object",
  "properties": {
    "deploy_id": {
      "type": "string",
      "minLength": 1,
      "description": "The deployment id (required)."
    },
    "from": {
      "type": "string",
      "minLength": 1,
      "description": "Range start (required) — a unix timestamp or a relative expression like 'now-5h' (units: s, m, h, d, w)."
    },
    "to": {
      "description": "Range end — unix timestamp or relative expression; defaults to now.",
      "type": "string"
    }
  },
  "required": [
    "deploy_id",
    "from"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_gpu_availability(source, base_model)

Get GPU availability for LLM deployments (which hardware can currently be provisioned). DeepInfra REST: GET /deploy/llm/gpu_availability.

输入模式

{
  "type": "object",
  "properties": {
    "source": {
      "description": "Filter by source.",
      "type": "string"
    },
    "base_model": {
      "description": "Filter by base model.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_list_models

List the DeepInfra model catalog (all available models). DeepInfra REST: GET /models/list.

输入模式

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_model(model_name, version)

Get one model's catalog entry (pricing, context length, capabilities). DeepInfra REST: GET /models/{model_name}.

输入模式

{
  "type": "object",
  "properties": {
    "model_name": {
      "type": "string",
      "minLength": 1,
      "description": "The model name (required). May contain a slash, e.g. 'meta-llama/Meta-Llama-3-8B' — the slash is preserved in the path."
    },
    "version": {
      "description": "A specific model version (defaults to latest).",
      "type": "string"
    }
  },
  "required": [
    "model_name"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_hardware(model)

Get the hardware options / GPU configuration available for a given model. DeepInfra REST: GET /v2/hardware.

输入模式

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "The model name to query hardware for (required)."
    }
  },
  "required": [
    "model"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_query_logs(deploy_id, from, to, limit)

Query inference request logs for a dedicated deployment over a time window. DeepInfra REST: GET /v1/logs/query.

输入模式

{
  "type": "object",
  "properties": {
    "deploy_id": {
      "type": "string",
      "minLength": 1,
      "description": "The deployment id to query logs for (required)."
    },
    "from": {
      "description": "Window start — fractional-seconds unix timestamp, inclusive.",
      "type": "string"
    },
    "to": {
      "description": "Window end — fractional-seconds unix timestamp, exclusive.",
      "type": "string"
    },
    "limit": {
      "description": "Max log lines to return (default 100, range 1..1000).",
      "type": "integer",
      "minimum": 1,
      "maximum": 1000
    }
  },
  "required": [
    "deploy_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢deepinfra_get_live_metrics

Get global live inference metrics across the account (real-time throughput / activity). DeepInfra REST: GET /v1/metrics/live.

输入模式

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴deepinfra_start_deployment(deploy_id)

Start (resume) a dedicated deployment. WARNING: this resumes a dedicated deployment and may incur GPU charges while it runs. DeepInfra REST: POST /deploy/{deploy_id}/start.

输入模式

{
  "type": "object",
  "properties": {
    "deploy_id": {
      "type": "string",
      "minLength": 1,
      "description": "The deployment id to start (required)."
    }
  },
  "required": [
    "deploy_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴deepinfra_stop_deployment(deploy_id)

Stop (pause) a running dedicated deployment. This halts inference and stops accruing GPU charges for it. DeepInfra REST: POST /deploy/{deploy_id}/stop.

输入模式

{
  "type": "object",
  "properties": {
    "deploy_id": {
      "type": "string",
      "minLength": 1,
      "description": "The deployment id to stop (required)."
    }
  },
  "required": [
    "deploy_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

社区

评价此服务器

证据

最近观测

已验证未记录版本18 个工具