deepinfra
Run DeepInfra inference, list models and read account rate limits.
Should I use this
Quality & Safety
Based on automated analysis of tool definitions and protocol compliance.
Context Cost
This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.
Install
One-Click Install
Add this to your `claude_desktop_config.json` file:
{
"mcpServers": {
"deepinfra": {
"url": "https://deepinfra.usefulapi.io/mcp"
}
}
}Remote endpoints
https://deepinfra.usefulapi.io/mcpstreamable-httpWhat it can do
Tool inventory
Tools (18)
🟢deepinfra_get_account(checklist)
Get the current DeepInfra account / user details (identity, email, quotas). DeepInfra REST: GET /v1/me.
Input Schema
{
"type": "object",
"properties": {
"checklist": {
"description": "Include the onboarding checklist state in the response.",
"type": "boolean"
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_rate_limit
Get the account's current rate limits (per-model / per-endpoint request and token limits). DeepInfra REST: GET /v1/me/rate_limit.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_list_api_tokens
List the account's API tokens (metadata only, not secret values). DeepInfra REST: GET /v1/api-tokens.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_usage(from, to)
Get spend / usage for a billing period. DeepInfra REST: GET /payment/usage.
Input Schema
{
"type": "object",
"properties": {
"from": {
"type": "string",
"minLength": 1,
"description": "Period start (required). Format 'YYYY.MM', or 'current', or 'current(-N)' for N months back, or a unix timestamp in seconds."
},
"to": {
"description": "Period end (same format as `from`); defaults to the current period.",
"type": "string"
}
},
"required": [
"from"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_usage_tokens(from, to)
Get per-model token usage for a period. DeepInfra REST: GET /payment/usage/tokens.
Input Schema
{
"type": "object",
"properties": {
"from": {
"type": "string",
"minLength": 1,
"description": "Period start (required). Format 'YYYY.MM', or 'current', or 'current(-N)' for N months back, or a unix timestamp in seconds."
},
"to": {
"description": "Period end (same format as `from`); defaults to the current period.",
"type": "string"
}
},
"required": [
"from"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_usage_rent(from, to)
Get GPU rental (dedicated hardware) usage for a time range. DeepInfra REST: GET /payment/usage/rent.
Input Schema
{
"type": "object",
"properties": {
"from": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "Range start as a unix timestamp in seconds (required)."
},
"to": {
"description": "Range end as a unix timestamp in seconds; defaults to now.",
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991
}
},
"required": [
"from"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_list_invoices(limit, starting_after, invoice_type)
List billing invoices for the account. DeepInfra REST: GET /payment/invoices.
Input Schema
{
"type": "object",
"properties": {
"limit": {
"description": "Max number of invoices to return.",
"type": "integer",
"minimum": 1,
"maximum": 9007199254740991
},
"starting_after": {
"description": "Cursor — return invoices after this invoice id (pagination).",
"type": "string"
},
"invoice_type": {
"description": "Filter by invoice type.",
"type": "string"
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_list_deployments(status)
List the account's dedicated deployments. DeepInfra REST: GET /deploy/list.
Input Schema
{
"type": "object",
"properties": {
"status": {
"description": "Comma-separated statuses to filter by, e.g. 'running,initializing'.",
"type": "string"
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_deployment(deploy_id)
Get one dedicated deployment by id (config + current status). DeepInfra REST: GET /deploy/{deploy_id}.
Input Schema
{
"type": "object",
"properties": {
"deploy_id": {
"type": "string",
"minLength": 1,
"description": "The deployment id (required)."
}
},
"required": [
"deploy_id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_deployment_stats(deploy_id, from, to)
Get time-series stats (throughput / latency / replicas) for a dedicated deployment. DeepInfra REST: GET /deploy/{deploy_id}/stats2.
Input Schema
{
"type": "object",
"properties": {
"deploy_id": {
"type": "string",
"minLength": 1,
"description": "The deployment id (required)."
},
"from": {
"type": "string",
"minLength": 1,
"description": "Range start (required) — a unix timestamp or a relative expression like 'now-5h' (units: s, m, h, d, w)."
},
"to": {
"description": "Range end — unix timestamp or relative expression; defaults to now.",
"type": "string"
}
},
"required": [
"deploy_id",
"from"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_gpu_availability(source, base_model)
Get GPU availability for LLM deployments (which hardware can currently be provisioned). DeepInfra REST: GET /deploy/llm/gpu_availability.
Input Schema
{
"type": "object",
"properties": {
"source": {
"description": "Filter by source.",
"type": "string"
},
"base_model": {
"description": "Filter by base model.",
"type": "string"
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_list_models
List the DeepInfra model catalog (all available models). DeepInfra REST: GET /models/list.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_model(model_name, version)
Get one model's catalog entry (pricing, context length, capabilities). DeepInfra REST: GET /models/{model_name}.
Input Schema
{
"type": "object",
"properties": {
"model_name": {
"type": "string",
"minLength": 1,
"description": "The model name (required). May contain a slash, e.g. 'meta-llama/Meta-Llama-3-8B' — the slash is preserved in the path."
},
"version": {
"description": "A specific model version (defaults to latest).",
"type": "string"
}
},
"required": [
"model_name"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_hardware(model)
Get the hardware options / GPU configuration available for a given model. DeepInfra REST: GET /v2/hardware.
Input Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"minLength": 1,
"description": "The model name to query hardware for (required)."
}
},
"required": [
"model"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_query_logs(deploy_id, from, to, limit)
Query inference request logs for a dedicated deployment over a time window. DeepInfra REST: GET /v1/logs/query.
Input Schema
{
"type": "object",
"properties": {
"deploy_id": {
"type": "string",
"minLength": 1,
"description": "The deployment id to query logs for (required)."
},
"from": {
"description": "Window start — fractional-seconds unix timestamp, inclusive.",
"type": "string"
},
"to": {
"description": "Window end — fractional-seconds unix timestamp, exclusive.",
"type": "string"
},
"limit": {
"description": "Max log lines to return (default 100, range 1..1000).",
"type": "integer",
"minimum": 1,
"maximum": 1000
}
},
"required": [
"deploy_id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢deepinfra_get_live_metrics
Get global live inference metrics across the account (real-time throughput / activity). DeepInfra REST: GET /v1/metrics/live.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🔴deepinfra_start_deployment(deploy_id)
Start (resume) a dedicated deployment. WARNING: this resumes a dedicated deployment and may incur GPU charges while it runs. DeepInfra REST: POST /deploy/{deploy_id}/start.
Input Schema
{
"type": "object",
"properties": {
"deploy_id": {
"type": "string",
"minLength": 1,
"description": "The deployment id to start (required)."
}
},
"required": [
"deploy_id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🔴deepinfra_stop_deployment(deploy_id)
Stop (pause) a running dedicated deployment. This halts inference and stops accruing GPU charges for it. DeepInfra REST: POST /deploy/{deploy_id}/stop.
Input Schema
{
"type": "object",
"properties": {
"deploy_id": {
"type": "string",
"minLength": 1,
"description": "The deployment id to stop (required)."
}
},
"required": [
"deploy_id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Community
Evidence