FastGPU

Compare live GPU cloud rental prices and match workloads to the cheapest provider.

我该使用它吗

质量与安全性

A
描述质量
100%
模式完整度
90%
命名质量
90%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~1,074token 数(工具定义)
~5.2 KB典型响应大小
对注意力有中等影响(占 128k 上下文窗口的 0.84%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "fastgpu": {
      "url": "https://fastgpu.co/api/mcp"
    }
  }
}

远程端点

https://fastgpu.co/api/mcpstreamable-http

它能做什么

工具清单

工具(2)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢list_gpu_prices(vendor, tier)

One entry per GPU model with the current cheapest live rental price across the whole market (RunPod, Vast.ai, Lambda, hyperscalers and more). No key required. Use this to compare GPU prices.

输入模式

{
  "type": "object",
  "properties": {
    "vendor": {
      "type": "string",
      "description": "Filter by GPU vendor.",
      "enum": [
        "NVIDIA",
        "AMD"
      ]
    },
    "tier": {
      "type": "string",
      "description": "Filter by tier.",
      "enum": [
        "flagship",
        "datacenter",
        "prosumer",
        "entry"
      ]
    }
  }
}

输出模式

{
  "type": "object",
  "properties": {
    "updated_at": {
      "type": [
        "string",
        "null"
      ]
    },
    "stale": {
      "type": "boolean"
    },
    "count": {
      "type": "integer"
    },
    "gpus": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "gpu": {
            "type": "string"
          },
          "vram_gb": {
            "type": [
              "number",
              "null"
            ]
          },
          "vendor": {
            "type": [
              "string",
              "null"
            ]
          },
          "arch": {
            "type": [
              "string",
              "null"
            ]
          },
          "tier": {
            "type": [
              "string",
              "null"
            ]
          },
          "cheapest_usd_hr": {
            "type": [
              "number",
              "null"
            ]
          },
          "cheapest_provider": {
            "type": [
              "string",
              "null"
            ]
          },
          "cheapest_min_gpu_count": {
            "type": [
              "integer",
              "null"
            ]
          },
          "provider_count": {
            "type": "integer"
          },
          "offer_count": {
            "type": "integer"
          },
          "url": {
            "type": "string"
          }
        }
      }
    }
  },
  "required": [
    "count",
    "gpus"
  ]
}
🟢match_workload(query, model, params_b, vram_gb, task, ...)

The routing DECISION: describe a job (a model, size, or GPU need) and get the ranked, reasoned recommendation for the cheapest place to run it across the live market, with the required VRAM, GPU count, effective $/hr, and how much cheaper it is than a hyperscaler. No key required. Results mirror the site and apply a small, disclosed partner tie-break between otherwise-equal offers (each match reports partner true/false).

输入模式

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": "Plain-language job, e.g. \"cheapest to serve Llama 3 70B\" or \"2x H100 for fine-tuning\". Provide this OR a structured spec below."
    },
    "model": {
      "type": "string",
      "description": "Open model name to size against, e.g. \"Llama 3 70B\", \"Qwen 72B\", \"Mixtral\"."
    },
    "params_b": {
      "type": "number",
      "description": "Model size in billions of parameters when no exact model is named."
    },
    "vram_gb": {
      "type": "integer",
      "description": "Rough VRAM the job needs, in GB, if you already know it."
    },
    "task": {
      "type": "string",
      "description": "What the job does.",
      "enum": [
        "inference",
        "finetune-lora",
        "finetune-full",
        "generate",
        "transcribe",
        "embed"
      ]
    },
    "precision": {
      "type": "string",
      "description": "Numeric precision to size the model at.",
      "enum": [
        "fp16",
        "int8",
        "int4"
      ]
    },
    "gpu_count": {
      "type": "integer",
      "description": "Exact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing."
    },
    "budget_usd_hr": {
      "type": "number",
      "description": "Only recommend configs at or under this hourly budget."
    },
    "region": {
      "type": "string",
      "description": "Restrict to a data-residency region.",
      "enum": [
        "US",
        "EU",
        "ASIA"
      ]
    },
    "spot": {
      "type": "string",
      "description": "Set true to include interruptible spot capacity for a cheaper rate.",
      "enum": [
        "true",
        "false"
      ]
    },
    "reserved": {
      "type": "string",
      "description": "Set true to include reserved / committed-term capacity for a lower rate.",
      "enum": [
        "true",
        "false"
      ]
    }
  }
}

输出模式

{
  "type": "object",
  "properties": {
    "updated_at": {
      "type": [
        "string",
        "null"
      ]
    },
    "stale": {
      "type": "boolean"
    },
    "workload": {
      "type": "object",
      "properties": {
        "interpretation": {
          "type": "string"
        },
        "identified": {
          "type": "boolean"
        },
        "task": {
          "type": "string"
        },
        "precision": {
          "type": "string"
        },
        "vram_required_gb": {
          "type": "number"
        },
        "gpu_count": {
          "type": [
            "integer",
            "null"
          ]
        }
      }
    },
    "hero": {
      "type": [
        "object",
        "null"
      ],
      "properties": {
        "hyperscaler_ceiling": {
          "type": [
            "string",
            "null"
          ]
        },
        "savings_pct": {
          "type": [
            "number",
            "null"
          ]
        },
        "same_model": {
          "type": [
            "boolean",
            "null"
          ]
        }
      }
    },
    "count": {
      "type": "integer"
    },
    "matches": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "gpu": {
            "type": "string"
          },
          "gpu_count": {
            "type": "integer"
          },
          "effective_usd_hr": {
            "type": "number"
          },
          "monthly_usd": {
            "type": "number"
          },
          "provider": {
            "type": "string"
          },
          "provider_label": {
            "type": [
              "string",
              "null"
            ]
          },
          "offer_type": {
            "type": "string"
          },
          "reliability": {
            "type": "string"
          },
          "fits_single_card": {
            "type": "boolean"
          },
          "over_budget": {
            "type": "boolean"
          },
          "tokens_per_sec": {
            "type": [
              "number",
              "null"
            ]
          },
          "usd_per_million_tokens": {
            "type": [
              "number",
              "null"
            ]
          },
          "score": {
            "type": [
              "number",
              "null"
            ]
          },
          "partner": {
            "type": "boolean"
          },
          "reason": {
            "type": "string"
          },
          "url": {
            "type": "string"
          }
        }
      }
    }
  },
  "required": [
    "count",
    "matches"
  ]
}

推荐提示词

list_items
List all [items] available in FastGPU
预期工具: list_gpu_prices
browse_collection
Show me the [collection] from FastGPU
预期工具: list_gpu_prices

社区

评价此服务器

证据

最近观测

已验证未记录版本2 个工具
已验证未记录版本2 个工具