FastGPU

Compare live GPU cloud rental prices and match workloads to the cheapest provider.

Should I use this

Quality & Safety

A
Description quality
100%
Schema completeness
90%
Naming quality
90%
Poisoning risk
100%
Permission match
100%
Protocol compliance
100%

Based on automated analysis of tool definitions and protocol compliance.

Context Cost

~1,074Tokens (tool definitions)
~5.2 KBTypical response size
Moderate attention impact (0.84% of 128k context)

This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.

Install

One-Click Install

Add this to your `claude_desktop_config.json` file:

{
  "mcpServers": {
    "fastgpu": {
      "url": "https://fastgpu.co/api/mcp"
    }
  }
}

Remote endpoints

https://fastgpu.co/api/mcpstreamable-http

What it can do

Tool inventory

Tools (2)

🟢 Read-only🟡 Write🔴 Delete⚪ Unknown
🟢list_gpu_prices(vendor, tier)

One entry per GPU model with the current cheapest live rental price across the whole market (RunPod, Vast.ai, Lambda, hyperscalers and more). No key required. Use this to compare GPU prices.

Input Schema

{
  "type": "object",
  "properties": {
    "vendor": {
      "type": "string",
      "description": "Filter by GPU vendor.",
      "enum": [
        "NVIDIA",
        "AMD"
      ]
    },
    "tier": {
      "type": "string",
      "description": "Filter by tier.",
      "enum": [
        "flagship",
        "datacenter",
        "prosumer",
        "entry"
      ]
    }
  }
}

Output Schema

{
  "type": "object",
  "properties": {
    "updated_at": {
      "type": [
        "string",
        "null"
      ]
    },
    "stale": {
      "type": "boolean"
    },
    "count": {
      "type": "integer"
    },
    "gpus": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "gpu": {
            "type": "string"
          },
          "vram_gb": {
            "type": [
              "number",
              "null"
            ]
          },
          "vendor": {
            "type": [
              "string",
              "null"
            ]
          },
          "arch": {
            "type": [
              "string",
              "null"
            ]
          },
          "tier": {
            "type": [
              "string",
              "null"
            ]
          },
          "cheapest_usd_hr": {
            "type": [
              "number",
              "null"
            ]
          },
          "cheapest_provider": {
            "type": [
              "string",
              "null"
            ]
          },
          "cheapest_min_gpu_count": {
            "type": [
              "integer",
              "null"
            ]
          },
          "provider_count": {
            "type": "integer"
          },
          "offer_count": {
            "type": "integer"
          },
          "url": {
            "type": "string"
          }
        }
      }
    }
  },
  "required": [
    "count",
    "gpus"
  ]
}
🟢match_workload(query, model, params_b, vram_gb, task, ...)

The routing DECISION: describe a job (a model, size, or GPU need) and get the ranked, reasoned recommendation for the cheapest place to run it across the live market, with the required VRAM, GPU count, effective $/hr, and how much cheaper it is than a hyperscaler. No key required. Results mirror the site and apply a small, disclosed partner tie-break between otherwise-equal offers (each match reports partner true/false).

Input Schema

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": "Plain-language job, e.g. \"cheapest to serve Llama 3 70B\" or \"2x H100 for fine-tuning\". Provide this OR a structured spec below."
    },
    "model": {
      "type": "string",
      "description": "Open model name to size against, e.g. \"Llama 3 70B\", \"Qwen 72B\", \"Mixtral\"."
    },
    "params_b": {
      "type": "number",
      "description": "Model size in billions of parameters when no exact model is named."
    },
    "vram_gb": {
      "type": "integer",
      "description": "Rough VRAM the job needs, in GB, if you already know it."
    },
    "task": {
      "type": "string",
      "description": "What the job does.",
      "enum": [
        "inference",
        "finetune-lora",
        "finetune-full",
        "generate",
        "transcribe",
        "embed"
      ]
    },
    "precision": {
      "type": "string",
      "description": "Numeric precision to size the model at.",
      "enum": [
        "fp16",
        "int8",
        "int4"
      ]
    },
    "gpu_count": {
      "type": "integer",
      "description": "Exact positive GPU count. Overrides a count in query text. Returns no matches if no supported configuration fits; omit for automatic sizing."
    },
    "budget_usd_hr": {
      "type": "number",
      "description": "Only recommend configs at or under this hourly budget."
    },
    "region": {
      "type": "string",
      "description": "Restrict to a data-residency region.",
      "enum": [
        "US",
        "EU",
        "ASIA"
      ]
    },
    "spot": {
      "type": "string",
      "description": "Set true to include interruptible spot capacity for a cheaper rate.",
      "enum": [
        "true",
        "false"
      ]
    },
    "reserved": {
      "type": "string",
      "description": "Set true to include reserved / committed-term capacity for a lower rate.",
      "enum": [
        "true",
        "false"
      ]
    }
  }
}

Output Schema

{
  "type": "object",
  "properties": {
    "updated_at": {
      "type": [
        "string",
        "null"
      ]
    },
    "stale": {
      "type": "boolean"
    },
    "workload": {
      "type": "object",
      "properties": {
        "interpretation": {
          "type": "string"
        },
        "identified": {
          "type": "boolean"
        },
        "task": {
          "type": "string"
        },
        "precision": {
          "type": "string"
        },
        "vram_required_gb": {
          "type": "number"
        },
        "gpu_count": {
          "type": [
            "integer",
            "null"
          ]
        }
      }
    },
    "hero": {
      "type": [
        "object",
        "null"
      ],
      "properties": {
        "hyperscaler_ceiling": {
          "type": [
            "string",
            "null"
          ]
        },
        "savings_pct": {
          "type": [
            "number",
            "null"
          ]
        },
        "same_model": {
          "type": [
            "boolean",
            "null"
          ]
        }
      }
    },
    "count": {
      "type": "integer"
    },
    "matches": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "gpu": {
            "type": "string"
          },
          "gpu_count": {
            "type": "integer"
          },
          "effective_usd_hr": {
            "type": "number"
          },
          "monthly_usd": {
            "type": "number"
          },
          "provider": {
            "type": "string"
          },
          "provider_label": {
            "type": [
              "string",
              "null"
            ]
          },
          "offer_type": {
            "type": "string"
          },
          "reliability": {
            "type": "string"
          },
          "fits_single_card": {
            "type": "boolean"
          },
          "over_budget": {
            "type": "boolean"
          },
          "tokens_per_sec": {
            "type": [
              "number",
              "null"
            ]
          },
          "usd_per_million_tokens": {
            "type": [
              "number",
              "null"
            ]
          },
          "score": {
            "type": [
              "number",
              "null"
            ]
          },
          "partner": {
            "type": "boolean"
          },
          "reason": {
            "type": "string"
          },
          "url": {
            "type": "string"
          }
        }
      }
    }
  },
  "required": [
    "count",
    "matches"
  ]
}

Recommended Prompts

list_items
List all [items] available in FastGPU
Expected tools: list_gpu_prices
browse_collection
Show me the [collection] from FastGPU
Expected tools: list_gpu_prices

Community

Rate this Server

Evidence

Recent observations

verifiedversion not recorded2 tools
verifiedversion not recorded2 tools