StudioTV LLM VRAM Calculator

Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.

Sollte ich dies verwenden

Qualität und Sicherheit

A
Qualität der Beschreibung
93%
Vollständigkeit des Schemas
97%
Qualität der Benennung
80%
Risiko der Vergiftung
80%
Übereinstimmung der Berechtigungen
100%
Einhaltung des Protokolls
100%

Befunde (2)

  • HIGHTool poisoning patterns detected
  • MEDIUMTool description contains URL to non-standard domainin estimate_vram

Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.

Kontextkosten

~877Tokens (Tool-Definitionen)
~2.2 KBTypische Antwortgröße
Mittlere Auswirkung auf die Aufmerksamkeit (0.69% von 128k Kontext)

Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.

Installieren

Installation mit einem Klick

Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:

{
  "mcpServers": {
    "llm-vram": {
      "url": "https://studiotvai.com/api/mcp"
    }
  }
}

Remote-Endpunkte

https://studiotvai.com/api/mcpstreamable-http

Was es kann

Tool-Inventar

Tools (3)

🟢 Nur lesen🟡 Schreiben🔴 Löschen⚪ Unbekannt
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)

GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
    },
    "gpu": {
      "type": "string",
      "description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
    },
    "gpu_count": {
      "type": [
        "integer",
        "string"
      ],
      "description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
    },
    "context": {
      "type": "integer",
      "minimum": 1,
      "description": "Tokens per request (prompt + output). Default 8192."
    },
    "concurrent_requests": {
      "type": "integer",
      "minimum": 1,
      "description": "Requests served at the same time, each with its own KV cache. Default 1."
    },
    "weights": {
      "type": "string",
      "enum": [
        "native",
        "bf16",
        "fp8",
        "int8",
        "nvfp4",
        "mxfp4",
        "int4",
        "q8_0",
        "q6_k",
        "q5_k_m",
        "q4_k_m",
        "q3_k_m",
        "fp32"
      ],
      "description": "Weight format. Default: the official checkpoint (native) or BF16."
    },
    "kv_cache": {
      "type": "string",
      "enum": [
        "bf16",
        "fp8",
        "q8_0",
        "q4_0"
      ],
      "description": "KV cache precision. Default bf16."
    },
    "engine": {
      "type": "string",
      "enum": [
        "vllm",
        "sglang",
        "trtllm",
        "llamacpp",
        "transformers"
      ],
      "description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
    }
  },
  "required": [
    "model"
  ]
}
🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)

Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "gpu": {
      "type": "string",
      "description": "GPU id or name, e.g. \"rtx-4090\"."
    },
    "gpu_count": {
      "type": "integer",
      "minimum": 1,
      "description": "Default 1."
    },
    "context": {
      "type": "integer",
      "minimum": 1,
      "description": "Tokens per request (prompt + output). Default 8192."
    },
    "concurrent_requests": {
      "type": "integer",
      "minimum": 1,
      "description": "Requests served at the same time, each with its own KV cache. Default 1."
    },
    "weights": {
      "type": "string",
      "enum": [
        "native",
        "bf16",
        "fp8",
        "int8",
        "nvfp4",
        "mxfp4",
        "int4",
        "q8_0",
        "q6_k",
        "q5_k_m",
        "q4_k_m",
        "q3_k_m",
        "fp32"
      ],
      "description": "Weight format. Default: the official checkpoint (native) or BF16."
    },
    "kv_cache": {
      "type": "string",
      "enum": [
        "bf16",
        "fp8",
        "q8_0",
        "q4_0"
      ],
      "description": "KV cache precision. Default bf16."
    },
    "engine": {
      "type": "string",
      "enum": [
        "vllm",
        "sglang",
        "trtllm",
        "llamacpp",
        "transformers"
      ],
      "description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
    }
  },
  "required": [
    "gpu"
  ]
}
🟢gpu_prices(gpu)

Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "gpu": {
      "type": "string",
      "description": "GPU id or name. Omit for every GPU."
    }
  }
}

Community

Diesen Server bewerten

Nachweis

Aktuelle Beobachtungen

verifiziertVersion nicht aufgezeichnet3 Tools