StudioTV LLM VRAM Calculator
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
Sollte ich dies verwenden
Qualität und Sicherheit
Befunde (2)
- HIGH
- MEDIUMin estimate_vram
Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.
Kontextkosten
Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.
Installieren
Installation mit einem Klick
Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:
{
"mcpServers": {
"llm-vram": {
"url": "https://studiotvai.com/api/mcp"
}
}
}Remote-Endpunkte
https://studiotvai.com/api/mcpstreamable-httpWas es kann
Tool-Inventar
Tools (3)
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)
GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
},
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
},
"gpu_count": {
"type": [
"integer",
"string"
],
"description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"model"
]
}🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)
Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.
Eingabe-Schema
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\"."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"description": "Default 1."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"gpu"
]
}🟢gpu_prices(gpu)
Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.
Eingabe-Schema
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name. Omit for every GPU."
}
}
}Community
Nachweis