StudioTV LLM VRAM Calculator
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
¿Debería usar esto?
Calidad y seguridad
Hallazgos (2)
- HIGH
- MEDIUMen estimate_vram
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"llm-vram": {
"url": "https://studiotvai.com/api/mcp"
}
}
}Puntos de conexión remotos
https://studiotvai.com/api/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (3)
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)
GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).
Esquema de entrada
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
},
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
},
"gpu_count": {
"type": [
"integer",
"string"
],
"description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"model"
]
}🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)
Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.
Esquema de entrada
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\"."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"description": "Default 1."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"gpu"
]
}🟢gpu_prices(gpu)
Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.
Esquema de entrada
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name. Omit for every GPU."
}
}
}Comunidad
Evidencia