StudioTV LLM VRAM Calculator
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
Should I use this
Quality & Safety
Findings (2)
- HIGH
- MEDIUMin estimate_vram
Based on automated analysis of tool definitions and protocol compliance.
Context Cost
This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.
Install
One-Click Install
Add this to your `claude_desktop_config.json` file:
{
"mcpServers": {
"llm-vram": {
"url": "https://studiotvai.com/api/mcp"
}
}
}Remote endpoints
https://studiotvai.com/api/mcpstreamable-httpWhat it can do
Tool inventory
Tools (3)
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)
GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).
Input Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
},
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
},
"gpu_count": {
"type": [
"integer",
"string"
],
"description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"model"
]
}🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)
Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.
Input Schema
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\"."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"description": "Default 1."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"gpu"
]
}🟢gpu_prices(gpu)
Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.
Input Schema
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name. Omit for every GPU."
}
}
}Community
Evidence