StudioTV LLM VRAM Calculator
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
使うべきか
品質と安全性
検出事項(2)
- HIGH
- MEDIUMestimate_vram 内
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"llm-vram": {
"url": "https://studiotvai.com/api/mcp"
}
}
}リモートエンドポイント
https://studiotvai.com/api/mcpstreamable-httpできること
ツール一覧
ツール(3)
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)
GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
},
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
},
"gpu_count": {
"type": [
"integer",
"string"
],
"description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"model"
]
}🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)
Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.
入力スキーマ
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\"."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"description": "Default 1."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"gpu"
]
}🟢gpu_prices(gpu)
Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.
入力スキーマ
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name. Omit for every GPU."
}
}
}コミュニティ
エビデンス