StudioTV LLM VRAM Calculator
Does an LLM fit on your GPU? VRAM, KV cache and GPU count for any Hugging Face model.
사용해야 할까요
품질 및 안전성
발견 사항 (2)
- HIGH
- MEDIUMestimate_vram에서
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"llm-vram": {
"url": "https://studiotvai.com/api/mcp"
}
}
}원격 엔드포인트
https://studiotvai.com/api/mcpstreamable-http할 수 있는 일
도구 목록
도구 (3)
🟢estimate_vram(model, gpu, gpu_count, context, concurrent_requests, ...)
GPU memory, number of GPUs, speed and rental cost to run an open LLM. Works for the models listed at https://studiotvai.com/api/models.json and any Hugging Face model id or link. Uses the real KV cache of each architecture (sliding window, hybrid linear attention, MLA).
입력 스키마
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name or id (\"Llama 3.3 70B\", \"qwen3.8-27b\"), Hugging Face id (\"Qwen/Qwen3-32B\") or link."
},
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\", \"H100\", \"mac-m4-max-128\". Default h100. List: https://studiotvai.com/api/gpus.json"
},
"gpu_count": {
"type": [
"integer",
"string"
],
"description": "Number of GPUs, or \"auto\" (default) for the fewest that fit."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"model"
]
}🟢models_that_fit(gpu, gpu_count, context, concurrent_requests, weights, ...)
Every listed open model that fits on the given GPU(s), largest first, with the most faithful weight format that fits.
입력 스키마
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name, e.g. \"rtx-4090\"."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"description": "Default 1."
},
"context": {
"type": "integer",
"minimum": 1,
"description": "Tokens per request (prompt + output). Default 8192."
},
"concurrent_requests": {
"type": "integer",
"minimum": 1,
"description": "Requests served at the same time, each with its own KV cache. Default 1."
},
"weights": {
"type": "string",
"enum": [
"native",
"bf16",
"fp8",
"int8",
"nvfp4",
"mxfp4",
"int4",
"q8_0",
"q6_k",
"q5_k_m",
"q4_k_m",
"q3_k_m",
"fp32"
],
"description": "Weight format. Default: the official checkpoint (native) or BF16."
},
"kv_cache": {
"type": "string",
"enum": [
"bf16",
"fp8",
"q8_0",
"q4_0"
],
"description": "KV cache precision. Default bf16."
},
"engine": {
"type": "string",
"enum": [
"vllm",
"sglang",
"trtllm",
"llamacpp",
"transformers"
],
"description": "Serving engine. Default: vllm on data-center GPUs, llamacpp elsewhere."
}
},
"required": [
"gpu"
]
}🟢gpu_prices(gpu)
Cheapest on-demand price per GPU-hour from RunPod, Vast.ai, Verda and Azure, checked every hour.
입력 스키마
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU id or name. Omit for every GPU."
}
}
}커뮤니티
증거