Nodegrove VRAM: can I run it?
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
我該用這個嗎
品質與安全性
根據工具定義與協定合規性的自動化分析。
上下文成本
這是每次將伺服器的工具載入模型上下文時所消耗的約略 token 數量。數量越高,可用於其他工作的注意力就越少。
安裝
一鍵安裝
將以下內容加入你的 `claude_desktop_config.json` 檔案:
{
"mcpServers": {
"vram-mcp": {
"command": "npx",
"args": [
"@nodegrove/vram-mcp"
]
}
}
}可執行的套件
1.1.1stdio遠端端點
https://mcp.nodegrove.io/mcpstreamable-http它能做什麼
工具清單
工具(6)
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
輸入結構描述
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
輸入結構描述
{
"type": "object",
"properties": {
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
輸入結構描述
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"quant": {
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
],
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
輸入結構描述
{
"type": "object",
"properties": {
"repo": {
"type": "string",
"minLength": 3,
"maxLength": 200,
"description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"required": [
"repo"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_models(search)
The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.
輸入結構描述
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_gpus(search)
The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.
輸入結構描述
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}社群
證據