Nodegrove VRAM: can I run it?
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
我该使用它吗
质量与安全性
基于对工具定义和协议合规性的自动分析。
上下文开销
这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。
安装
一键安装
将以下内容添加到你的 `claude_desktop_config.json` 文件中:
{
"mcpServers": {
"vram-mcp": {
"command": "npx",
"args": [
"@nodegrove/vram-mcp"
]
}
}
}可运行的软件包
1.1.1stdio远程端点
https://mcp.nodegrove.io/mcpstreamable-http它能做什么
工具清单
工具(6)
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
输入模式
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
输入模式
{
"type": "object",
"properties": {
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
输入模式
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"quant": {
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
],
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
输入模式
{
"type": "object",
"properties": {
"repo": {
"type": "string",
"minLength": 3,
"maxLength": 200,
"description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"required": [
"repo"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_models(search)
The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.
输入模式
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_gpus(search)
The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.
输入模式
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}社区
证据