Nodegrove VRAM: can I run it?
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
Sollte ich dies verwenden
Qualität und Sicherheit
Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.
Kontextkosten
Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.
Installieren
Installation mit einem Klick
Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:
{
"mcpServers": {
"vram-mcp": {
"command": "npx",
"args": [
"@nodegrove/vram-mcp"
]
}
}
}Ausführbare Pakete
1.1.1stdioRemote-Endpunkte
https://mcp.nodegrove.io/mcpstreamable-httpWas es kann
Tool-Inventar
Tools (6)
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
Eingabe-Schema
{
"type": "object",
"properties": {
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"quant": {
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
],
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
Eingabe-Schema
{
"type": "object",
"properties": {
"repo": {
"type": "string",
"minLength": 3,
"maxLength": 200,
"description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"required": [
"repo"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_models(search)
The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.
Eingabe-Schema
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_gpus(search)
The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.
Eingabe-Schema
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Community
Nachweis