Nodegrove VRAM: can I run it?
Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.
使うべきか
品質と安全性
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"vram-mcp": {
"command": "npx",
"args": [
"@nodegrove/vram-mcp"
]
}
}
}実行可能なパッケージ
1.1.1stdioリモートエンドポイント
https://mcp.nodegrove.io/mcpstreamable-httpできること
ツール一覧
ツール(6)
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)
Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)
Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.
入力スキーマ
{
"type": "object",
"properties": {
"gpu": {
"description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
"type": "string",
"minLength": 1,
"maxLength": 100
},
"vram_gb": {
"description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 4096
},
"bandwidth_gb_s": {
"description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 100000
},
"apple_silicon": {
"description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
"type": "boolean"
},
"quant": {
"default": "q4",
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
]
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)
How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
"type": "string",
"minLength": 1,
"maxLength": 200
},
"architecture": {
"description": "A model described by its config.json values instead of a name.",
"type": "object",
"properties": {
"params_b": {
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000,
"description": "Total parameters, billions; all experts for a mixture-of-experts model."
},
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1000,
"description": "num_hidden_layers"
},
"kv_heads": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 1024,
"description": "num_key_value_heads"
},
"head_dim": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 4096,
"description": "head_dim, or hidden_size ÷ num_attention_heads"
},
"active_params_b": {
"description": "Parameters read per token, billions (mixture-of-experts only).",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"native_context": {
"description": "The context window the model supports, tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"kv_groups": {
"description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
"maxItems": 8,
"type": "array",
"items": {
"type": "object",
"properties": {
"layers": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
},
"values_per_token": {
"type": "number",
"exclusiveMinimum": 0,
"description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
},
"window_tokens": {
"description": "Sliding window: these layers keep only this many tokens.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"layers",
"values_per_token"
]
}
},
"fixed_state_gb": {
"description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
"type": "number",
"minimum": 0,
"maximum": 100
}
},
"required": [
"params_b",
"layers",
"kv_heads",
"head_dim"
]
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"quant": {
"type": "string",
"enum": [
"fp16",
"q8",
"q6",
"q5",
"q4",
"q3"
],
"description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)
Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.
入力スキーマ
{
"type": "object",
"properties": {
"repo": {
"type": "string",
"minLength": 3,
"maxLength": 200,
"description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
},
"active_params_b": {
"description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
"type": "number",
"exclusiveMinimum": 0,
"maximum": 10000
},
"context": {
"default": 8192,
"description": "Tokens held in context: prompt plus conversation.",
"type": "integer",
"minimum": 1,
"maximum": 10000000
},
"kv_cache": {
"default": "fp16",
"description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
"type": "string",
"enum": [
"fp16",
"q8"
]
}
},
"required": [
"repo"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_models(search)
The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.
入力スキーマ
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_gpus(search)
The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.
入力スキーマ
{
"type": "object",
"properties": {
"search": {
"description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
"type": "string",
"maxLength": 100
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}コミュニティ
エビデンス