Nodegrove VRAM: can I run it?

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

我该使用它吗

质量与安全性

A
描述质量
100%
模式完整度
92%
命名质量
87%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~2,922token 数(工具定义)
~4.2 KB典型响应大小
对注意力有显著影响(占 128k 上下文窗口的 2.28%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "vram-mcp": {
      "command": "npx",
      "args": [
        "@nodegrove/vram-mcp"
      ]
    }
  }
}

可运行的软件包

npm@nodegrove/vram-mcp1.1.1stdio

远程端点

https://mcp.nodegrove.io/mcpstreamable-http

它能做什么

工具清单

工具(6)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)

Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.

输入模式

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)

Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.

输入模式

{
  "type": "object",
  "properties": {
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)

How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).

输入模式

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "quant": {
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ],
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)

Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.

输入模式

{
  "type": "object",
  "properties": {
    "repo": {
      "type": "string",
      "minLength": 3,
      "maxLength": 200,
      "description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "required": [
    "repo"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_models(search)

The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.

输入模式

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_gpus(search)

The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.

输入模式

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

社区

评价此服务器

证据

最近观测

已验证未记录版本6 个工具