Nodegrove VRAM: can I run it?

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

Should I use this

Quality & Safety

A
Description quality
100%
Schema completeness
92%
Naming quality
87%
Poisoning risk
100%
Permission match
100%
Protocol compliance
100%

Based on automated analysis of tool definitions and protocol compliance.

Context Cost

~2,922Tokens (tool definitions)
~4.2 KBTypical response size
Significant attention impact (2.28% of 128k context)

This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.

Install

One-Click Install

Add this to your `claude_desktop_config.json` file:

{
  "mcpServers": {
    "vram-mcp": {
      "command": "npx",
      "args": [
        "@nodegrove/vram-mcp"
      ]
    }
  }
}

Runnable packages

npm@nodegrove/vram-mcp1.1.1stdio

Remote endpoints

https://mcp.nodegrove.io/mcpstreamable-http

What it can do

Tool inventory

Tools (6)

๐ŸŸข Read-only๐ŸŸก Write๐Ÿ”ด Deleteโšช Unknown
๐ŸŸขcan_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)

Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.

Input Schema

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size รท num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers ร— kv_heads ร— head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 ร— KV heads ร— head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
๐ŸŸขwhat_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)

Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.

Input Schema

{
  "type": "object",
  "properties": {
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
๐ŸŸขestimate_vram(model, architecture, active_params_b, quant, context, ...)

How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).

Input Schema

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size รท num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers ร— kv_heads ร— head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 ร— KV heads ร— head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "quant": {
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ],
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
๐ŸŸขestimate_from_hf_repo(repo, active_params_b, context, kv_cache)

Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.

Input Schema

{
  "type": "object",
  "properties": {
    "repo": {
      "type": "string",
      "minLength": 3,
      "maxLength": 200,
      "description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "required": [
    "repo"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
๐ŸŸขlist_models(search)

The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.

Input Schema

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
๐ŸŸขlist_gpus(search)

The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.

Input Schema

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Community

Rate this Server

Evidence

Recent observations

verifiedversion not recorded6 tools