Nodegrove VRAM: can I run it?

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
92%
Calidad de los nombres
87%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~2,922Tokens (definiciones de herramientas)
~4.2 KBTamaño de respuesta típico
Impacto significativo en la atención (2.28% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "vram-mcp": {
      "command": "npx",
      "args": [
        "@nodegrove/vram-mcp"
      ]
    }
  }
}

Paquetes ejecutables

npm@nodegrove/vram-mcp1.1.1stdio

Puntos de conexión remotos

https://mcp.nodegrove.io/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (6)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)

Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)

Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)

How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).

Esquema de entrada

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "quant": {
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ],
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)

Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "repo": {
      "type": "string",
      "minLength": 3,
      "maxLength": 200,
      "description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "required": [
    "repo"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_models(search)

The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_gpus(search)

The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada6 herramientas