Nodegrove VRAM: can I run it?

Can this LLM run on my GPU? VRAM, speed ceiling and what fits instead, for any model and GPU.

사용해야 할까요

품질 및 안전성

A
설명 품질
100%
스키마 완전성
92%
이름 품질
87%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~2,922토큰 (도구 정의)
~4.2 KB일반적인 응답 크기
상당한 주의 영향 (128k 컨텍스트의 2.28%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "vram-mcp": {
      "command": "npx",
      "args": [
        "@nodegrove/vram-mcp"
      ]
    }
  }
}

실행 가능한 패키지

npm@nodegrove/vram-mcp1.1.1stdio

원격 엔드포인트

https://mcp.nodegrove.io/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (6)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢can_i_run(model, architecture, active_params_b, gpu, vram_gb, ...)

Can this GPU run this open-weight LLM? Returns fits, tight or no, the memory split (weights, KV cache, overhead), a decode-speed ceiling, the longest context that fits and, on a no, every change that would make it fit: quantisation, KV cache, context, another card or a smaller model. Model: a name or id from list_models, any Hugging Face repo id, or its architecture. GPU: a name or id from list_gpus, or vram_gb for any other card.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢what_fits(gpu, vram_gb, bandwidth_gb_s, apple_silicon, quant, ...)

Which open-weight LLMs fit this GPU: every model in list_models checked at one quantisation and context, with a recommended everyday model (the biggest class that fits with room for context at conversational speed), the largest that fits, the best at Q8 and the first out of reach. GPU: a name or id from list_gpus, or vram_gb for any other card.

입력 스키마

{
  "type": "object",
  "properties": {
    "gpu": {
      "description": "A GPU from list_gpus (id or name, e.g. \"rtx-4090\", \"4090\" or \"M4 Max\").",
      "type": "string",
      "minLength": 1,
      "maxLength": 100
    },
    "vram_gb": {
      "description": "Memory of a card not in list_gpus, GB. For a Mac, its unified memory with apple_silicon: true.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 4096
    },
    "bandwidth_gb_s": {
      "description": "Memory bandwidth from the maker's spec, GB/s, for a speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 100000
    },
    "apple_silicon": {
      "description": "vram_gb is Apple unified memory; the GPU can use about 75% of it by default.",
      "type": "boolean"
    },
    "quant": {
      "default": "q4",
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). q4 is the common default.",
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ]
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_vram(model, architecture, active_params_b, quant, context, ...)

How much memory an LLM needs: weights + KV cache + overhead at each quantisation (or one), at a given context, and the smallest common card class that holds each. Model: a name or id from list_models, any Hugging Face repo id, or its architecture (params_b, layers, kv_heads, head_dim).

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "description": "A model from list_models (id or name, e.g. \"llama-3.3-70b\" or \"Llama 3.3 70B\"), or any Hugging Face repo id (e.g. \"Qwen/Qwen3-8B\"), read live from its config.json.",
      "type": "string",
      "minLength": 1,
      "maxLength": 200
    },
    "architecture": {
      "description": "A model described by its config.json values instead of a name.",
      "type": "object",
      "properties": {
        "params_b": {
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000,
          "description": "Total parameters, billions; all experts for a mixture-of-experts model."
        },
        "layers": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1000,
          "description": "num_hidden_layers"
        },
        "kv_heads": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 1024,
          "description": "num_key_value_heads"
        },
        "head_dim": {
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 4096,
          "description": "head_dim, or hidden_size ÷ num_attention_heads"
        },
        "active_params_b": {
          "description": "Parameters read per token, billions (mixture-of-experts only).",
          "type": "number",
          "exclusiveMinimum": 0,
          "maximum": 10000
        },
        "native_context": {
          "description": "The context window the model supports, tokens.",
          "type": "integer",
          "exclusiveMinimum": 0,
          "maximum": 9007199254740991
        },
        "kv_groups": {
          "description": "Only for non-standard attention: one entry per group of layers that cache the same way. Replaces layers × kv_heads × head_dim.",
          "maxItems": 8,
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "layers": {
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              },
              "values_per_token": {
                "type": "number",
                "exclusiveMinimum": 0,
                "description": "Values each layer caches per token: 2 × KV heads × head dim, or the latent width for MLA."
              },
              "window_tokens": {
                "description": "Sliding window: these layers keep only this many tokens.",
                "type": "integer",
                "exclusiveMinimum": 0,
                "maximum": 9007199254740991
              }
            },
            "required": [
              "layers",
              "values_per_token"
            ]
          }
        },
        "fixed_state_gb": {
          "description": "Fixed recurrent state of linear-attention or Mamba layers, GB.",
          "type": "number",
          "minimum": 0,
          "maximum": 100
        }
      },
      "required": [
        "params_b",
        "layers",
        "kv_heads",
        "head_dim"
      ]
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "quant": {
      "type": "string",
      "enum": [
        "fp16",
        "q8",
        "q6",
        "q5",
        "q4",
        "q3"
      ],
      "description": "Weight quantisation: fp16 (FP16 / BF16), q8 (Q8_0), q6 (Q6_K), q5 (Q5_K_M), q4 (Q4_K_M), q3 (Q3_K_M). Omit it for all 6."
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢estimate_from_hf_repo(repo, active_params_b, context, kv_cache)

Reads any Hugging Face model repo's config.json and parameter count and estimates its memory: the attention layout found (standard, sliding-window, hybrid or latent), how much each 1,000 tokens of context costs, and weights + KV cache + overhead at every quantisation. For models nodegrove.io has not reviewed; anything the reader cannot model is listed in warnings.

입력 스키마

{
  "type": "object",
  "properties": {
    "repo": {
      "type": "string",
      "minLength": 3,
      "maxLength": 200,
      "description": "Hugging Face repo id, e.g. \"Qwen/Qwen3-8B\", or its huggingface.co URL."
    },
    "active_params_b": {
      "description": "Parameters read per token, billions, for a mixture-of-experts model read from Hugging Face (from its model card). Sets the speed ceiling.",
      "type": "number",
      "exclusiveMinimum": 0,
      "maximum": 10000
    },
    "context": {
      "default": 8192,
      "description": "Tokens held in context: prompt plus conversation.",
      "type": "integer",
      "minimum": 1,
      "maximum": 10000000
    },
    "kv_cache": {
      "default": "fp16",
      "description": "KV cache precision. fp16 is what most runtimes use; q8 halves the cache.",
      "type": "string",
      "enum": [
        "fp16",
        "q8"
      ]
    }
  },
  "required": [
    "repo"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_models(search)

The open-weight LLMs nodegrove.io has verified against their config.json (data version 2026-10-06): id, size, attention design, native context, licence, memory at Q4 with 8k context and each model's page.

입력 스키마

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_gpus(search)

The GPUs and machines nodegrove.io covers: memory, the memory a runtime can use and bandwidth, from the makers' specs, with each one's page.

입력 스키마

{
  "type": "object",
  "properties": {
    "search": {
      "description": "Words to filter by, e.g. \"qwen\" or \"24 GB\".",
      "type": "string",
      "maxLength": 100
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 6개