FitLLM

Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.

사용해야 할까요

품질 및 안전성

A
설명 품질
100%
스키마 완전성
73%
이름 품질
93%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~963토큰 (도구 정의)
~1.9 KB일반적인 응답 크기
중간 정도의 주의 영향 (128k 컨텍스트의 0.75%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "fitllm": {
      "url": "https://fitllm.run/api/mcp"
    }
  }
}

원격 엔드포인트

https://fitllm.run/api/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (3)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢check_llm_fit(model, gpu, gpu_count, mac_ram_gb, quant, ...)

Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""
    },
    "gpu": {
      "type": "string",
      "description": "GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."
    },
    "gpu_count": {
      "type": "integer",
      "minimum": 1,
      "maximum": 8,
      "description": "Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."
    },
    "mac_ram_gb": {
      "type": "integer",
      "minimum": 8,
      "maximum": 2048,
      "description": "Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."
    },
    "quant": {
      "type": "string",
      "description": "Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."
    },
    "context_tokens": {
      "type": "integer",
      "minimum": 1024,
      "description": "Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."
    },
    "ctx": {
      "type": "integer",
      "minimum": 1024,
      "description": "Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."
    },
    "kv_bits": {
      "type": "number",
      "enum": [
        16,
        8,
        4
      ],
      "description": "KV-cache quantization bits (default 16 = F16)"
    }
  },
  "required": [
    "model"
  ],
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢what_fits_on_hardware(gpu, gpu_count, mac_ram_gb)

Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.

입력 스키마

{
  "type": "object",
  "properties": {
    "gpu": {
      "type": "string",
      "description": "GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."
    },
    "gpu_count": {
      "type": "integer",
      "minimum": 1,
      "maximum": 8,
      "description": "Number of identical copies of the gpu. Default 1."
    },
    "mac_ram_gb": {
      "type": "integer",
      "minimum": 8,
      "maximum": 2048,
      "description": "Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."
    }
  },
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢list_supported

List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "http://json-schema.org/draft-07/schema#"
}

권장 프롬프트

list_items
List all [items] available in FitLLM
예상 도구: list_supported
browse_collection
Show me the [collection] from FitLLM
예상 도구: list_supported

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 3개
검증됨버전이 기록되지 않음도구 3개