FitLLM
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
使うべきか
品質と安全性
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"fitllm": {
"url": "https://fitllm.run/api/mcp"
}
}
}リモートエンドポイント
https://fitllm.run/api/mcpstreamable-httpできること
ツール一覧
ツール(3)
🟢check_llm_fit(model, gpu, gpu_count, mac_ram_gb, quant, ...)
Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""
},
"gpu": {
"type": "string",
"description": "GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."
},
"quant": {
"type": "string",
"description": "Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."
},
"context_tokens": {
"type": "integer",
"minimum": 1024,
"description": "Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."
},
"ctx": {
"type": "integer",
"minimum": 1024,
"description": "Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."
},
"kv_bits": {
"type": "number",
"enum": [
16,
8,
4
],
"description": "KV-cache quantization bits (default 16 = F16)"
}
},
"required": [
"model"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢what_fits_on_hardware(gpu, gpu_count, mac_ram_gb)
Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.
入力スキーマ
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu. Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_supported
List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.
入力スキーマ
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}推奨プロンプト
list_supportedlist_supportedコミュニティ
エビデンス