FitLLM
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
Sollte ich dies verwenden
Qualität und Sicherheit
Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.
Kontextkosten
Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.
Installieren
Installation mit einem Klick
Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:
{
"mcpServers": {
"fitllm": {
"url": "https://fitllm.run/api/mcp"
}
}
}Remote-Endpunkte
https://fitllm.run/api/mcpstreamable-httpWas es kann
Tool-Inventar
Tools (3)
🟢check_llm_fit(model, gpu, gpu_count, mac_ram_gb, quant, ...)
Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""
},
"gpu": {
"type": "string",
"description": "GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."
},
"quant": {
"type": "string",
"description": "Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."
},
"context_tokens": {
"type": "integer",
"minimum": 1024,
"description": "Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."
},
"ctx": {
"type": "integer",
"minimum": 1024,
"description": "Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."
},
"kv_bits": {
"type": "number",
"enum": [
16,
8,
4
],
"description": "KV-cache quantization bits (default 16 = F16)"
}
},
"required": [
"model"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢what_fits_on_hardware(gpu, gpu_count, mac_ram_gb)
Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.
Eingabe-Schema
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu. Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_supported
List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.
Eingabe-Schema
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}Empfohlene Prompts
list_supportedlist_supportedCommunity
Nachweis