FitLLM
Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.
¿Debería usar esto?
Calidad y seguridad
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"fitllm": {
"url": "https://fitllm.run/api/mcp"
}
}
}Puntos de conexión remotos
https://fitllm.run/api/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (3)
🟢check_llm_fit(model, gpu, gpu_count, mac_ram_gb, quant, ...)
Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).
Esquema de entrada
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""
},
"gpu": {
"type": "string",
"description": "GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."
},
"quant": {
"type": "string",
"description": "Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."
},
"context_tokens": {
"type": "integer",
"minimum": 1024,
"description": "Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."
},
"ctx": {
"type": "integer",
"minimum": 1024,
"description": "Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."
},
"kv_bits": {
"type": "number",
"enum": [
16,
8,
4
],
"description": "KV-cache quantization bits (default 16 = F16)"
}
},
"required": [
"model"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢what_fits_on_hardware(gpu, gpu_count, mac_ram_gb)
Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.
Esquema de entrada
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."
},
"gpu_count": {
"type": "integer",
"minimum": 1,
"maximum": 8,
"description": "Number of identical copies of the gpu. Default 1."
},
"mac_ram_gb": {
"type": "integer",
"minimum": 8,
"maximum": 2048,
"description": "Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_supported
List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.
Esquema de entrada
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}Prompts recomendados
list_supportedlist_supportedComunidad
Evidencia