FitLLM

Will this LLM fit on your GPU, multi-GPU rig or Mac? Exact VRAM & KV-cache math. Read-only.

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
73%
Calidad de los nombres
93%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~963Tokens (definiciones de herramientas)
~1.9 KBTamaño de respuesta típico
Impacto moderado en la atención (0.75% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "fitllm": {
      "url": "https://fitllm.run/api/mcp"
    }
  }
}

Puntos de conexión remotos

https://fitllm.run/api/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (3)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟢check_llm_fit(model, gpu, gpu_count, mac_ram_gb, quant, ...)

Check whether a specific local LLM fits in the memory of a specific GPU or Apple Silicon Mac. Returns fits/tight/won't-fit verdict with the memory breakdown (weights, KV cache, linear-attention state when present, runtime overhead, reserve), max context, and a concrete fix if it doesn't fit. Use this whenever a user asks anything like "can I run <model> on my <GPU/Mac>?", "will <model> fit in <N>GB?", or "what do I need to run <model>?". Estimates using curated, config-derived architecture fields (MLA, sliding-window, hybrid attention, MoE modeled).

Esquema de entrada

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "LLM name, fuzzy — e.g. \"GLM-4.7-Flash\", \"gpt-oss-20b\", \"gemma 31b\""
    },
    "gpu": {
      "type": "string",
      "description": "GPU name, fuzzy — e.g. \"RTX 4090\", \"RX 7900 XTX\", \"A100 80GB\". Multi-GPU rigs: join with + — e.g. \"RTX 5090 + RTX 3090\" (VRAM pools across cards). Provide gpu OR mac_ram_gb."
    },
    "gpu_count": {
      "type": "integer",
      "minimum": 1,
      "maximum": 8,
      "description": "Number of identical copies of the gpu (e.g. gpu=\"RTX 3090\", gpu_count=2 for a 2×3090 rig). Default 1."
    },
    "mac_ram_gb": {
      "type": "integer",
      "minimum": 8,
      "maximum": 2048,
      "description": "Apple Silicon unified memory in GB — e.g. 16, 64, 512. Provide gpu OR mac_ram_gb."
    },
    "quant": {
      "type": "string",
      "description": "Weight quantization. GPU: Q4_K_M(default)/Q5_K_M/Q6_K/Q8_0/FP16. Mac: 4/8(default)/16 (bits)."
    },
    "context_tokens": {
      "type": "integer",
      "minimum": 1024,
      "description": "Context length in tokens (default 8192). Alias: ctx (same field as the REST API)."
    },
    "ctx": {
      "type": "integer",
      "minimum": 1024,
      "description": "Alias of context_tokens — accepted because the REST API uses this name. Do not pass both with different values."
    },
    "kv_bits": {
      "type": "number",
      "enum": [
        16,
        8,
        4
      ],
      "description": "KV-cache quantization bits (default 16 = F16)"
    }
  },
  "required": [
    "model"
  ],
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢what_fits_on_hardware(gpu, gpu_count, mac_ram_gb)

Rank which popular local LLMs fit on a given GPU or Apple Silicon Mac (at ~4-bit quantization, 8K context) — models that fit come first, biggest first, with max context each. Use when a user asks "what can I run on my <GPU/Mac/N GB>?", "best local model for my machine?", or gives hardware without naming a model.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "gpu": {
      "type": "string",
      "description": "GPU name, fuzzy. Multi-GPU rigs: join with + (e.g. \"RTX 5090 + RTX 3090\"). Provide gpu OR mac_ram_gb."
    },
    "gpu_count": {
      "type": "integer",
      "minimum": 1,
      "maximum": 8,
      "description": "Number of identical copies of the gpu. Default 1."
    },
    "mac_ram_gb": {
      "type": "integer",
      "minimum": 8,
      "maximum": 2048,
      "description": "Apple Silicon unified memory GB. Provide gpu OR mac_ram_gb."
    }
  },
  "additionalProperties": false,
  "$schema": "http://json-schema.org/draft-07/schema#"
}
🟢list_supported

List the built-in model names and hardware names this fit-checker knows (for mapping user wording to exact names). Standard text-only HuggingFace transformer configs can also be checked via fitllm.run; unsupported architectures are rejected.

Esquema de entrada

{
  "type": "object",
  "properties": {},
  "$schema": "http://json-schema.org/draft-07/schema#"
}

Prompts recomendados

list_items
List all [items] available in FitLLM
Herramientas esperadas: list_supported
browse_collection
Show me the [collection] from FitLLM
Herramientas esperadas: list_supported

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada3 herramientas
verificadoversión no registrada3 herramientas