vetted-consumer
Will a local LLM run on your hardware? GGUF quant, buy-vs-rent-vs-API cost, used-GPU prices.
사용해야 할까요
품질 및 안전성
발견 사항 (2)
- LOWget_used_gpu_prices에서
- LOWcompare_hardware에서
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"vetted-consumer": {
"url": "https://vettedconsumer.com/mcp"
}
}
}원격 엔드포인트
https://vettedconsumer.com/mcpstreamable-http할 수 있는 일
도구 목록
도구 (9)
⚪can_i_run_it(model, total_b, active_b, mxfp4, hardware, ...)
Will a given local LLM run on given hardware? Returns fit, the best quant that fits, theoretical tok/s, and real owner-measured tok/s where available.
입력 스키마
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name, e.g. 'Llama 70B', 'gpt-oss-120B', 'Qwen 32B'. Use list_models to see known names."
},
"total_b": {
"type": "number",
"description": "For an unlisted model: total parameters in billions"
},
"active_b": {
"type": "number",
"description": "For an unlisted model: active params in billions (= total for dense, less for MoE)"
},
"mxfp4": {
"type": "boolean",
"description": "True if the model ships natively in MXFP4 (e.g. gpt-oss)"
},
"hardware": {
"type": "string",
"description": "Hardware name/id, e.g. 'rtx-3090', 'Mac 128GB', 'Strix Halo'. Use list_hardware to see known ones."
},
"vram_gb": {
"type": "number",
"description": "For custom hardware: VRAM or unified memory in GB"
},
"bandwidth_gbps": {
"type": "number",
"description": "For custom hardware: memory bandwidth in GB/s"
},
"unified": {
"type": "boolean",
"description": "True for unified-memory machines (Macs, Strix Halo, CPU+RAM)"
},
"context": {
"type": "number",
"description": "Context window in tokens (default 8192)"
},
"kv_precision": {
"type": "string",
"enum": [
"f16",
"q8",
"q4"
],
"description": "KV cache precision (default f16)"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢recommend_quant(model, total_b, active_b, mxfp4, hardware, ...)
Which GGUF quantization to download for a model on given hardware: the full quant ladder with file size, max context, and tok/s for each, plus the recommended pick.
입력 스키마
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name, e.g. 'Llama 70B', 'gpt-oss-120B', 'Qwen 32B'. Use list_models to see known names."
},
"total_b": {
"type": "number",
"description": "For an unlisted model: total parameters in billions"
},
"active_b": {
"type": "number",
"description": "For an unlisted model: active params in billions (= total for dense, less for MoE)"
},
"mxfp4": {
"type": "boolean",
"description": "True if the model ships natively in MXFP4 (e.g. gpt-oss)"
},
"hardware": {
"type": "string",
"description": "Hardware name/id, e.g. 'rtx-3090', 'Mac 128GB', 'Strix Halo'. Use list_hardware to see known ones."
},
"vram_gb": {
"type": "number",
"description": "For custom hardware: VRAM or unified memory in GB"
},
"bandwidth_gbps": {
"type": "number",
"description": "For custom hardware: memory bandwidth in GB/s"
},
"unified": {
"type": "boolean",
"description": "True for unified-memory machines (Macs, Strix Halo, CPU+RAM)"
},
"context": {
"type": "number",
"description": "Context window in tokens (default 8192)"
},
"kv_precision": {
"type": "string",
"enum": [
"f16",
"q8",
"q4"
],
"description": "KV cache precision (default f16)"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪cheapest_hardware_for_model(model, total_b, active_b, mxfp4, context)
The cheapest catalogued, buyable machine that runs a given model at Q4 with the requested context.
입력 스키마
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name, e.g. 'Llama 70B', 'gpt-oss-120B', 'Qwen 32B'. Use list_models to see known names."
},
"total_b": {
"type": "number",
"description": "For an unlisted model: total parameters in billions"
},
"active_b": {
"type": "number",
"description": "For an unlisted model: active params in billions (= total for dense, less for MoE)"
},
"mxfp4": {
"type": "boolean",
"description": "True if the model ships natively in MXFP4 (e.g. gpt-oss)"
},
"context": {
"type": "number",
"description": "Context window in tokens (default 8192)"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_models
List the local LLM model classes the tools know about (params, dense/MoE, native context).
입력 스키마
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢list_hardware
List the machines the tools know about (memory, bandwidth, price, buy link).
입력 스키마
{
"type": "object",
"properties": {},
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪cost_compare(hardware, price_usd, tdp_w, hours, tokens, ...)
Buy vs rent vs API cost to run a model locally: monthly/1y/3y totals, break-even months, and the energy cost per 1M tokens. Same math as /cost-calculator/.
입력 스키마
{
"type": "object",
"properties": {
"hardware": {
"type": "string",
"description": "Catalogued hardware name/id (see list_hardware), e.g. 'rtx-3090-used'"
},
"price_usd": {
"type": "number",
"description": "For custom hardware: price in USD"
},
"tdp_w": {
"type": "number",
"description": "For custom hardware: board power draw in watts"
},
"hours": {
"type": "number",
"description": "Active hours per day (default 3)"
},
"tokens": {
"type": "number",
"description": "Tokens generated per day, for the API comparison (default 300000)"
},
"kwh": {
"type": "number",
"description": "Electricity $/kWh (default 0.16)"
},
"rent": {
"type": "number",
"description": "Cloud GPU $/hour (default 0.59)"
},
"api": {
"type": "number",
"description": "API $/million tokens (default 1.0)"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢recommend_hardware(model, total_b, active_b, mxfp4, context, ...)
Ranked list of catalogued, buyable machines that run a model at the requested context, cheapest first, with an optional budget cap.
입력 스키마
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name, e.g. 'Llama 70B', 'gpt-oss-120B', 'Qwen 32B'. Use list_models to see known names."
},
"total_b": {
"type": "number",
"description": "For an unlisted model: total parameters in billions"
},
"active_b": {
"type": "number",
"description": "For an unlisted model: active params in billions (= total for dense, less for MoE)"
},
"mxfp4": {
"type": "boolean",
"description": "True if the model ships natively in MXFP4 (e.g. gpt-oss)"
},
"context": {
"type": "number",
"description": "Context window in tokens (default 8192)"
},
"kv_precision": {
"type": "string",
"enum": [
"f16",
"q8",
"q4"
],
"description": "KV cache precision (default f16)"
},
"budget": {
"type": "number",
"description": "Optional max price in USD"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢get_used_gpu_prices(gpu)
Current typical used-GPU prices for local-AI rigs (eBay Browse API median asking + hand-verified, monthly).
입력 스키마
{
"type": "object",
"properties": {
"gpu": {
"type": "string",
"description": "Optional name/id filter, e.g. \"3090\""
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪compare_hardware(hardware, model, total_b, active_b, mxfp4, ...)
Side-by-side memory, bandwidth, price, and (with a model) fit + tok/s for 2 to 4 machines.
입력 스키마
{
"type": "object",
"properties": {
"hardware": {
"type": "string",
"description": "2 to 4 hardware names/ids, comma-separated"
},
"model": {
"type": "string",
"description": "Model name, e.g. 'Llama 70B', 'gpt-oss-120B', 'Qwen 32B'. Use list_models to see known names."
},
"total_b": {
"type": "number",
"description": "For an unlisted model: total parameters in billions"
},
"active_b": {
"type": "number",
"description": "For an unlisted model: active params in billions (= total for dense, less for MoE)"
},
"mxfp4": {
"type": "boolean",
"description": "True if the model ships natively in MXFP4 (e.g. gpt-oss)"
},
"context": {
"type": "number",
"description": "Context window in tokens (default 8192)"
},
"kv_precision": {
"type": "string",
"enum": [
"f16",
"q8",
"q4"
],
"description": "KV cache precision (default f16)"
}
},
"required": [
"hardware"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}커뮤니티
증거