Nonobench

An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.

Sollte ich dies verwenden

Qualität und Sicherheit

A
Qualität der Beschreibung
74%
Vollständigkeit des Schemas
85%
Qualität der Benennung
98%
Risiko der Vergiftung
100%
Übereinstimmung der Berechtigungen
100%
Einhaltung des Protokolls
100%

Befunde (7)

  • LOWTool 'get_leaderboard' description lacks action verbin get_leaderboard
  • LOWTool 'list_families' description lacks action verbin list_families
  • LOWTool 'compare_models' description lacks action verbin compare_models
  • LOWTool 'get_model_results' description lacks action verbin get_model_results
  • LOWTool 'list_puzzles' description lacks action verbin list_puzzles
  • LOWTool 'get_puzzle' description lacks action verbin get_puzzle
  • LOWTool 'get_puzzle_results' description lacks action verbin get_puzzle_results

Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.

Kontextkosten

~3,234Tokens (Tool-Definitionen)
~2.9 KBTypische Antwortgröße
Erhebliche Auswirkung auf die Aufmerksamkeit (2.53% von 128k Kontext)

Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.

Installieren

Installation mit einem Klick

Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:

{
  "mcpServers": {
    "nonobench": {
      "url": "https://www.nonobench.com/mcp"
    }
  }
}

Remote-Endpunkte

https://www.nonobench.com/mcpstreamable-http

Was es kann

Tool-Inventar

Tools (11)

🟢 Nur lesen🟡 Schreiben🔴 Löschen⚪ Unbekannt
🟢get_leaderboard(size, provider, family, version, effort, ...)

Models ranked by accuracy; defaults to all effort levels for compatibility.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "size": {
      "description": "Grid size to filter on",
      "type": "string",
      "enum": [
        "5x5",
        "10x10",
        "15x15",
        "20x20"
      ]
    },
    "provider": {
      "description": "Comma-separated provider ids; empty means no filter",
      "type": "string"
    },
    "family": {
      "description": "Comma-separated family ids; empty means no filter",
      "type": "string"
    },
    "version": {
      "description": "Comma-separated benchmark versions: 1.0, 1.1, 1.2; empty means all",
      "type": "string"
    },
    "effort": {
      "description": "best, all (default), or one effort level; empty means all",
      "type": "string"
    },
    "reasoning": {
      "description": "Only reasoning (true) or non-reasoning (false) variants",
      "type": "boolean"
    },
    "open_weights": {
      "description": "Only open-weight (true) or closed (false) models",
      "type": "boolean"
    },
    "min_correct": {
      "description": "Minimum puzzles solved in the selected tier; default 0 includes unsolved variants",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "updatedAt": {
      "type": "string"
    },
    "size": {
      "type": "string"
    },
    "models": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "model": {
            "type": "string",
            "description": "Model variant id, e.g. claude-opus-5.5-high"
          },
          "displayName": {
            "type": "string"
          },
          "family": {
            "type": "string"
          },
          "effort": {
            "type": [
              "string",
              "null"
            ]
          },
          "provider": {
            "type": [
              "string",
              "null"
            ]
          },
          "reasoning": {
            "type": "boolean"
          },
          "accuracy": {
            "type": "number",
            "description": "Percentage of puzzles solved"
          },
          "rank": {
            "type": "number"
          },
          "correct": {
            "type": "number"
          },
          "total": {
            "type": "number"
          },
          "totalCostUsd": {
            "type": "number"
          }
        },
        "required": [
          "model",
          "displayName",
          "family",
          "effort",
          "provider",
          "reasoning",
          "accuracy",
          "rank",
          "correct",
          "total",
          "totalCostUsd"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "updatedAt",
    "size",
    "models"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢list_providers

Provider ids, names, families and variant counts.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "providers": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": {
            "type": "string"
          },
          "name": {
            "type": "string"
          },
          "variantCount": {
            "type": "number"
          },
          "families": {
            "type": "array",
            "items": {
              "type": "string"
            }
          }
        },
        "required": [
          "id",
          "name",
          "variantCount",
          "families"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "providers"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢list_families

Model families, available efforts and best variants.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "families": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "family": {
            "type": "string"
          },
          "displayName": {
            "type": "string"
          },
          "provider": {
            "type": [
              "string",
              "null"
            ]
          },
          "bestVariant": {
            "type": "string"
          },
          "efforts": {
            "type": "array",
            "items": {
              "type": "string"
            }
          }
        },
        "required": [
          "family",
          "displayName",
          "provider",
          "bestVariant",
          "efforts"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "families"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢compare_models(models)

Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "models": {
      "minItems": 2,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh"
    }
  },
  "required": [
    "models"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "models": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "model": {
            "type": "string",
            "description": "Model variant id, e.g. claude-opus-5.5-high"
          },
          "displayName": {
            "type": "string"
          },
          "family": {
            "type": "string"
          },
          "effort": {
            "type": [
              "string",
              "null"
            ]
          },
          "provider": {
            "type": [
              "string",
              "null"
            ]
          },
          "reasoning": {
            "type": "boolean"
          },
          "accuracy": {
            "type": "number",
            "description": "Percentage of puzzles solved"
          },
          "correct": {
            "type": "number"
          },
          "total": {
            "type": "number"
          },
          "failedRuns": {
            "type": "number"
          },
          "bySize": {
            "type": "array",
            "items": {
              "type": "object",
              "properties": {
                "size": {
                  "type": "string"
                }
              },
              "required": [
                "size"
              ],
              "additionalProperties": {}
            }
          }
        },
        "required": [
          "model",
          "displayName",
          "family",
          "effort",
          "provider",
          "reasoning",
          "accuracy",
          "correct",
          "total",
          "failedRuns",
          "bySize"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "models"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢get_model_results(model)

Accuracy, cost, latency and token use for one model, broken down by grid size.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "Model name as listed on the leaderboard, e.g. gpt-5.4-xhigh"
    }
  },
  "required": [
    "model"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "Model variant id, e.g. claude-opus-5.5-high"
    },
    "displayName": {
      "type": "string"
    },
    "family": {
      "type": "string"
    },
    "effort": {
      "type": [
        "string",
        "null"
      ]
    },
    "provider": {
      "type": [
        "string",
        "null"
      ]
    },
    "reasoning": {
      "type": "boolean"
    },
    "accuracy": {
      "type": "number",
      "description": "Percentage of puzzles solved"
    },
    "correct": {
      "type": "number"
    },
    "total": {
      "type": "number"
    },
    "failedRuns": {
      "type": "number"
    },
    "bySize": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "size": {
            "type": "string"
          }
        },
        "required": [
          "size"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "model",
    "displayName",
    "family",
    "effort",
    "provider",
    "reasoning",
    "accuracy",
    "correct",
    "total",
    "failedRuns",
    "bySize"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": {}
}
🟢list_puzzles(size)

The benchmark puzzles with their ids and row/column clues.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "size": {
      "description": "Grid size to filter on",
      "type": "string",
      "enum": [
        "5x5",
        "10x10",
        "15x15",
        "20x20"
      ]
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "puzzles": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": {
            "type": "string"
          },
          "index": {
            "type": "number"
          },
          "size": {
            "type": "string"
          },
          "width": {
            "type": "number"
          },
          "height": {
            "type": "number"
          },
          "rowClues": {
            "type": "array",
            "items": {
              "type": "array",
              "items": {
                "type": "number"
              }
            }
          },
          "columnClues": {
            "type": "array",
            "items": {
              "type": "array",
              "items": {
                "type": "number"
              }
            }
          },
          "url": {
            "type": "string"
          }
        },
        "required": [
          "id",
          "index",
          "size",
          "width",
          "height",
          "rowClues",
          "columnClues",
          "url"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "puzzles"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢get_puzzle(id, include_solution)

One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "Puzzle id from list_puzzles"
    },
    "include_solution": {
      "description": "Include a reference solution",
      "type": "boolean"
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string"
    },
    "index": {
      "type": "number"
    },
    "size": {
      "type": "string"
    },
    "width": {
      "type": "number"
    },
    "height": {
      "type": "number"
    },
    "rowClues": {
      "type": "array",
      "items": {
        "type": "array",
        "items": {
          "type": "number"
        }
      }
    },
    "columnClues": {
      "type": "array",
      "items": {
        "type": "array",
        "items": {
          "type": "number"
        }
      }
    },
    "url": {
      "type": "string"
    },
    "prompt": {
      "type": "string",
      "description": "Clue text as models received it"
    },
    "referenceSolution": {
      "type": "string"
    }
  },
  "required": [
    "id",
    "index",
    "size",
    "width",
    "height",
    "rowClues",
    "columnClues",
    "url",
    "prompt"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": {}
}
🟢check_solution(id, grid)

Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "Puzzle id from list_puzzles"
    },
    "grid": {
      "type": "string",
      "description": "Row-major string of 0 (empty) and 1 (filled), width × height characters"
    }
  },
  "required": [
    "id",
    "grid"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "correct": {
      "type": "boolean"
    },
    "error": {
      "type": "string"
    },
    "rowViolations": {
      "type": "array",
      "items": {}
    },
    "columnViolations": {
      "type": "array",
      "items": {}
    }
  },
  "required": [
    "correct",
    "rowViolations",
    "columnViolations"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": {}
}
🟢get_puzzle_results(id, provider, family, effort, reasoning, ...)

Per-model outcomes for one puzzle. Answers are omitted unless requested.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "Puzzle id from list_puzzles"
    },
    "provider": {
      "description": "Comma-separated provider ids; empty means no filter",
      "type": "string"
    },
    "family": {
      "description": "Comma-separated family ids; empty means no filter",
      "type": "string"
    },
    "effort": {
      "description": "best, all (default), or one effort level",
      "type": "string"
    },
    "reasoning": {
      "description": "Only reasoning (true) or non-reasoning (false) variants",
      "type": "boolean"
    },
    "open_weights": {
      "description": "Only open-weight (true) or closed (false) models",
      "type": "boolean"
    },
    "include_answers": {
      "description": "Include each model's answer grid",
      "type": "boolean"
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "updatedAt": {
      "type": "string"
    },
    "puzzleId": {
      "type": "string"
    },
    "index": {
      "type": "number"
    },
    "size": {
      "type": "string"
    },
    "attempts": {
      "type": "number"
    },
    "solved": {
      "type": "number"
    },
    "runs": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "model": {
            "type": "string"
          },
          "correct": {
            "type": "boolean"
          },
          "status": {
            "type": "string"
          },
          "displayName": {
            "type": "string"
          },
          "answer": {
            "type": [
              "string",
              "null"
            ]
          }
        },
        "required": [
          "model",
          "correct",
          "status",
          "displayName"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "updatedAt",
    "puzzleId",
    "index",
    "size",
    "attempts",
    "solved",
    "runs"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢get_model_puzzles(model)

Which puzzles one model solved, missed, timed out on, or has not run.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "description": "Model variant id as listed on the leaderboard, e.g. claude-opus-5.5-high"
    }
  },
  "required": [
    "model"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string"
    },
    "displayName": {
      "type": "string"
    },
    "solved": {
      "type": "number"
    },
    "attempted": {
      "type": "number"
    },
    "puzzles": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": {
            "type": "string"
          },
          "index": {
            "type": "number"
          },
          "size": {
            "type": "string"
          },
          "state": {
            "type": "string",
            "enum": [
              "solved",
              "wrong",
              "cut-off",
              "not-run"
            ]
          }
        },
        "required": [
          "id",
          "index",
          "size",
          "state"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "model",
    "displayName",
    "solved",
    "attempted",
    "puzzles"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}
🟢list_runs(model, puzzle_id, size, include_output, limit, ...)

Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "description": "Only runs of this model variant id",
      "type": "string"
    },
    "puzzle_id": {
      "description": "Only runs on this puzzle id from list_puzzles",
      "type": "string"
    },
    "size": {
      "description": "Grid size to filter on",
      "type": "string",
      "enum": [
        "5x5",
        "10x10",
        "15x15",
        "20x20"
      ]
    },
    "include_output": {
      "description": "Include raw prompt and model output (large)",
      "type": "boolean"
    },
    "limit": {
      "description": "Default 100",
      "type": "integer",
      "minimum": 1,
      "maximum": 500
    },
    "offset": {
      "description": "Runs to skip, for paging; default 0",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "total": {
      "type": "number"
    },
    "limit": {
      "type": "number"
    },
    "offset": {
      "type": "number"
    },
    "runs": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "model": {
            "type": "string"
          },
          "puzzleId": {
            "type": "string"
          },
          "size": {
            "type": "string"
          },
          "correct": {
            "type": "boolean"
          },
          "status": {
            "type": "string"
          }
        },
        "required": [
          "model",
          "puzzleId",
          "size",
          "correct",
          "status"
        ],
        "additionalProperties": {}
      }
    }
  },
  "required": [
    "total",
    "limit",
    "offset",
    "runs"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Community

Diesen Server bewerten

Nachweis

Aktuelle Beobachtungen

verifiziertVersion nicht aufgezeichnet11 Tools