Nonobench
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
Sollte ich dies verwenden
Qualität und Sicherheit
Befunde (7)
- LOWin get_leaderboard
- LOWin list_families
- LOWin compare_models
- LOWin get_model_results
- LOWin list_puzzles
- LOWin get_puzzle
- LOWin get_puzzle_results
Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.
Kontextkosten
Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.
Installieren
Installation mit einem Klick
Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:
{
"mcpServers": {
"nonobench": {
"url": "https://www.nonobench.com/mcp"
}
}
}Remote-Endpunkte
https://www.nonobench.com/mcpstreamable-httpWas es kann
Tool-Inventar
Tools (11)
🟢get_leaderboard(size, provider, family, version, effort, ...)
Models ranked by accuracy; defaults to all effort levels for compatibility.
Eingabe-Schema
{
"type": "object",
"properties": {
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
},
"provider": {
"description": "Comma-separated provider ids; empty means no filter",
"type": "string"
},
"family": {
"description": "Comma-separated family ids; empty means no filter",
"type": "string"
},
"version": {
"description": "Comma-separated benchmark versions: 1.0, 1.1, 1.2; empty means all",
"type": "string"
},
"effort": {
"description": "best, all (default), or one effort level; empty means all",
"type": "string"
},
"reasoning": {
"description": "Only reasoning (true) or non-reasoning (false) variants",
"type": "boolean"
},
"open_weights": {
"description": "Only open-weight (true) or closed (false) models",
"type": "boolean"
},
"min_correct": {
"description": "Minimum puzzles solved in the selected tier; default 0 includes unsolved variants",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"updatedAt": {
"type": "string"
},
"size": {
"type": "string"
},
"models": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"rank": {
"type": "number"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"totalCostUsd": {
"type": "number"
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"rank",
"correct",
"total",
"totalCostUsd"
],
"additionalProperties": {}
}
}
},
"required": [
"updatedAt",
"size",
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_providers
Provider ids, names, families and variant counts.
Eingabe-Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"providers": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"name": {
"type": "string"
},
"variantCount": {
"type": "number"
},
"families": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"id",
"name",
"variantCount",
"families"
],
"additionalProperties": {}
}
}
},
"required": [
"providers"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_families
Model families, available efforts and best variants.
Eingabe-Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"families": {
"type": "array",
"items": {
"type": "object",
"properties": {
"family": {
"type": "string"
},
"displayName": {
"type": "string"
},
"provider": {
"type": [
"string",
"null"
]
},
"bestVariant": {
"type": "string"
},
"efforts": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"family",
"displayName",
"provider",
"bestVariant",
"efforts"
],
"additionalProperties": {}
}
}
},
"required": [
"families"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢compare_models(models)
Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.
Eingabe-Schema
{
"type": "object",
"properties": {
"models": {
"minItems": 2,
"maxItems": 20,
"type": "array",
"items": {
"type": "string"
},
"description": "2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh"
}
},
"required": [
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"models": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"failedRuns": {
"type": "number"
},
"bySize": {
"type": "array",
"items": {
"type": "object",
"properties": {
"size": {
"type": "string"
}
},
"required": [
"size"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"correct",
"total",
"failedRuns",
"bySize"
],
"additionalProperties": {}
}
}
},
"required": [
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_model_results(model)
Accuracy, cost, latency and token use for one model, broken down by grid size.
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name as listed on the leaderboard, e.g. gpt-5.4-xhigh"
}
},
"required": [
"model"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"failedRuns": {
"type": "number"
},
"bySize": {
"type": "array",
"items": {
"type": "object",
"properties": {
"size": {
"type": "string"
}
},
"required": [
"size"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"correct",
"total",
"failedRuns",
"bySize"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢list_puzzles(size)
The benchmark puzzles with their ids and row/column clues.
Eingabe-Schema
{
"type": "object",
"properties": {
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"puzzles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"width": {
"type": "number"
},
"height": {
"type": "number"
},
"rowClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"columnClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"url": {
"type": "string"
}
},
"required": [
"id",
"index",
"size",
"width",
"height",
"rowClues",
"columnClues",
"url"
],
"additionalProperties": {}
}
}
},
"required": [
"puzzles"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_puzzle(id, include_solution)
One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.
Eingabe-Schema
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"include_solution": {
"description": "Include a reference solution",
"type": "boolean"
}
},
"required": [
"id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"width": {
"type": "number"
},
"height": {
"type": "number"
},
"rowClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"columnClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"url": {
"type": "string"
},
"prompt": {
"type": "string",
"description": "Clue text as models received it"
},
"referenceSolution": {
"type": "string"
}
},
"required": [
"id",
"index",
"size",
"width",
"height",
"rowClues",
"columnClues",
"url",
"prompt"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢check_solution(id, grid)
Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.
Eingabe-Schema
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"grid": {
"type": "string",
"description": "Row-major string of 0 (empty) and 1 (filled), width × height characters"
}
},
"required": [
"id",
"grid"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"correct": {
"type": "boolean"
},
"error": {
"type": "string"
},
"rowViolations": {
"type": "array",
"items": {}
},
"columnViolations": {
"type": "array",
"items": {}
}
},
"required": [
"correct",
"rowViolations",
"columnViolations"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢get_puzzle_results(id, provider, family, effort, reasoning, ...)
Per-model outcomes for one puzzle. Answers are omitted unless requested.
Eingabe-Schema
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"provider": {
"description": "Comma-separated provider ids; empty means no filter",
"type": "string"
},
"family": {
"description": "Comma-separated family ids; empty means no filter",
"type": "string"
},
"effort": {
"description": "best, all (default), or one effort level",
"type": "string"
},
"reasoning": {
"description": "Only reasoning (true) or non-reasoning (false) variants",
"type": "boolean"
},
"open_weights": {
"description": "Only open-weight (true) or closed (false) models",
"type": "boolean"
},
"include_answers": {
"description": "Include each model's answer grid",
"type": "boolean"
}
},
"required": [
"id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"updatedAt": {
"type": "string"
},
"puzzleId": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"attempts": {
"type": "number"
},
"solved": {
"type": "number"
},
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string"
},
"correct": {
"type": "boolean"
},
"status": {
"type": "string"
},
"displayName": {
"type": "string"
},
"answer": {
"type": [
"string",
"null"
]
}
},
"required": [
"model",
"correct",
"status",
"displayName"
],
"additionalProperties": {}
}
}
},
"required": [
"updatedAt",
"puzzleId",
"index",
"size",
"attempts",
"solved",
"runs"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_model_puzzles(model)
Which puzzles one model solved, missed, timed out on, or has not run.
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id as listed on the leaderboard, e.g. claude-opus-5.5-high"
}
},
"required": [
"model"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"model": {
"type": "string"
},
"displayName": {
"type": "string"
},
"solved": {
"type": "number"
},
"attempted": {
"type": "number"
},
"puzzles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"state": {
"type": "string",
"enum": [
"solved",
"wrong",
"cut-off",
"not-run"
]
}
},
"required": [
"id",
"index",
"size",
"state"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"solved",
"attempted",
"puzzles"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_runs(model, puzzle_id, size, include_output, limit, ...)
Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.
Eingabe-Schema
{
"type": "object",
"properties": {
"model": {
"description": "Only runs of this model variant id",
"type": "string"
},
"puzzle_id": {
"description": "Only runs on this puzzle id from list_puzzles",
"type": "string"
},
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
},
"include_output": {
"description": "Include raw prompt and model output (large)",
"type": "boolean"
},
"limit": {
"description": "Default 100",
"type": "integer",
"minimum": 1,
"maximum": 500
},
"offset": {
"description": "Runs to skip, for paging; default 0",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Ausgabe-Schema
{
"type": "object",
"properties": {
"total": {
"type": "number"
},
"limit": {
"type": "number"
},
"offset": {
"type": "number"
},
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string"
},
"puzzleId": {
"type": "string"
},
"size": {
"type": "string"
},
"correct": {
"type": "boolean"
},
"status": {
"type": "string"
}
},
"required": [
"model",
"puzzleId",
"size",
"correct",
"status"
],
"additionalProperties": {}
}
}
},
"required": [
"total",
"limit",
"offset",
"runs"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}Community
Nachweis