Nonobench
An open-source benchmark of how well LLMs solve nonogram puzzles, from 5x5 to 20x20.
使うべきか
品質と安全性
検出事項(7)
- LOWget_leaderboard 内
- LOWlist_families 内
- LOWcompare_models 内
- LOWget_model_results 内
- LOWlist_puzzles 内
- LOWget_puzzle 内
- LOWget_puzzle_results 内
ツール定義とプロトコルへの準拠に関する自動分析に基づいています。
コンテキストコスト
これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。
インストール
ワンクリックインストール
これを `claude_desktop_config.json` ファイルに追加してください:
{
"mcpServers": {
"nonobench": {
"url": "https://www.nonobench.com/mcp"
}
}
}リモートエンドポイント
https://www.nonobench.com/mcpstreamable-httpできること
ツール一覧
ツール(11)
🟢get_leaderboard(size, provider, family, version, effort, ...)
Models ranked by accuracy; defaults to all effort levels for compatibility.
入力スキーマ
{
"type": "object",
"properties": {
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
},
"provider": {
"description": "Comma-separated provider ids; empty means no filter",
"type": "string"
},
"family": {
"description": "Comma-separated family ids; empty means no filter",
"type": "string"
},
"version": {
"description": "Comma-separated benchmark versions: 1.0, 1.1, 1.2; empty means all",
"type": "string"
},
"effort": {
"description": "best, all (default), or one effort level; empty means all",
"type": "string"
},
"reasoning": {
"description": "Only reasoning (true) or non-reasoning (false) variants",
"type": "boolean"
},
"open_weights": {
"description": "Only open-weight (true) or closed (false) models",
"type": "boolean"
},
"min_correct": {
"description": "Minimum puzzles solved in the selected tier; default 0 includes unsolved variants",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"updatedAt": {
"type": "string"
},
"size": {
"type": "string"
},
"models": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"rank": {
"type": "number"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"totalCostUsd": {
"type": "number"
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"rank",
"correct",
"total",
"totalCostUsd"
],
"additionalProperties": {}
}
}
},
"required": [
"updatedAt",
"size",
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_providers
Provider ids, names, families and variant counts.
入力スキーマ
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"providers": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"name": {
"type": "string"
},
"variantCount": {
"type": "number"
},
"families": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"id",
"name",
"variantCount",
"families"
],
"additionalProperties": {}
}
}
},
"required": [
"providers"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_families
Model families, available efforts and best variants.
入力スキーマ
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"families": {
"type": "array",
"items": {
"type": "object",
"properties": {
"family": {
"type": "string"
},
"displayName": {
"type": "string"
},
"provider": {
"type": [
"string",
"null"
]
},
"bestVariant": {
"type": "string"
},
"efforts": {
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"family",
"displayName",
"provider",
"bestVariant",
"efforts"
],
"additionalProperties": {}
}
}
},
"required": [
"families"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢compare_models(models)
Side-by-side core overall and per-size accuracy, cost, latency and token results for model or family names.
入力スキーマ
{
"type": "object",
"properties": {
"models": {
"minItems": 2,
"maxItems": 20,
"type": "array",
"items": {
"type": "string"
},
"description": "2 to 20 model variant ids or family names, e.g. claude-opus-5.5 or gpt-6-astra-xhigh"
}
},
"required": [
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"models": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"failedRuns": {
"type": "number"
},
"bySize": {
"type": "array",
"items": {
"type": "object",
"properties": {
"size": {
"type": "string"
}
},
"required": [
"size"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"correct",
"total",
"failedRuns",
"bySize"
],
"additionalProperties": {}
}
}
},
"required": [
"models"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_model_results(model)
Accuracy, cost, latency and token use for one model, broken down by grid size.
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model name as listed on the leaderboard, e.g. gpt-5.4-xhigh"
}
},
"required": [
"model"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id, e.g. claude-opus-5.5-high"
},
"displayName": {
"type": "string"
},
"family": {
"type": "string"
},
"effort": {
"type": [
"string",
"null"
]
},
"provider": {
"type": [
"string",
"null"
]
},
"reasoning": {
"type": "boolean"
},
"accuracy": {
"type": "number",
"description": "Percentage of puzzles solved"
},
"correct": {
"type": "number"
},
"total": {
"type": "number"
},
"failedRuns": {
"type": "number"
},
"bySize": {
"type": "array",
"items": {
"type": "object",
"properties": {
"size": {
"type": "string"
}
},
"required": [
"size"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"family",
"effort",
"provider",
"reasoning",
"accuracy",
"correct",
"total",
"failedRuns",
"bySize"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢list_puzzles(size)
The benchmark puzzles with their ids and row/column clues.
入力スキーマ
{
"type": "object",
"properties": {
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"puzzles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"width": {
"type": "number"
},
"height": {
"type": "number"
},
"rowClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"columnClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"url": {
"type": "string"
}
},
"required": [
"id",
"index",
"size",
"width",
"height",
"rowClues",
"columnClues",
"url"
],
"additionalProperties": {}
}
}
},
"required": [
"puzzles"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_puzzle(id, include_solution)
One puzzle, including the clue text models were prompted with. The reference solution is only included on request; some puzzles have several valid solutions.
入力スキーマ
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"include_solution": {
"description": "Include a reference solution",
"type": "boolean"
}
},
"required": [
"id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"width": {
"type": "number"
},
"height": {
"type": "number"
},
"rowClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"columnClues": {
"type": "array",
"items": {
"type": "array",
"items": {
"type": "number"
}
}
},
"url": {
"type": "string"
},
"prompt": {
"type": "string",
"description": "Clue text as models received it"
},
"referenceSolution": {
"type": "string"
}
},
"required": [
"id",
"index",
"size",
"width",
"height",
"rowClues",
"columnClues",
"url",
"prompt"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢check_solution(id, grid)
Check a grid against a puzzle's clues, using the same rule as the benchmark grader. Reports which rows and columns do not match.
入力スキーマ
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"grid": {
"type": "string",
"description": "Row-major string of 0 (empty) and 1 (filled), width × height characters"
}
},
"required": [
"id",
"grid"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"correct": {
"type": "boolean"
},
"error": {
"type": "string"
},
"rowViolations": {
"type": "array",
"items": {}
},
"columnViolations": {
"type": "array",
"items": {}
}
},
"required": [
"correct",
"rowViolations",
"columnViolations"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": {}
}🟢get_puzzle_results(id, provider, family, effort, reasoning, ...)
Per-model outcomes for one puzzle. Answers are omitted unless requested.
入力スキーマ
{
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Puzzle id from list_puzzles"
},
"provider": {
"description": "Comma-separated provider ids; empty means no filter",
"type": "string"
},
"family": {
"description": "Comma-separated family ids; empty means no filter",
"type": "string"
},
"effort": {
"description": "best, all (default), or one effort level",
"type": "string"
},
"reasoning": {
"description": "Only reasoning (true) or non-reasoning (false) variants",
"type": "boolean"
},
"open_weights": {
"description": "Only open-weight (true) or closed (false) models",
"type": "boolean"
},
"include_answers": {
"description": "Include each model's answer grid",
"type": "boolean"
}
},
"required": [
"id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"updatedAt": {
"type": "string"
},
"puzzleId": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"attempts": {
"type": "number"
},
"solved": {
"type": "number"
},
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string"
},
"correct": {
"type": "boolean"
},
"status": {
"type": "string"
},
"displayName": {
"type": "string"
},
"answer": {
"type": [
"string",
"null"
]
}
},
"required": [
"model",
"correct",
"status",
"displayName"
],
"additionalProperties": {}
}
}
},
"required": [
"updatedAt",
"puzzleId",
"index",
"size",
"attempts",
"solved",
"runs"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢get_model_puzzles(model)
Which puzzles one model solved, missed, timed out on, or has not run.
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string",
"description": "Model variant id as listed on the leaderboard, e.g. claude-opus-5.5-high"
}
},
"required": [
"model"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"model": {
"type": "string"
},
"displayName": {
"type": "string"
},
"solved": {
"type": "number"
},
"attempted": {
"type": "number"
},
"puzzles": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string"
},
"index": {
"type": "number"
},
"size": {
"type": "string"
},
"state": {
"type": "string",
"enum": [
"solved",
"wrong",
"cut-off",
"not-run"
]
}
},
"required": [
"id",
"index",
"size",
"state"
],
"additionalProperties": {}
}
}
},
"required": [
"model",
"displayName",
"solved",
"attempted",
"puzzles"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}🟢list_runs(model, puzzle_id, size, include_output, limit, ...)
Individual benchmark runs (one model on one puzzle), optionally with the raw prompt and model output.
入力スキーマ
{
"type": "object",
"properties": {
"model": {
"description": "Only runs of this model variant id",
"type": "string"
},
"puzzle_id": {
"description": "Only runs on this puzzle id from list_puzzles",
"type": "string"
},
"size": {
"description": "Grid size to filter on",
"type": "string",
"enum": [
"5x5",
"10x10",
"15x15",
"20x20"
]
},
"include_output": {
"description": "Include raw prompt and model output (large)",
"type": "boolean"
},
"limit": {
"description": "Default 100",
"type": "integer",
"minimum": 1,
"maximum": 500
},
"offset": {
"description": "Runs to skip, for paging; default 0",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}出力スキーマ
{
"type": "object",
"properties": {
"total": {
"type": "number"
},
"limit": {
"type": "number"
},
"offset": {
"type": "number"
},
"runs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"model": {
"type": "string"
},
"puzzleId": {
"type": "string"
},
"size": {
"type": "string"
},
"correct": {
"type": "boolean"
},
"status": {
"type": "string"
}
},
"required": [
"model",
"puzzleId",
"size",
"correct",
"status"
],
"additionalProperties": {}
}
}
},
"required": [
"total",
"limit",
"offset",
"runs"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}コミュニティ
エビデンス