Misata Studio: verified synthetic data

Hosted: plan and build multi-table test and demo data, export 25 formats, with checks shown.

Sollte ich dies verwenden

Qualität und Sicherheit

A
Qualität der Beschreibung
100%
Vollständigkeit des Schemas
66%
Qualität der Benennung
95%
Risiko der Vergiftung
100%
Übereinstimmung der Berechtigungen
100%
Einhaltung des Protokolls
100%

Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.

Kontextkosten

~3,169Tokens (Tool-Definitionen)
~762 BTypische Antwortgröße
Erhebliche Auswirkung auf die Aufmerksamkeit (2.48% von 128k Kontext)

Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.

Installieren

Installation mit einem Klick

Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:

{
  "mcpServers": {
    "misata": {
      "url": "https://api.misata.studio/mcp"
    }
  }
}

Remote-Endpunkte

https://api.misata.studio/mcpstreamable-http

Was es kann

Tool-Inventar

Tools (12)

🟢 Nur lesen🟡 Schreiben🔴 Löschen⚪ Unbekannt
🟢whoami

Which account this connection is, and what it may do: signed in with a Studio MCP key or anonymous, the row cap, and whether a model key is available for plain-English requests.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "title": "whoamiArguments"
}
🟢plan_dataset(request, schema, ddl)

See the tables, sizes and relationships the engine would build, before any rows exist. Free (no rows are made), so it is worth calling before generate_dataset on anything non-trivial: review what it understood and assumed, then adjust your request or schema before spending a real call. Args: request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given — then it is still used to ground realism, e.g. locale and what columns mean). schema: A Misata schema dict (see the server instructions for the format). No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. Returns: route: "chat" (not a dataset request — see `reply`), "design" (the model designed the tables) or "pack" (matched a built-in shape). tables: name, estimated rows, columns, foreign keys for each table the engine would build. understanding: what the engine read the request as (business, archetype, assumptions). requirements: every specific thing the request asked for, so you can see what was understood.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "request": {
      "default": "",
      "title": "Request",
      "type": "string"
    },
    "schema": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Schema"
    },
    "ddl": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Ddl"
    }
  },
  "title": "plan_datasetArguments"
}
🟢blueprint_guide

The reference for the engine's full design language (the `blueprint` argument of generate_dataset, start_generation and validate_blueprint): roles, distributions, formulas that read parent columns and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "title": "blueprint_guideArguments"
}

Ausgabe-Schema

{
  "type": "object",
  "properties": {
    "result": {
      "title": "Result",
      "type": "string"
    }
  },
  "required": [
    "result"
  ],
  "title": "blueprint_guideOutput"
}
🟢validate_blueprint(blueprint, preview)

Check a blueprint before generating it: every error said as what to change, the design traps it falls into (a column that will round away, a copy of a value not made yet, a window that counts nothing), its size against your row cap, and a small preview run (a few thousand rows) with the verifier's findings and sample rows, so you can see the data behave before the real call. Free. Returns: valid, errors (fix these), warnings (read these), estimated_rows, row_cap, preview: {passed, findings, tables: {name: first rows}} when `preview`.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "blueprint": {
      "additionalProperties": true,
      "title": "Blueprint",
      "type": "object"
    },
    "preview": {
      "default": true,
      "title": "Preview",
      "type": "boolean"
    }
  },
  "required": [
    "blueprint"
  ],
  "title": "validate_blueprintArguments"
}
🟡generate_dataset(request, schema, ddl, seed, research, ...)

Make a verified relational dataset and wait for it: every foreign key checked, dates correctly ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed. Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine designs the tables with a model, which can take several minutes: call start_generation instead and poll get_status, so the call does not sit open and time out. Args: request: Plain-English description (needs an LLM key unless `schema`/`ddl` is also given). schema: A Misata schema dict for exact structural control. No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. seed: Reproducibility seed — the same request/schema and seed always produce the same rows. research: Ground realistic numbers (prices, growth rates) in real published facts via a web search. Off by default (an anonymous caller's request should not trigger external calls unless asked for); needs a key regardless of `schema`/`ddl`. blueprint: The engine's full design language (blueprint_guide has the reference): readings around a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact aggregates. Use it whenever the data must behave like the real process. No key needed; run validate_blueprint on it first. Returns: dataset_id: pass this to get_certificate / query_dataset / export_dataset. passed: whether every check held. A dataset that did not pass is still returned, with `certificate.findings` saying what failed — inspect before trusting it. verification: a short, plain-language account of what was checked and what held — show it to the person as the proof, in place of asking them to take the data on trust. tables: {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every row is in the stored dataset, reachable by query_dataset/export_dataset. certificate: the short form (claims, requirements, findings). get_certificate returns the rest (patterns, realism scorecard, every proof chart).

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "request": {
      "default": "",
      "title": "Request",
      "type": "string"
    },
    "schema": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Schema"
    },
    "ddl": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Ddl"
    },
    "seed": {
      "default": 42,
      "title": "Seed",
      "type": "integer"
    },
    "research": {
      "default": false,
      "title": "Research",
      "type": "boolean"
    },
    "blueprint": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Blueprint"
    }
  },
  "title": "generate_datasetArguments"
}
🟢start_generation(request, schema, ddl, seed, research, ...)

Start a generation in the background and return a job_id at once. Use it for any plain-English `request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every 20-30 seconds until it says done. Arguments are the same as generate_dataset. Returns: job_id, status "running". get_status gives the stage and, when finished, the dataset.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "request": {
      "default": "",
      "title": "Request",
      "type": "string"
    },
    "schema": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Schema"
    },
    "ddl": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Ddl"
    },
    "seed": {
      "default": 42,
      "title": "Seed",
      "type": "integer"
    },
    "research": {
      "default": false,
      "title": "Research",
      "type": "boolean"
    },
    "blueprint": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Blueprint"
    }
  },
  "title": "start_generationArguments"
}
🟢get_status(job_id)

Where a start_generation job is. While running: the stage it is in and how long it has run. When done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate).

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "job_id": {
      "title": "Job Id",
      "type": "string"
    }
  },
  "required": [
    "job_id"
  ],
  "title": "get_statusArguments"
}
🔴cancel_generation(job_id)

Stop a running start_generation job. It stops at the next stage boundary and keeps nothing.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "job_id": {
      "title": "Job Id",
      "type": "string"
    }
  },
  "required": [
    "job_id"
  ],
  "title": "cancel_generationArguments"
}
🟢get_certificate(dataset_id)

The full certificate for a dataset made by generate_dataset: every claim (stated vs actual), every requirement's status and evidence, realism findings, planted defects/anomalies if any, and the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "dataset_id": {
      "title": "Dataset Id",
      "type": "string"
    }
  },
  "required": [
    "dataset_id"
  ],
  "title": "get_certificateArguments"
}
🟢query_dataset(dataset_id, sql, limit)

Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view over that dataset's own files — nothing else on the server is reachable this way). Args: dataset_id: from a prior generate_dataset call. sql: a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or statements that write or read outside the dataset (checked before running). limit: rows returned (capped at 5000). Returns: columns, rows, truncated (whether more rows existed than `limit`).

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "dataset_id": {
      "title": "Dataset Id",
      "type": "string"
    },
    "sql": {
      "title": "Sql",
      "type": "string"
    },
    "limit": {
      "default": 1000,
      "title": "Limit",
      "type": "integer"
    }
  },
  "required": [
    "dataset_id",
    "sql"
  ],
  "title": "query_datasetArguments"
}
🟡export_dataset(dataset_id, format, dialect, inline)

Export a generated dataset as a file. Returns a `download_url` the person can open (or you can fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link rather than pasting file contents into the chat. Args: dataset_id: from a prior generate_dataset call. format: data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql. code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy, typescript, jsonschema, expectations, django, openapi, mockapi, demo. `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a real database: run the returned SQL through your own database connection, since this server never holds a database credential itself. dialect: for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake. inline: also return the file itself as `base64` (only for files under a few MB). Use it when you must write the file yourself and cannot fetch a URL. Returns: filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when `inline` and small enough.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "dataset_id": {
      "title": "Dataset Id",
      "type": "string"
    },
    "format": {
      "default": "csv",
      "title": "Format",
      "type": "string"
    },
    "dialect": {
      "default": "postgres",
      "title": "Dialect",
      "type": "string"
    },
    "inline": {
      "default": false,
      "title": "Inline",
      "type": "boolean"
    }
  },
  "required": [
    "dataset_id"
  ],
  "title": "export_datasetArguments"
}
🟢find_ready_dataset(query)

The ready-made datasets Misata publishes: free sample databases (direct download, public domain) and premium datasets with a full answer key (a free preview, then a one-off price). Check this first when someone wants sample, demo, practice or teaching data for a common scenario (retail, e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML, clinical, network security): handing over a dataset that already exists is instant. If none fits their tables, generate one instead. Args: query: what the person needs, in their words. Only orders the list (closest first); every dataset is still returned, so judge the fit yourself from the tables and summary. Returns: datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and download_url (free) or free_preview_url + buy_url + price_usd (premium)}].

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "query": {
      "default": "",
      "title": "Query",
      "type": "string"
    }
  },
  "title": "find_ready_datasetArguments"
}

Community

Diesen Server bewerten

Nachweis

Aktuelle Beobachtungen

verifiziertVersion nicht aufgezeichnet12 Tools