Misata Studio: verified synthetic data
Hosted: plan and build multi-table test and demo data, export 25 formats, with checks shown.
¿Debería usar esto?
Calidad y seguridad
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"misata": {
"url": "https://api.misata.studio/mcp"
}
}
}Puntos de conexión remotos
https://api.misata.studio/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (12)
🟢whoami
Which account this connection is, and what it may do: signed in with a Studio MCP key or anonymous, the row cap, and whether a model key is available for plain-English requests.
Esquema de entrada
{
"type": "object",
"properties": {},
"title": "whoamiArguments"
}🟢plan_dataset(request, schema, ddl)
See the tables, sizes and relationships the engine would build, before any rows exist. Free (no rows are made), so it is worth calling before generate_dataset on anything non-trivial: review what it understood and assumed, then adjust your request or schema before spending a real call. Args: request: Plain-English description (needs an LLM key, unless `schema`/`ddl` is also given — then it is still used to ground realism, e.g. locale and what columns mean). schema: A Misata schema dict (see the server instructions for the format). No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. Returns: route: "chat" (not a dataset request — see `reply`), "design" (the model designed the tables) or "pack" (matched a built-in shape). tables: name, estimated rows, columns, foreign keys for each table the engine would build. understanding: what the engine read the request as (business, archetype, assumptions). requirements: every specific thing the request asked for, so you can see what was understood.
Esquema de entrada
{
"type": "object",
"properties": {
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
},
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
}
},
"title": "plan_datasetArguments"
}🟢blueprint_guide
The reference for the engine's full design language (the `blueprint` argument of generate_dataset, start_generation and validate_blueprint): roles, distributions, formulas that read parent columns and draw noise, per-group sequences (seq, lag, cumsum, ar1 drift, random walk), aggregates, event windows, cause-and-effect, lifecycles, with patterns. Read it before writing a blueprint.
Esquema de entrada
{
"type": "object",
"properties": {},
"title": "blueprint_guideArguments"
}Esquema de salida
{
"type": "object",
"properties": {
"result": {
"title": "Result",
"type": "string"
}
},
"required": [
"result"
],
"title": "blueprint_guideOutput"
}🟢validate_blueprint(blueprint, preview)
Check a blueprint before generating it: every error said as what to change, the design traps it falls into (a column that will round away, a copy of a value not made yet, a window that counts nothing), its size against your row cap, and a small preview run (a few thousand rows) with the verifier's findings and sample rows, so you can see the data behave before the real call. Free. Returns: valid, errors (fix these), warnings (read these), estimated_rows, row_cap, preview: {passed, findings, tables: {name: first rows}} when `preview`.
Esquema de entrada
{
"type": "object",
"properties": {
"blueprint": {
"additionalProperties": true,
"title": "Blueprint",
"type": "object"
},
"preview": {
"default": true,
"title": "Preview",
"type": "boolean"
}
},
"required": [
"blueprint"
],
"title": "validate_blueprintArguments"
}🟡generate_dataset(request, schema, ddl, seed, research, ...)
Make a verified relational dataset and wait for it: every foreign key checked, dates correctly ordered, declared aggregates and rates exact, before anything is returned. Deterministic for a seed. Use this when you give a `schema` or `ddl` (seconds). For a plain-English `request` the engine designs the tables with a model, which can take several minutes: call start_generation instead and poll get_status, so the call does not sit open and time out. Args: request: Plain-English description (needs an LLM key unless `schema`/`ddl` is also given). schema: A Misata schema dict for exact structural control. No key needed for structure. ddl: CREATE TABLE statements. No key needed for structure. seed: Reproducibility seed — the same request/schema and seed always produce the same rows. research: Ground realistic numbers (prices, growth rates) in real published facts via a web search. Off by default (an anonymous caller's request should not trigger external calls unless asked for); needs a key regardless of `schema`/`ddl`. blueprint: The engine's full design language (blueprint_guide has the reference): readings around a parent's nominal, drift, autocorrelation, cause-and-effect, event windows, exact aggregates. Use it whenever the data must behave like the real process. No key needed; run validate_blueprint on it first. Returns: dataset_id: pass this to get_certificate / query_dataset / export_dataset. passed: whether every check held. A dataset that did not pass is still returned, with `certificate.findings` saying what failed — inspect before trusting it. verification: a short, plain-language account of what was checked and what held — show it to the person as the proof, in place of asking them to take the data on trust. tables: {name: {columns, preview_rows (first 20), total_rows}} — the preview only; every row is in the stored dataset, reachable by query_dataset/export_dataset. certificate: the short form (claims, requirements, findings). get_certificate returns the rest (patterns, realism scorecard, every proof chart).
Esquema de entrada
{
"type": "object",
"properties": {
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
},
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
},
"seed": {
"default": 42,
"title": "Seed",
"type": "integer"
},
"research": {
"default": false,
"title": "Research",
"type": "boolean"
},
"blueprint": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Blueprint"
}
},
"title": "generate_datasetArguments"
}🟢start_generation(request, schema, ddl, seed, research, ...)
Start a generation in the background and return a job_id at once. Use it for any plain-English `request` (a model designs the tables, which takes minutes) — then call get_status(job_id) every 20-30 seconds until it says done. Arguments are the same as generate_dataset. Returns: job_id, status "running". get_status gives the stage and, when finished, the dataset.
Esquema de entrada
{
"type": "object",
"properties": {
"request": {
"default": "",
"title": "Request",
"type": "string"
},
"schema": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Schema"
},
"ddl": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Ddl"
},
"seed": {
"default": 42,
"title": "Seed",
"type": "integer"
},
"research": {
"default": false,
"title": "Research",
"type": "boolean"
},
"blueprint": {
"anyOf": [
{
"additionalProperties": true,
"type": "object"
},
{
"type": "null"
}
],
"default": null,
"title": "Blueprint"
}
},
"title": "start_generationArguments"
}🟢get_status(job_id)
Where a start_generation job is. While running: the stage it is in and how long it has run. When done: the same answer generate_dataset returns (dataset_id, verification, tables, certificate).
Esquema de entrada
{
"type": "object",
"properties": {
"job_id": {
"title": "Job Id",
"type": "string"
}
},
"required": [
"job_id"
],
"title": "get_statusArguments"
}🔴cancel_generation(job_id)
Stop a running start_generation job. It stops at the next stage boundary and keeps nothing.
Esquema de entrada
{
"type": "object",
"properties": {
"job_id": {
"title": "Job Id",
"type": "string"
}
},
"required": [
"job_id"
],
"title": "cancel_generationArguments"
}🟢get_certificate(dataset_id)
The full certificate for a dataset made by generate_dataset: every claim (stated vs actual), every requirement's status and evidence, realism findings, planted defects/anomalies if any, and the pattern charts. This is the whole answer key, not the trimmed form generate_dataset returns.
Esquema de entrada
{
"type": "object",
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
}
},
"required": [
"dataset_id"
],
"title": "get_certificateArguments"
}🟢query_dataset(dataset_id, sql, limit)
Read-only SQL over a dataset's own tables (one SELECT/WITH statement; every table name is a view over that dataset's own files — nothing else on the server is reachable this way). Args: dataset_id: from a prior generate_dataset call. sql: a single SELECT (or WITH ... SELECT) statement. No semicolons, file paths, or statements that write or read outside the dataset (checked before running). limit: rows returned (capped at 5000). Returns: columns, rows, truncated (whether more rows existed than `limit`).
Esquema de entrada
{
"type": "object",
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
},
"sql": {
"title": "Sql",
"type": "string"
},
"limit": {
"default": 1000,
"title": "Limit",
"type": "integer"
}
},
"required": [
"dataset_id",
"sql"
],
"title": "query_datasetArguments"
}🟡export_dataset(dataset_id, format, dialect, inline)
Export a generated dataset as a file. Returns a `download_url` the person can open (or you can fetch, e.g. with curl) for as long as the dataset is held, about 2 hours. Give the person the link rather than pasting file contents into the chat. Args: dataset_id: from a prior generate_dataset call. format: data: csv, parquet, jsonl, json, avro, xlsx, feather, orc, sqlite, duckdb, sql. code and docs: dbt, notebook, dictionary, dbml, mermaid, prisma, sqlalchemy, typescript, jsonschema, expectations, django, openapi, mockapi, demo. `sql` is schema.sql (DDL with keys) + data.sql (COPY/INSERT) — the way to seed a real database: run the returned SQL through your own database connection, since this server never holds a database credential itself. dialect: for `sql` only: postgres, mysql, sqlite, mssql, oracle, bigquery, snowflake. inline: also return the file itself as `base64` (only for files under a few MB). Use it when you must write the file yourself and cannot fetch a URL. Returns: filename, content_type, bytes, download_url, expires_at (unix seconds), and `base64` when `inline` and small enough.
Esquema de entrada
{
"type": "object",
"properties": {
"dataset_id": {
"title": "Dataset Id",
"type": "string"
},
"format": {
"default": "csv",
"title": "Format",
"type": "string"
},
"dialect": {
"default": "postgres",
"title": "Dialect",
"type": "string"
},
"inline": {
"default": false,
"title": "Inline",
"type": "boolean"
}
},
"required": [
"dataset_id"
],
"title": "export_datasetArguments"
}🟢find_ready_dataset(query)
The ready-made datasets Misata publishes: free sample databases (direct download, public domain) and premium datasets with a full answer key (a free preview, then a one-off price). Check this first when someone wants sample, demo, practice or teaching data for a common scenario (retail, e-commerce, SaaS, manufacturing SPC, predictive maintenance, insurance claims, fraud/AML, clinical, network security): handing over a dataset that already exists is instant. If none fits their tables, generate one instead. Args: query: what the person needs, in their words. Only orders the list (closest first); every dataset is still returned, so judge the fit yourself from the tables and summary. Returns: datasets: [{slug, kind (free|premium), title, summary, rows, tables, page_url, and download_url (free) or free_preview_url + buy_url + price_usd (premium)}].
Esquema de entrada
{
"type": "object",
"properties": {
"query": {
"default": "",
"title": "Query",
"type": "string"
}
},
"title": "find_ready_datasetArguments"
}Comunidad
Evidencia