Mozilla Data Collective

Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
77%
Calidad de los nombres
100%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~984Tokens (definiciones de herramientas)
~2.6 KBTamaño de respuesta típico
Impacto moderado en la atención (0.77% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "datasets": {
      "url": "https://mozilladatacollective.com/api/mcp"
    }
  }
}

Puntos de conexión remotos

https://mozilladatacollective.com/api/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (3)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟢search(query, limit, task, locale, license, ...)

Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords."
    },
    "limit": {
      "default": 10,
      "description": "Maximum number of results to return (1-25).",
      "type": "integer",
      "minimum": 1,
      "maximum": 25
    },
    "task": {
      "description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "enum": [
          "N/A",
          "NLP",
          "ASR",
          "LID",
          "TTS",
          "MT",
          "LM",
          "LLM",
          "NLU",
          "NLG",
          "CALL",
          "RAG",
          "CV",
          "ML",
          "OTH"
        ]
      }
    },
    "locale": {
      "description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "license": {
      "description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "format": {
      "description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "isPaid": {
      "description": "true returns only paid datasets, false only free ones. Omit to include both.",
      "type": "boolean"
    },
    "hasSample": {
      "description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
      "type": "boolean"
    },
    "sort": {
      "description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
      "type": "string",
      "enum": [
        "relevance",
        "newest",
        "size"
      ]
    },
    "sortDirection": {
      "description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
      "type": "string",
      "enum": [
        "asc",
        "desc"
      ]
    },
    "uploadDate": {
      "description": "Restrict to datasets published within this recent window.",
      "type": "string",
      "enum": [
        "today",
        "thisWeek",
        "thisMonth",
        "thisYear"
      ]
    }
  },
  "required": [
    "query"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢fetch(id)

Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Dataset id or slug, as returned in the id field of search results."
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_filters

List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.

Esquema de entrada

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada3 herramientas
verificadoversión no registrada3 herramientas
verificadoversión no registrada3 herramientas