Mozilla Data Collective

Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

Should I use this

Quality & Safety

A
Description quality
100%
Schema completeness
77%
Naming quality
100%
Poisoning risk
100%
Permission match
100%
Protocol compliance
100%

Based on automated analysis of tool definitions and protocol compliance.

Context Cost

~984Tokens (tool definitions)
~2.6 KBTypical response size
Moderate attention impact (0.77% of 128k context)

This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.

Install

One-Click Install

Add this to your `claude_desktop_config.json` file:

{
  "mcpServers": {
    "datasets": {
      "url": "https://mozilladatacollective.com/api/mcp"
    }
  }
}

Remote endpoints

https://mozilladatacollective.com/api/mcpstreamable-http

What it can do

Tool inventory

Tools (3)

🟢 Read-only🟡 Write🔴 Delete⚪ Unknown
🟢search(query, limit, task, locale, license, ...)

Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.

Input Schema

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords."
    },
    "limit": {
      "default": 10,
      "description": "Maximum number of results to return (1-25).",
      "type": "integer",
      "minimum": 1,
      "maximum": 25
    },
    "task": {
      "description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "enum": [
          "N/A",
          "NLP",
          "ASR",
          "LID",
          "TTS",
          "MT",
          "LM",
          "LLM",
          "NLU",
          "NLG",
          "CALL",
          "RAG",
          "CV",
          "ML",
          "OTH"
        ]
      }
    },
    "locale": {
      "description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "license": {
      "description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "format": {
      "description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "isPaid": {
      "description": "true returns only paid datasets, false only free ones. Omit to include both.",
      "type": "boolean"
    },
    "hasSample": {
      "description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
      "type": "boolean"
    },
    "sort": {
      "description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
      "type": "string",
      "enum": [
        "relevance",
        "newest",
        "size"
      ]
    },
    "sortDirection": {
      "description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
      "type": "string",
      "enum": [
        "asc",
        "desc"
      ]
    },
    "uploadDate": {
      "description": "Restrict to datasets published within this recent window.",
      "type": "string",
      "enum": [
        "today",
        "thisWeek",
        "thisMonth",
        "thisYear"
      ]
    }
  },
  "required": [
    "query"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢fetch(id)

Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.

Input Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Dataset id or slug, as returned in the id field of search results."
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_filters

List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.

Input Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Community

Rate this Server

Evidence

Recent observations

verifiedversion not recorded3 tools
verifiedversion not recorded3 tools