Mozilla Data Collective
Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.
Should I use this
Quality & Safety
Based on automated analysis of tool definitions and protocol compliance.
Context Cost
This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.
Install
One-Click Install
Add this to your `claude_desktop_config.json` file:
{
"mcpServers": {
"datasets": {
"url": "https://mozilladatacollective.com/api/mcp"
}
}
}Remote endpoints
https://mozilladatacollective.com/api/mcpstreamable-httpWhat it can do
Tool inventory
Tools (3)
🟢search(query, limit, task, locale, license, ...)
Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.
Input Schema
{
"type": "object",
"properties": {
"query": {
"type": "string",
"minLength": 1,
"maxLength": 500,
"description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords."
},
"limit": {
"default": 10,
"description": "Maximum number of results to return (1-25).",
"type": "integer",
"minimum": 1,
"maximum": 25
},
"task": {
"description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
"minItems": 1,
"maxItems": 20,
"type": "array",
"items": {
"type": "string",
"enum": [
"N/A",
"NLP",
"ASR",
"LID",
"TTS",
"MT",
"LM",
"LLM",
"NLU",
"NLG",
"CALL",
"RAG",
"CV",
"ML",
"OTH"
]
}
},
"locale": {
"description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"minItems": 1,
"maxItems": 20,
"type": "array",
"items": {
"type": "string",
"minLength": 1
}
},
"license": {
"description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"minItems": 1,
"maxItems": 20,
"type": "array",
"items": {
"type": "string",
"minLength": 1
}
},
"format": {
"description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
"minItems": 1,
"maxItems": 20,
"type": "array",
"items": {
"type": "string",
"minLength": 1
}
},
"isPaid": {
"description": "true returns only paid datasets, false only free ones. Omit to include both.",
"type": "boolean"
},
"hasSample": {
"description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
"type": "boolean"
},
"sort": {
"description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
"type": "string",
"enum": [
"relevance",
"newest",
"size"
]
},
"sortDirection": {
"description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
"type": "string",
"enum": [
"asc",
"desc"
]
},
"uploadDate": {
"description": "Restrict to datasets published within this recent window.",
"type": "string",
"enum": [
"today",
"thisWeek",
"thisMonth",
"thisYear"
]
}
},
"required": [
"query"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢fetch(id)
Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.
Input Schema
{
"type": "object",
"properties": {
"id": {
"type": "string",
"minLength": 1,
"maxLength": 500,
"description": "Dataset id or slug, as returned in the id field of search results."
}
},
"required": [
"id"
],
"$schema": "https://json-schema.org/draft/2020-12/schema"
}🟢list_filters
List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.
Input Schema
{
"type": "object",
"properties": {},
"$schema": "https://json-schema.org/draft/2020-12/schema"
}Community
Evidence