Mozilla Data Collective

Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

Sollte ich dies verwenden

Qualität und Sicherheit

A
Qualität der Beschreibung
100%
Vollständigkeit des Schemas
77%
Qualität der Benennung
100%
Risiko der Vergiftung
100%
Übereinstimmung der Berechtigungen
100%
Einhaltung des Protokolls
100%

Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.

Kontextkosten

~984Tokens (Tool-Definitionen)
~2.6 KBTypische Antwortgröße
Mittlere Auswirkung auf die Aufmerksamkeit (0.77% von 128k Kontext)

Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.

Installieren

Installation mit einem Klick

Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:

{
  "mcpServers": {
    "datasets": {
      "url": "https://mozilladatacollective.com/api/mcp"
    }
  }
}

Remote-Endpunkte

https://mozilladatacollective.com/api/mcpstreamable-http

Was es kann

Tool-Inventar

Tools (3)

🟢 Nur lesen🟡 Schreiben🔴 Löschen⚪ Unbekannt
🟢search(query, limit, task, locale, license, ...)

Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords."
    },
    "limit": {
      "default": 10,
      "description": "Maximum number of results to return (1-25).",
      "type": "integer",
      "minimum": 1,
      "maximum": 25
    },
    "task": {
      "description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "enum": [
          "N/A",
          "NLP",
          "ASR",
          "LID",
          "TTS",
          "MT",
          "LM",
          "LLM",
          "NLU",
          "NLG",
          "CALL",
          "RAG",
          "CV",
          "ML",
          "OTH"
        ]
      }
    },
    "locale": {
      "description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "license": {
      "description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "format": {
      "description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "isPaid": {
      "description": "true returns only paid datasets, false only free ones. Omit to include both.",
      "type": "boolean"
    },
    "hasSample": {
      "description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
      "type": "boolean"
    },
    "sort": {
      "description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
      "type": "string",
      "enum": [
        "relevance",
        "newest",
        "size"
      ]
    },
    "sortDirection": {
      "description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
      "type": "string",
      "enum": [
        "asc",
        "desc"
      ]
    },
    "uploadDate": {
      "description": "Restrict to datasets published within this recent window.",
      "type": "string",
      "enum": [
        "today",
        "thisWeek",
        "thisMonth",
        "thisYear"
      ]
    }
  },
  "required": [
    "query"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢fetch(id)

Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Dataset id or slug, as returned in the id field of search results."
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_filters

List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Community

Diesen Server bewerten

Nachweis

Aktuelle Beobachtungen

verifiziertVersion nicht aufgezeichnet3 Tools
verifiziertVersion nicht aufgezeichnet3 Tools