Mozilla Data Collective

Search the Mozilla Data Collective catalog of ethically sourced AI training datasets.

使うべきか

品質と安全性

A
説明の品質
100%
スキーマの完全性
77%
命名の品質
100%
ポイズニングのリスク
100%
権限の一致
100%
プロトコルへの準拠
100%

ツール定義とプロトコルへの準拠に関する自動分析に基づいています。

コンテキストコスト

~984トークン数(ツール定義)
~2.6 KB一般的なレスポンスサイズ
注意への影響は中程度(128k コンテキストの 0.77%)

これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。

インストール

ワンクリックインストール

これを `claude_desktop_config.json` ファイルに追加してください:

{
  "mcpServers": {
    "datasets": {
      "url": "https://mozilladatacollective.com/api/mcp"
    }
  }
}

リモートエンドポイント

https://mozilladatacollective.com/api/mcpstreamable-http

できること

ツール一覧

ツール(3)

🟢 読み取り専用🟡 書き込み🔴 削除⚪ 不明
🟢search(query, limit, task, locale, license, ...)

Search the Mozilla Data Collective catalog of AI training datasets by natural-language query, optionally narrowed by task, language, license, format, price, sample availability or publish date. Returns matching datasets as {id, title, url}; pass an id to the fetch tool for full details. Call list_filters first if you intend to filter — filter values must match the catalog exactly.

入力スキーマ

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Natural-language search query describing the datasets you are looking for, e.g. 'Spanish speech recordings for TTS training'. Descriptive phrases retrieve better than single keywords."
    },
    "limit": {
      "default": 10,
      "description": "Maximum number of results to return (1-25).",
      "type": "integer",
      "minimum": 1,
      "maximum": 25
    },
    "task": {
      "description": "Restrict to these machine-learning tasks, e.g. ['ASR', 'TTS'].",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "enum": [
          "N/A",
          "NLP",
          "ASR",
          "LID",
          "TTS",
          "MT",
          "LM",
          "LLM",
          "NLU",
          "NLG",
          "CALL",
          "RAG",
          "CV",
          "ML",
          "OTH"
        ]
      }
    },
    "locale": {
      "description": "Restrict to these language/locale codes, e.g. ['sw', 'pt-BR']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "license": {
      "description": "Restrict to these license abbreviations, e.g. ['CC0-1.0', 'CC-BY-4.0']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "format": {
      "description": "Restrict to these file formats, e.g. ['WAV', 'MP3']. Values must match exactly (case-sensitive); call the list_filters tool to get the valid ones.",
      "minItems": 1,
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1
      }
    },
    "isPaid": {
      "description": "true returns only paid datasets, false only free ones. Omit to include both.",
      "type": "boolean"
    },
    "hasSample": {
      "description": "true returns only datasets that publish a downloadable sample, useful when the user wants to try data before committing. false behaves the same as omitting it.",
      "type": "boolean"
    },
    "sort": {
      "description": "Result ordering. Defaults to 'relevance'; use 'newest' or 'size' only when the user asks for it.",
      "type": "string",
      "enum": [
        "relevance",
        "newest",
        "size"
      ]
    },
    "sortDirection": {
      "description": "Direction for the sort field. Only meaningful alongside sort='newest' or sort='size'.",
      "type": "string",
      "enum": [
        "asc",
        "desc"
      ]
    },
    "uploadDate": {
      "description": "Restrict to datasets published within this recent window.",
      "type": "string",
      "enum": [
        "today",
        "thisWeek",
        "thisMonth",
        "thisYear"
      ]
    }
  },
  "required": [
    "query"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢fetch(id)

Fetch the full public details of one Mozilla Data Collective dataset by id or slug: description, organization, task, locale, license, format, size, pricing, and its page URL.

入力スキーマ

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "maxLength": 500,
      "description": "Dataset id or slug, as returned in the id field of search results."
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢list_filters

List every value the search tool's filters accept: the tasks, locales, licenses and formats present in the catalog, plus the sort and date-range options. Task and license values are abbreviations, so taskLabels and licenseLabels spell them out. Filter values are matched exactly, so call this before filtering a search rather than guessing values. Takes no arguments.

入力スキーマ

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

コミュニティ

このサーバーを評価する

エビデンス

最近の観測

検証済みバージョンは記録されていませんツール 3 件
検証済みバージョンは記録されていませんツール 3 件
検証済みバージョンは記録されていませんツール 3 件