Dataset Cleaner & Exporter

Dedupe, flatten and clean messy JSON rows (emails, phones, URLs, HTML) in one call, as JSON or CSV.

사용해야 할까요

품질 및 안전성

A
설명 품질
100%
스키마 완전성
64%
이름 품질
90%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~1,055토큰 (도구 정의)
~4.8 KB일반적인 응답 크기
중간 정도의 주의 영향 (128k 컨텍스트의 0.82%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "dataset-cleaner-exporter": {
      "url": "https://dataset-cleaner-exporter.nerolabs.workers.dev/mcp"
    }
  }
}

원격 엔드포인트

https://dataset-cleaner-exporter.nerolabs.workers.dev/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (2)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢list_capabilities

Returns the exact cleaning rules (how emails, phone numbers and URLs are detected and normalized), the dedup modes and keep strategies, the order the steps run in, and the maximum rows per call. Call this first if you are unsure how a field will be treated. Free, processes no data.

입력 스키마

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
🔴clean_rows(rows, dedupMode, dedupKeys, keepStrategy, similarityThreshold, ...)

Deduplicates, flattens and cleans a list of JSON rows in one call and returns spreadsheet-ready rows (or CSV text) plus a summary with exact counts: rows in, rows added by expansion, duplicates removed, rows dropped by maxItems, rows out, the final column list, per-column fill rates and warnings. Steps, in order: optionally explode one array field into one row per entry; flatten nested objects into columns (address.city becomes address_city); trim text; lowercase valid emails; reduce phone numbers to digits with any leading +; lowercase URL hosts and drop the trailing slash; optionally strip HTML and turn numeric or true/false text into numbers and booleans; blank text becomes null; keep, remove or rename columns; then remove duplicates (normalized by default, comparing the whole row unless dedupKeys is set) keeping the most complete row. Deterministic, no AI, nothing guessed. Use it on scraped leads, CRM exports or API results before loading them anywhere. There is a row limit per call (see list_capabilities); split bigger lists across several calls.

입력 스키마

{
  "type": "object",
  "properties": {
    "rows": {
      "type": "array",
      "description": "The rows to clean. Each row is a JSON object; keys may differ between rows and values may be nested.",
      "items": {
        "type": "object"
      }
    },
    "dedupMode": {
      "type": "string",
      "enum": [
        "none",
        "exact",
        "normalized",
        "fuzzy"
      ],
      "description": "How duplicates are found. normalized (default) ignores case and whitespace; exact needs identical values; fuzzy also merges near-duplicates (up to 100 rows and 1000 characters of key text, so name a short field in dedupKeys); none keeps every row."
    },
    "dedupKeys": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "Fields that identify a duplicate, for example [\"Email\"]. Empty compares the whole row. Use the final column names: flattened (Details_founded) and renamed. Exact and case-sensitive. Rows where every key is empty count as duplicates of each other."
    },
    "keepStrategy": {
      "type": "string",
      "enum": [
        "most_complete",
        "first",
        "last"
      ],
      "description": "Which duplicate survives: most_complete (default, fewest empty fields), first or last."
    },
    "similarityThreshold": {
      "type": "number",
      "minimum": 0.5,
      "maximum": 0.99,
      "description": "Fuzzy mode only. 0.5 to 0.99, default 0.9. Higher is stricter."
    },
    "flatten": {
      "type": "boolean",
      "description": "Default true. Turn nested objects into flat columns. Arrays become one JSON-text cell."
    },
    "flattenSeparator": {
      "type": "string",
      "description": "Joins nested key paths when flattening. Default \"_\"."
    },
    "expandArrayField": {
      "type": "string",
      "description": "Optional. One top-level array field (for example \"offers\") to explode into one row per entry, repeating the other fields. Object entries become columns. The expanded total must stay within the row limit."
    },
    "cleanFields": {
      "type": "boolean",
      "description": "Default true. Normalize emails, phone numbers and URLs, detected by field name or value shape."
    },
    "stripHtml": {
      "type": "boolean",
      "description": "Default false. Remove HTML tags and decode common entities in text."
    },
    "coerceTypes": {
      "type": "boolean",
      "description": "Default false. Turn \"42\" into 42 and \"true\" into true. Leading-zero values like \"007\" stay text."
    },
    "emptyToNull": {
      "type": "boolean",
      "description": "Default true. Blank text becomes null."
    },
    "dropEmptyFields": {
      "type": "boolean",
      "description": "Default false. Remove null and empty fields from each row."
    },
    "columnsToKeep": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "Keep only these columns (flattened names). Takes priority over columnsToRemove."
    },
    "columnsToRemove": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "Drop these columns (flattened names). Ignored if columnsToKeep is set."
    },
    "columnRenameMap": {
      "description": "Rename columns after keep/remove, as [\"oldName:newName\"] or {\"oldName\":\"newName\"}, for example {\"Details_founded\":\"founded\"}.",
      "anyOf": [
        {
          "type": "array",
          "items": {
            "type": "string"
          }
        },
        {
          "type": "object",
          "additionalProperties": {
            "type": "string"
          }
        }
      ]
    },
    "maxItems": {
      "type": "integer",
      "minimum": 0,
      "description": "Optional cap: read at most this many rows and return at most this many. 0 (default) means no cap."
    },
    "outputFormat": {
      "type": "string",
      "enum": [
        "json",
        "csv"
      ],
      "description": "json (default) returns rows; csv returns the same result as CSV text in \"csv\" instead."
    }
  },
  "required": [
    "rows"
  ],
  "additionalProperties": false
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 2개
검증됨버전이 기록되지 않음도구 2개