doc.page PDF Extraction

Extract PDFs to Markdown, RAG chunks and cited tables; publish tracked Doc Links with read stats.

사용해야 할까요

품질 및 안전성

A
설명 품질
96%
스키마 완전성
87%
이름 품질
94%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

발견 사항 (1)

  • LOWTool 'list_tables' description lacks action verblist_tables에서

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~996토큰 (도구 정의)
~992 B일반적인 응답 크기
중간 정도의 주의 영향 (128k 컨텍스트의 0.78%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "pdf-extract": {
      "url": "https://doc.page/api/mcp"
    }
  }
}

원격 엔드포인트

https://doc.page/api/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (7)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢extract_pdf(url, outputs, mode, chunkTokens)

Extract a PDF into clean Markdown and structured elements (headings, paragraphs). Returns the canonical ExtractedDocument object. mode "hybrid" runs a heavier semantic engine that also reconstructs tables and bounding boxes; the default "fast" engine is prose-only (low confidence.tables).

입력 스키마

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "http(s) URL of the PDF to extract."
    },
    "outputs": {
      "type": "array",
      "description": "Subset of outputs to include. Default: markdown and elements.",
      "items": {
        "type": "string",
        "enum": [
          "markdown",
          "elements",
          "chunks",
          "images"
        ]
      }
    },
    "mode": {
      "type": "string",
      "enum": [
        "fast",
        "hybrid"
      ],
      "description": "fast = prose engine. hybrid = semantic engine with tables + bounding boxes when deployed; falls back to fast with a warning otherwise."
    },
    "chunkTokens": {
      "type": "integer",
      "description": "Target chunk size in tokens (when chunks are requested). Default 512."
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}
🟢get_chunks(url, maxTokens)

Split a PDF into semantic chunks ready for embeddings (RAG). Each chunk carries its text, estimated tokens, starting page, section heading and the source element ids for citation.

입력 스키마

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "http(s) URL of the PDF to chunk."
    },
    "maxTokens": {
      "type": "integer",
      "description": "Target chunk size in tokens. Default 512."
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}
🟢list_tables(url)

Return every table in a PDF as structured JSON (reconstructed rows and columns) with page and bounding box for verifiable citations. Uses the semantic (hybrid) engine.

입력 스키마

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "http(s) URL of the PDF."
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}
🟡create_doc_link(url, name, slug, expiresAt, notifyOnOpen)

Publish a PDF as a tracked doc.page Doc Link and get back a shareable URL. The link belongs to the API key's account and also appears in its doc.page library. Requires an API key. Free plan: up to 3 active links; custom vanity slugs are premium-only. Optional expiry and open-notification toggle.

입력 스키마

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "http(s) URL of the PDF to publish (max 25 MB)."
    },
    "name": {
      "type": "string",
      "description": "Display name in the library. Defaults to the filename."
    },
    "slug": {
      "type": "string",
      "description": "Custom vanity slug (premium plans only). Lowercase letters, digits and hyphens."
    },
    "expiresAt": {
      "type": "string",
      "description": "ISO 8601 date-time after which the link stops working. Omit for no expiry."
    },
    "notifyOnOpen": {
      "type": "boolean",
      "description": "Email the account owner on the first open. Default true."
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}
🟢list_doc_links

List the Doc Links of the API key's account (id, slug, URL, name, disabled/expiry state, total views, last view). Use this to recover links created in earlier sessions before querying stats. Requires an API key.

입력 스키마

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
🟢get_doc_link_stats(id, slug, include)

Reading analytics for one Doc Link of the API key's account, by id or slug. Always returns the summary (total views, unique visitors, last visit). Premium plans additionally get countries, visitor companies (as_org) and per-page views + average dwell time; pass include:["visits"] for the recent visit rows. Requires an API key.

입력 스키마

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "Doc Link item id (from create_doc_link or list_doc_links)."
    },
    "slug": {
      "type": "string",
      "description": "Doc Link slug — alternative to id."
    },
    "include": {
      "type": "array",
      "description": "Extra sections. \"visits\" adds the recent visit rows (enriched on premium plans).",
      "items": {
        "type": "string",
        "enum": [
          "visits"
        ]
      }
    }
  },
  "additionalProperties": false
}
⚪revoke_doc_link(id, slug)

Disable a Doc Link of the API key's account (by id or slug) so the public URL stops serving. The item and its stats remain in the library; on the free plan this frees an active-link slot. Requires an API key.

입력 스키마

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "Doc Link item id."
    },
    "slug": {
      "type": "string",
      "description": "Doc Link slug — alternative to id."
    }
  },
  "additionalProperties": false
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 7개
검증됨버전이 기록되지 않음도구 7개