pdfintact

Extract tables, text and formulas from PDFs, including scanned pages and broken text layers.

使うべきか

品質と安全性

A
説明の品質
100%
スキーマの完全性
70%
命名の品質
100%
ポイズニングのリスク
100%
権限の一致
100%
プロトコルへの準拠
100%

ツール定義とプロトコルへの準拠に関する自動分析に基づいています。

コンテキストコスト

~985トークン数(ツール定義)
~999 B一般的なレスポンスサイズ
注意への影響は中程度(128k コンテキストの 0.77%)

これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。

インストール

ワンクリックインストール

これを `claude_desktop_config.json` ファイルに追加してください:

{
  "mcpServers": {
    "pdfintact": {
      "url": "https://mcp.pdfintact.com/mcp"
    }
  }
}

リモートエンドポイント

https://mcp.pdfintact.com/mcpstreamable-http

できること

ツール一覧

ツール(3)

🟢 読み取り専用🟡 書き込み🔴 削除⚪ 不明
🟢convert_pdf(source, idempotency_key)

Convert a PDF into structured content (tables, charts, formulas, headings, body text) using a two-stage pipeline (layout detection, then a vision-language model) rather than a single VLM call on the raw PDF -- calling a VLM on a raw PDF directly is a known-unreliable pattern for numeric tables. Measured accuracy (500-page real-world benchmark of government/corporate reports, ~51,000 table values checked): tables 95.2% digit-exact, body text 88.8%. This tool reads PDFs a VLM cannot read directly, including scanned pages and PDFs with corrupted/garbled text layers (common in older Japanese academic PDFs). For scanned Japanese documents the numbers hold up (99.4% on the same benchmark). For scanned Arabic, body text does NOT: characters are dropped mid-sentence and quantities can turn into different quantities, so body blocks from scanned Arabic are always flagged confidence:"estimated" -- tables in the same documents stayed exact in our measurement. Strong on Japanese-language documents specifically; the accuracy figures above were measured on Japanese material and are not a claim about every language. Chart values are extracted but are best-effort estimates (about 52% exact match, excluding axis tick labels) and are always flagged confidence:"estimated" in the result -- do not treat estimated chart numbers as authoritative. This is a PAID, ASYNCHRONOUS, per-page-billed operation: credits are reserved from the caller's PDFIntact balance before processing starts, and the response's _meta.credits_remaining shows the balance right after reservation. Processing takes real wall-clock time (roughly 7 seconds/page; a 500-page PDF takes about 42 minutes including a multi-minute cold start), so this tool returns a job_handle immediately without waiting -- call get_result with that job_handle to poll for completion instead of calling convert_pdf again. Always pass idempotency_key; reuse the exact same value if you retry the same request, otherwise retries can double-charge and double-process. Provide the PDF either as a public https URL (source.type="url", up to ~200MB) or inline base64 (source.type="base64", up to ~20MB) -- prefer the URL form for large files. Requires sign-in (OAuth): this session is not authenticated, so calling this tool will fail until the PDFIntact account is connected and authorized.

入力スキーマ

{
  "type": "object",
  "properties": {
    "source": {
      "oneOf": [
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "url"
            },
            "url": {
              "type": "string",
              "format": "uri",
              "description": "HTTPS URL to fetch the PDF from"
            }
          },
          "required": [
            "type",
            "url"
          ]
        },
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "base64"
            },
            "data": {
              "type": "string",
              "description": "Base64-encoded PDF file contents"
            }
          },
          "required": [
            "type",
            "data"
          ]
        }
      ]
    },
    "idempotency_key": {
      "type": "string",
      "minLength": 1,
      "maxLength": 200,
      "description": "Unique key for this request. Reuse the same value on retry of the same PDF to avoid double charging."
    }
  },
  "required": [
    "source",
    "idempotency_key"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢get_result(job_handle)

Fetch the status and (once finished) the structured result of a job previously created by convert_pdf. Poll this with the job_handle convert_pdf returned until _meta.state is "done" or "failed" -- do not call convert_pdf again while waiting. job_handle is scoped to the account that created it; handles belonging to a different account are rejected as not found. Requires sign-in (OAuth): this session is not authenticated, so calling this tool will fail until the PDFIntact account is connected and authorized.

入力スキーマ

{
  "type": "object",
  "properties": {
    "job_handle": {
      "type": "string",
      "minLength": 1,
      "description": "The job_handle returned by convert_pdf"
    }
  },
  "required": [
    "job_handle"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢get_balance

Check the current PDFIntact credit balance for the authenticated account (1 page = 1 credit), plus the soonest-expiring credit lot. Useful before submitting a large PDF to convert_pdf, or after an insufficient_credits error to see how many credits are needed and get a purchase link. Requires sign-in (OAuth): this session is not authenticated, so calling this tool will fail until the PDFIntact account is connected and authorized.

入力スキーマ

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

コミュニティ

このサーバーを評価する

エビデンス

最近の観測

検証済みバージョンは記録されていませんツール 3 件
検証済みバージョンは記録されていませんツール 3 件