PreteWorks API
One key for the document-to-web pipeline: scrape, AI-extract, build & edit PDFs, fill forms.
Should I use this
Quality & Safety
Findings (8)
- HIGH
- MEDIUMin scrape_page
- LOWin images_to_pdf
- LOWin split_pdf
- LOWin fill_pdf_form
- LOWin extract_data
- LOWin answer_from_document
- LOWin docx_to_text
Based on automated analysis of tool definitions and protocol compliance.
Context Cost
This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.
Install
One-Click Install
Add this to your `claude_desktop_config.json` file:
{
"mcpServers": {
"preteworks-api": {
"url": "https://api.preteworks.com/mcp"
}
}
}Remote endpoints
https://api.preteworks.com/mcpstreamable-httpWhat it can do
Tool inventory
Tools (29)
🟢render_html_to_pdf(html)
Render a full HTML document to a PDF. External subresources are blocked for safety — inline images/fonts as data: URIs. Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"html": {
"type": "string",
"description": "The complete HTML document to render."
}
},
"required": [
"html"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪number_pdf(pdf, position, start, format)
Stamp a page number on every page. position: bottom-center (default) | bottom-left | bottom-right | top-center | top-left | top-right. format supports {n} and {total}. Returns a file_id + ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"position": {
"description": "Where to place it (default bottom-center).",
"type": "string"
},
"start": {
"description": "First page's number (default 1).",
"type": "number"
},
"format": {
"description": "e.g. '{n}' or '{n} / {total}' (default '{n}').",
"type": "string"
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢pdf_info(pdf)
Read a PDF's metadata (title, author, subject, keywords, creator, producer, dates) plus page count and page size, as JSON.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟡set_pdf_metadata(pdf, title, author, subject, keywords, ...)
Set a PDF's title, author, subject, keywords (array or comma string) and/or creator. Only the fields you provide change. Returns a file_id + ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"title": {
"type": "string"
},
"author": {
"type": "string"
},
"subject": {
"type": "string"
},
"keywords": {
"anyOf": [
{
"type": "string"
},
{
"type": "array",
"items": {
"type": "string"
}
}
]
},
"creator": {
"type": "string"
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢render_document(template, data)
Render a structured document to a PDF from a named template (invoice, receipt, report). Money/totals are computed for you. Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"template": {
"type": "string",
"enum": [
"invoice",
"receipt",
"report"
],
"description": "Document template."
},
"data": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {},
"description": "Template data. invoice/receipt: {seller:{name,address?,email?,logo_url?(data: URI)}, buyer?:{name,address?,email?}, number?, issued?/date?, due?, currency?, items:[{description,qty,unit_price}], tax_rate?, notes?, terms?}; receipt also takes amount_paid?, payment_method?. report: {title, subtitle?, author?, date?, sections:[{heading?,body?}]}. Money is computed server-side."
}
},
"required": [
"template",
"data"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢merge_pdfs(pdfs)
Combine 2+ PDFs into a single PDF, in the order given. Inputs are file_ids (from prior tools) or base64 PDFs. Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"pdfs": {
"minItems": 2,
"type": "array",
"items": {
"type": "string"
},
"description": "2+ PDF sources. Each: A file_id from a prior tool result, or a base64-encoded PDF."
}
},
"required": [
"pdfs"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🔴select_pages(pdf, pages)
Keep only the specified pages of a PDF (e.g. [1,3,5]) and drop the rest, preserving order. Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"pages": {
"type": "string",
"description": "Pages to keep, e.g. '1,3,5-7'."
}
},
"required": [
"pdf",
"pages"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢rotate_pdf(pdf, degrees)
Rotate every page of a PDF by 90, 180, or 270 degrees (clockwise). Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"degrees": {
"type": "number",
"description": "90, 180, or 270."
}
},
"required": [
"pdf",
"degrees"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢watermark_pdf(pdf, text)
Stamp a diagonal grey text watermark (e.g. "DRAFT", "CONFIDENTIAL") across every page of a PDF. Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"text": {
"type": "string",
"description": "The watermark text."
}
},
"required": [
"pdf",
"text"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢extract_pdf_text(pdf)
Extract the text content of a PDF — for RAG, summarization, or search. Accepts a file_id (from a prior tool) or a base64-encoded PDF, and returns the text inline. Not OCR: a scanned/image-only PDF returns little or no text.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪images_to_pdf(images)
Combine PNG/JPEG images (base64) into a PDF, one image per page. Returns a file_id and a ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"images": {
"type": "array",
"items": {
"type": "string"
},
"description": "Base64-encoded PNG/JPEG images, one per page (max 20)."
}
},
"required": [
"images"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪markdown_to_pdf(markdown, title)
Render Markdown to a clean, print-styled PDF. Returns a file_id and a ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"markdown": {
"type": "string",
"description": "The Markdown to render."
},
"title": {
"description": "Document title.",
"type": "string"
}
},
"required": [
"markdown"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪split_pdf(pdf, ranges)
Split a PDF into multiple PDFs — one per page by default, or by page ranges like '1-3;4-6'. Returns a file_id + URL per part.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"ranges": {
"description": "e.g. '1-3;4-6'. Omit to split every page.",
"type": "string"
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢read_pdf_form(pdf)
List a PDF's form fields (name, type, value, options) as JSON.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
}
},
"required": [
"pdf"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪fill_pdf_form(pdf, fields, flatten)
Fill a PDF's form fields from a map of field -> value; optionally flatten. Returns a file_id and a ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"pdf": {
"type": "string",
"description": "A file_id from a prior tool result, or a base64-encoded PDF."
},
"fields": {
"type": "object",
"propertyNames": {
"type": "string"
},
"additionalProperties": {},
"description": "Map of field name to value."
},
"flatten": {
"description": "Flatten the form after filling (default false).",
"type": "boolean"
}
},
"required": [
"pdf",
"fields"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢file_to_markdown(file, format)
Convert a Word (.docx), Excel (.xlsx) or CSV file (base64) into clean Markdown — for RAG, agents and pipelines. DOCX keeps headings/lists/tables; spreadsheets become Markdown tables (one per sheet). Returns Markdown inline. Optional 'format' (docx|xlsx|csv) overrides auto-detection.
Input Schema
{
"type": "object",
"properties": {
"file": {
"type": "string",
"description": "Base64-encoded .docx, .xlsx or .csv file."
},
"format": {
"description": "Optional format hint; auto-detected if omitted.",
"type": "string",
"enum": [
"docx",
"xlsx",
"csv"
]
}
},
"required": [
"file"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢data_to_spreadsheet(rows, format, sheet_name)
Turn rows of data into a downloadable XLSX (default) or CSV. 'rows' is an array of objects (keys → header row) or an array of arrays (first row is the header). Returns a file_id (for chaining) and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"rows": {
"minItems": 1,
"type": "array",
"items": {},
"description": "Array of objects (keys→columns) or array of arrays (first row = header)."
},
"format": {
"description": "Output format (default xlsx).",
"type": "string",
"enum": [
"xlsx",
"csv"
]
},
"sheet_name": {
"description": "Worksheet name (default 'Sheet1').",
"type": "string"
}
},
"required": [
"rows"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢batch_scrape(urls)
Fetch up to 10 public URLs and return each as clean Markdown, in one call — for research/RAG over several pages at once. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"urls": {
"minItems": 1,
"maxItems": 10,
"type": "array",
"items": {
"type": "string"
},
"description": "1–10 http/https URLs."
}
},
"required": [
"urls"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢crawl_site(url, limit)
Fetch a URL plus up to 7 more same-origin pages it links to (≤8 total), each as clean Markdown. Bounded and synchronous. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The root http/https URL."
},
"limit": {
"description": "Max pages incl. root (default 5, max 8).",
"type": "integer",
"minimum": 1,
"maximum": 8
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢map_site(url)
Discover the same-origin URLs linked from a page — the site map, with no page content fetched. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The http/https URL to map."
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢extract_data(fields, url, pdf, file)
Extract specified fields from a web page, PDF, or Office file as JSON, using AI. Provide 'fields' plus ONE source: a 'url', a 'pdf' (file_id or base64), or a 'file' (base64 .docx/.xlsx/.csv). Returns a JSON object mapping each field to its value, or null when absent.
Input Schema
{
"type": "object",
"properties": {
"fields": {
"type": "array",
"items": {
"type": "string"
},
"description": "Field names to extract, e.g. ['invoice total','due date']."
},
"url": {
"description": "A public http/https URL to read.",
"type": "string"
},
"pdf": {
"description": "A file_id from a prior tool result, or a base64-encoded PDF.",
"type": "string"
},
"file": {
"description": "A base64-encoded .docx, .xlsx or .csv file (converted to text first).",
"type": "string"
}
},
"required": [
"fields"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢summarize_document(url, pdf, file, text, max_words)
Summarize a web page, PDF, Office file (.docx/.xlsx/.csv), or raw text using AI. Provide ONE source: 'url', 'pdf' (file_id or base64), 'file' (base64), or 'text'. Optional 'max_words'. Returns the summary.
Input Schema
{
"type": "object",
"properties": {
"url": {
"description": "A public http/https URL to read.",
"type": "string"
},
"pdf": {
"description": "A file_id from a prior tool, or a base64-encoded PDF.",
"type": "string"
},
"file": {
"description": "A base64-encoded .docx, .xlsx or .csv file.",
"type": "string"
},
"text": {
"description": "Raw text to summarize directly.",
"type": "string"
},
"max_words": {
"description": "Approximate target length in words.",
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢answer_from_document(question, url, pdf, file, text)
Answer a question grounded ONLY in a source document — a web page, PDF, Office file, or raw text. Provide 'question' plus ONE source: 'url', 'pdf', 'file' (base64), or 'text'. Says so when the answer isn't in the document.
Input Schema
{
"type": "object",
"properties": {
"question": {
"type": "string",
"description": "The question to answer from the document."
},
"url": {
"description": "A public http/https URL to read.",
"type": "string"
},
"pdf": {
"description": "A file_id from a prior tool, or a base64-encoded PDF.",
"type": "string"
},
"file": {
"description": "A base64-encoded .docx, .xlsx or .csv file.",
"type": "string"
},
"text": {
"description": "Raw text to answer from directly.",
"type": "string"
}
},
"required": [
"question"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢docx_to_text(docx)
Extract the text of a .docx document (base64). Returns the text inline.
Input Schema
{
"type": "object",
"properties": {
"docx": {
"type": "string",
"description": "Base64-encoded .docx file."
}
},
"required": [
"docx"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}⚪docx_to_pdf(docx)
Convert a .docx document (base64) to a PDF. Returns a file_id and a ~1h URL.
Input Schema
{
"type": "object",
"properties": {
"docx": {
"type": "string",
"description": "Base64-encoded .docx file."
}
},
"required": [
"docx"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢read_url(url)
Fetch a public http/https URL and return its main content as clean Markdown — ideal for giving an agent readable web content for research or RAG. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The http/https URL to read."
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢url_to_pdf(url)
Fetch a public http/https URL and render the live page to a PDF. Returns a file_id (chainable into the PDF tools) and a ~1h download URL. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The http/https URL to snapshot."
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟡url_to_screenshot(url, full_page)
Fetch a public http/https URL and capture a PNG screenshot. Set full_page for the entire scroll height. Returns a file_id and a ~1h download URL.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The http/https URL to capture."
},
"full_page": {
"description": "Capture the full scroll height (default: viewport).",
"type": "boolean"
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢scrape_page(url, formats)
Fetch a public http/https URL and return its rendered HTML, the links on the page, and page metadata (title/description/OpenGraph/canonical/favicon) — in one render. For clean Markdown use read_url instead. Private/internal hosts are blocked.
Input Schema
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The http/https URL to scrape."
},
"formats": {
"description": "Which representations to return (default: all three).",
"type": "array",
"items": {
"type": "string",
"enum": [
"html",
"links",
"metadata"
]
}
}
},
"required": [
"url"
],
"$schema": "http://json-schema.org/draft-07/schema#"
}Community
Evidence