ScrapingBot
Public web data for AI agents: scrape any page, plus Google, TikTok, Instagram and Amazon as JSON.
¿Debería usar esto?
Calidad y seguridad
Hallazgos (4)
- LOWen getScrapeJob
- LOWen tiktokComments
- LOWen amazonProduct
- LOWen amazonRankings
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"scrapingbot": {
"url": "https://scrapingbot.io/api/mcp"
}
}
}Puntos de conexión remotos
https://scrapingbot.io/api/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (23)
🟢listCapabilities
List every ScrapingBot tool, the API routes behind them and their parameters. Free.
Esquema de entrada
{
"type": "object",
"properties": {}
}🟢scrapeWebsite(url, render_js, stealth_proxy, premium_proxy, wait, ...)
Fetch any public web page and return its content as markdown (format: "html" for the raw HTML). 1 credit; render_js (pages built with JavaScript) 5; premium_proxy 10; stealth_proxy 75. screenshot: true attaches a screenshot image.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri"
},
"render_js": {
"type": "boolean"
},
"stealth_proxy": {
"type": "boolean"
},
"premium_proxy": {
"type": "boolean"
},
"wait": {
"type": "integer",
"minimum": 0,
"maximum": 45000
},
"wait_for": {
"type": "string",
"minLength": 1
},
"wait_browser": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"cookies": {
"type": "string"
},
"screenshot": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "string",
"const": "fullPage"
}
]
},
"block_ads": {
"type": "boolean"
},
"block_resources": {
"type": "boolean"
},
"capture_runtime_issues": {
"type": "boolean"
},
"timeout": {
"type": "integer",
"minimum": 1,
"maximum": 120000
},
"format": {
"type": "string",
"enum": [
"markdown",
"html"
],
"description": "Page content as markdown (default, compact and readable) or the raw HTML"
}
},
"required": [
"url"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢extractStructuredData(url, render_js, stealth_proxy, premium_proxy, wait, ...)
Fetch a web page and have AI pull out exactly the fields you want, as JSON. Pass ai_query (plain English, e.g. "title, price and stock") or ai_schema ({"field": "description"}). The page cost plus 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri"
},
"render_js": {
"type": "boolean"
},
"stealth_proxy": {
"type": "boolean"
},
"premium_proxy": {
"type": "boolean"
},
"wait": {
"type": "integer",
"minimum": 0,
"maximum": 45000
},
"wait_for": {
"type": "string",
"minLength": 1
},
"wait_browser": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"cookies": {
"type": "string"
},
"screenshot": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "string",
"const": "fullPage"
}
]
},
"block_ads": {
"type": "boolean"
},
"block_resources": {
"type": "boolean"
},
"capture_runtime_issues": {
"type": "boolean"
},
"timeout": {
"type": "integer",
"minimum": 1,
"maximum": 120000
},
"ai_query": {
"type": "string",
"minLength": 1
},
"ai_schema": {
"anyOf": [
{
"type": "string",
"minLength": 1
},
{
"type": "object",
"additionalProperties": {}
}
]
}
},
"required": [
"url"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢getScrapeJob(job_id)
Look up an earlier scrape by the job_id it returned. Free.
Esquema de entrada
{
"type": "object",
"properties": {
"job_id": {
"type": "string",
"minLength": 1
}
},
"required": [
"job_id"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢pollJobUntilDone(job_id, timeout_ms, poll_interval_ms)
Wait for a scrape job to finish, then return it. Free.
Esquema de entrada
{
"type": "object",
"properties": {
"job_id": {
"type": "string",
"minLength": 1
},
"timeout_ms": {
"type": "integer",
"minimum": 1000,
"maximum": 300000
},
"poll_interval_ms": {
"type": "integer",
"minimum": 250,
"maximum": 30000
}
},
"required": [
"job_id"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢runBrowserScenario(url, render_js, stealth_proxy, premium_proxy, wait, ...)
Open a page in a real browser, run steps such as click, fill, scroll and wait, then return the resulting page as markdown. js_scenario: {"instructions": [{"fill": ["#q", "shoes"]}, {"click": "#go"}, {"wait": 1000}]}. From 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri"
},
"render_js": {
"type": "boolean"
},
"stealth_proxy": {
"type": "boolean"
},
"premium_proxy": {
"type": "boolean"
},
"wait": {
"type": "integer",
"minimum": 0,
"maximum": 45000
},
"wait_for": {
"type": "string",
"minLength": 1
},
"wait_browser": {
"type": "string",
"enum": [
"load",
"domcontentloaded",
"networkidle0",
"networkidle2"
]
},
"cookies": {
"type": "string"
},
"screenshot": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "string",
"const": "fullPage"
}
]
},
"block_ads": {
"type": "boolean"
},
"block_resources": {
"type": "boolean"
},
"capture_runtime_issues": {
"type": "boolean"
},
"timeout": {
"type": "integer",
"minimum": 1,
"maximum": 120000
},
"format": {
"type": "string",
"enum": [
"markdown",
"html"
],
"description": "Page content as markdown (default, compact and readable) or the raw HTML"
},
"js_scenario": {
"anyOf": [
{
"type": "array",
"items": {
"type": "object",
"additionalProperties": {}
}
},
{
"$ref": "#/properties/js_scenario/anyOf/0/items"
}
]
}
},
"required": [
"url",
"js_scenario"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢googleSearch(q, type, gl, hl, autocorrect, ...)
Google results as JSON: web (default), images, videos, news, shopping, places or maps (type), for any country (gl) and language (hl). Places and maps results include a cid for googleReviews. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"q": {
"type": "string",
"minLength": 1
},
"type": {
"type": "string",
"enum": [
"search",
"images",
"videos",
"news",
"shopping",
"places",
"maps"
],
"description": "What to search: web results (default), images, videos, news, shopping, places or maps."
},
"gl": {
"type": "string",
"minLength": 2,
"maxLength": 8
},
"hl": {
"type": "string",
"minLength": 2,
"maxLength": 8
},
"autocorrect": {
"type": "boolean"
},
"num": {
"anyOf": [
{
"type": "integer",
"minimum": 1,
"maximum": 100
},
{
"type": "string",
"minLength": 1
}
]
},
"page": {
"anyOf": [
{
"type": "integer",
"minimum": 1
},
{
"type": "string",
"minLength": 1
}
]
},
"tbs": {
"type": "string",
"minLength": 1
},
"ll": {
"type": "string",
"minLength": 1,
"description": "Maps/places viewport, e.g. @30.27,-97.74,12z"
},
"sortBy": {
"type": "string",
"minLength": 1
}
},
"required": [
"q"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢googleReviews(cid, fid, placeId, sortBy, nextPageToken, ...)
Reviews of a place from Google Maps, by the cid, fid or placeId in a googleSearch places or maps result. sortBy newest, mostRelevant, highestRating or lowestRating; pass nextPageToken for more. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"cid": {
"type": "string",
"minLength": 1
},
"fid": {
"type": "string",
"minLength": 1
},
"placeId": {
"type": "string",
"minLength": 1
},
"sortBy": {
"type": "string",
"enum": [
"mostRelevant",
"newest",
"highestRating",
"lowestRating"
]
},
"nextPageToken": {
"type": "string",
"minLength": 1
},
"gl": {
"type": "string",
"minLength": 2,
"maxLength": 8
},
"hl": {
"type": "string",
"minLength": 2,
"maxLength": 8
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢instagramUser(username, user_id, url)
An Instagram profile: bio, follower, following and post counts, links and the numeric user id other tools need. Pass one of username, user_id or url. 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"username": {
"type": "string",
"minLength": 1
},
"user_id": {
"anyOf": [
{
"type": "string",
"minLength": 1
},
{
"type": "integer",
"exclusiveMinimum": 0
}
]
},
"url": {
"type": "string",
"format": "uri"
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢instagramSearch(keyword, search_type)
Search Instagram accounts, hashtags, places or posts (search_type), or everything at once (global). 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"keyword": {
"type": "string",
"minLength": 1
},
"search_type": {
"type": "string",
"enum": [
"users",
"hashtags",
"places",
"global",
"posts"
]
}
},
"required": [
"keyword",
"search_type"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢instagramMedia(mode, user_id, shortcode, url, code_or_id_or_url, ...)
Instagram content by mode: an account's posts, reels or tagged posts (user_id), one post by shortcode or url, a post's comments, the replies to a comment (comment_id), or an account's active stories (username). 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"mode": {
"type": "string",
"enum": [
"user_posts",
"tagged_posts",
"reels",
"media_by_shortcode",
"media_by_url",
"comments",
"comment_replies",
"stories"
]
},
"user_id": {
"anyOf": [
{
"type": "string",
"minLength": 1
},
{
"type": "integer",
"exclusiveMinimum": 0
}
]
},
"shortcode": {
"type": "string",
"minLength": 1
},
"url": {
"type": "string",
"format": "uri"
},
"code_or_id_or_url": {
"type": "string",
"minLength": 1
},
"media_id": {
"anyOf": [
{
"type": "string",
"minLength": 1
},
{
"type": "integer",
"exclusiveMinimum": 0
}
]
},
"comment_id": {
"type": "string",
"minLength": 1,
"description": "Parent comment id (comment_replies)"
},
"username": {
"type": "string",
"minLength": 1,
"description": "Username (stories)"
},
"count": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 100
},
"end_cursor": {
"type": "string",
"minLength": 1
},
"pagination_token": {
"type": "string",
"minLength": 1
}
},
"required": [
"mode"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢instagramFollowers(user_id, direction, query, max_id)
List or search an Instagram account's followers or following, by numeric user_id from instagramUser. query searches the whole list; max_id pages. For very large accounts Instagram shows only about 50 recent followers, so use query to find someone. 5 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"user_id": {
"type": "string",
"pattern": "^\\d+$",
"description": "Numeric user id (from instagramUser)"
},
"direction": {
"type": "string",
"enum": [
"followers",
"following"
]
},
"query": {
"type": "string",
"minLength": 1,
"description": "Search the list by username or name"
},
"max_id": {
"type": "string",
"minLength": 1,
"description": "next_max_id from the previous page"
}
},
"required": [
"user_id",
"direction"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢tiktokVideo(url)
A TikTok video's stats, caption, author, sound and download links, by video URL. 1 credit.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri"
}
},
"required": [
"url"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢tiktokUser(unique_id, include_posts)
A TikTok profile by username: bio, counts, verified badge and numeric user id; include_posts returns the latest videos instead. 1 credit.
Esquema de entrada
{
"type": "object",
"properties": {
"unique_id": {
"type": "string",
"minLength": 1
},
"include_posts": {
"type": "boolean"
}
},
"required": [
"unique_id"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢tiktokSearch(query, search_type, count, cursor, search_id)
Search TikTok creators (search_type: users) or videos (videos) by keyword. Videos page: for page 2 onward pass the previous response's data.cursor as cursor and its data.searchId as search_id. Users search returns a single page and ignores cursor. 1 credit.
Esquema de entrada
{
"type": "object",
"properties": {
"query": {
"type": "string",
"minLength": 1
},
"search_type": {
"type": "string",
"enum": [
"users",
"videos"
]
},
"count": {
"type": "integer",
"exclusiveMinimum": 0,
"maximum": 100
},
"cursor": {
"anyOf": [
{
"type": "integer",
"minimum": 0
},
{
"type": "string",
"minLength": 1
}
],
"description": "Videos only: the previous response's data.cursor, for the next page"
},
"search_id": {
"type": "string",
"minLength": 1,
"description": "Videos only: the previous response's data.searchId; required with cursor for page 2 onward"
}
},
"required": [
"query",
"search_type"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢tiktokComments(url, video_id, comment_id, count, cursor)
Comments on a TikTok video (url), or the replies to one comment (comment_id plus the video url or video_id). 1 credit.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"type": "string",
"format": "uri",
"description": "Video URL (required for comments)"
},
"video_id": {
"type": "string",
"minLength": 1,
"description": "Video id (replies only; or pass url)"
},
"comment_id": {
"type": "string",
"minLength": 1,
"description": "Set to fetch the replies to this comment"
},
"count": {
"type": "integer",
"minimum": 1,
"maximum": 50
},
"cursor": {
"type": "integer",
"minimum": 0
}
},
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢tiktokFollowers(user_id, direction, count)
A TikTok account's followers or the accounts it follows, by numeric user_id (data.user.id from tiktokUser). Up to 50. 1 credit.
Esquema de entrada
{
"type": "object",
"properties": {
"user_id": {
"type": "string",
"pattern": "^\\d+$",
"description": "Numeric user id (data.user.id from tiktokUser)"
},
"direction": {
"type": "string",
"enum": [
"followers",
"following"
]
},
"count": {
"type": "integer",
"minimum": 1,
"maximum": 50
}
},
"required": [
"user_id",
"direction"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢amazonSearch(query, page, department)
Search Amazon US products by keyword: title, price, rating, Prime and photo for each. page for more. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"query": {
"type": "string",
"minLength": 1
},
"page": {
"anyOf": [
{
"type": "integer",
"minimum": 1,
"maximum": 100
},
{
"type": "string",
"minLength": 1
}
]
},
"department": {
"type": "string",
"minLength": 1
}
},
"required": [
"query"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢amazonProduct(asin)
Everything on an Amazon product page by ASIN: price, rating, availability, bullet points, specifications, photos and variations. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"asin": {
"type": "string",
"minLength": 1
}
},
"required": [
"asin"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢amazonProducts(asins)
Full details for up to 20 ASINs in one call, the same as amazonProduct for each. 10 credits per product returned; ASINs that fail are free.
Esquema de entrada
{
"type": "object",
"properties": {
"asins": {
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"minItems": 1,
"maxItems": 20,
"description": "Up to 20 ASINs; 10 credits per product returned"
}
},
"required": [
"asins"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢amazonSuggestions(query, department, limit)
The suggestions Amazon shows as a shopper types: useful for keyword research. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"query": {
"type": "string",
"minLength": 1,
"description": "What a shopper has typed so far"
},
"department": {
"type": "string",
"minLength": 1,
"description": "Search alias such as aps, electronics, stripbooks"
},
"limit": {
"type": "integer",
"minimum": 1,
"maximum": 20
}
},
"required": [
"query"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢amazonRankings(list, category, page)
Amazon's Best Sellers or New Releases for a department, in rank order, with price and rating. 10 credits.
Esquema de entrada
{
"type": "object",
"properties": {
"list": {
"type": "string",
"enum": [
"best_sellers",
"new_releases"
]
},
"category": {
"type": "string",
"enum": [
"electronics",
"pc",
"wireless",
"videogames",
"books",
"toys-and-games",
"home-garden",
"kitchen",
"appliances",
"beauty",
"hpc",
"grocery",
"fashion",
"sporting-goods",
"pet-supplies",
"baby-products",
"automotive",
"office-products",
"arts-crafts",
"musical-instruments"
],
"description": "Amazon department slug"
},
"page": {
"type": "integer",
"minimum": 1,
"maximum": 100
}
},
"required": [
"list",
"category"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}🟢providerRequest(provider, endpoint, params)
Call any Google, Instagram, TikTok or Amazon endpoint by path, for the few the other tools do not cover (listCapabilities lists them). Costs the same as that API.
Esquema de entrada
{
"type": "object",
"properties": {
"provider": {
"type": "string",
"enum": [
"google",
"instagram",
"tiktok",
"amazon"
]
},
"endpoint": {
"type": "string",
"minLength": 1
},
"params": {
"type": "object",
"additionalProperties": {}
}
},
"required": [
"provider",
"endpoint"
],
"additionalProperties": false,
"$schema": "http://json-schema.org/draft-07/schema#"
}Comunidad
Evidencia