scrapewright
Give it a URL, get structured rows. A model writes the parser once; replays are free.
¿Debería usar esto?
Calidad y seguridad
Hallazgos (1)
- LOWen account
Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.
Costo de contexto
Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.
Instalar
Instalación con un clic
Agrega esto a tu archivo `claude_desktop_config.json`:
{
"mcpServers": {
"scrapewright": {
"url": "https://scrapewright.app/mcp"
}
}
}Puntos de conexión remotos
https://scrapewright.app/mcpstreamable-httpQué puede hacer
Inventario de herramientas
Herramientas (5)
⚪detect_site(url)
Report what platform a site runs on and which strategy to use. Cheap; call it before a large job.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"title": "Url",
"type": "string"
}
},
"required": [
"url"
],
"title": "detect_siteArguments"
}Esquema de salida
{
"type": "object",
"additionalProperties": true,
"title": "detect_siteDictOutput"
}🟢extract_page(url, fields, js)
Extract structured data from ONE page. ``fields`` declares your own schema, e.g. ["title", "salary:number", "tags:list"]; omit it for the product schema. First call on a new site compiles a recipe (300 credits); later calls replay it for 1 credit per row.
Esquema de entrada
{
"type": "object",
"properties": {
"url": {
"title": "Url",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
},
"js": {
"default": false,
"title": "Js",
"type": "boolean"
}
},
"required": [
"url"
],
"title": "extract_pageArguments"
}Esquema de salida
{
"type": "object",
"additionalProperties": true,
"title": "extract_pageDictOutput"
}🟢crawl_site(listing_url, fields, max_items, js, scroll)
Walk a site from one listing URL and extract every item. Waits up to four minutes; a longer crawl returns a job_id to pass to crawl_status.
Esquema de entrada
{
"type": "object",
"properties": {
"listing_url": {
"title": "Listing Url",
"type": "string"
},
"fields": {
"anyOf": [
{
"items": {
"type": "string"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"title": "Fields"
},
"max_items": {
"default": 25,
"title": "Max Items",
"type": "integer"
},
"js": {
"default": false,
"title": "Js",
"type": "boolean"
},
"scroll": {
"default": 0,
"title": "Scroll",
"type": "integer"
}
},
"required": [
"listing_url"
],
"title": "crawl_siteArguments"
}Esquema de salida
{
"type": "object",
"additionalProperties": true,
"title": "crawl_siteDictOutput"
}🟢crawl_status(job_id)
Fetch a crawl that outlived its call.
Esquema de entrada
{
"type": "object",
"properties": {
"job_id": {
"title": "Job Id",
"type": "string"
}
},
"required": [
"job_id"
],
"title": "crawl_statusArguments"
}Esquema de salida
{
"type": "object",
"additionalProperties": true,
"title": "crawl_statusDictOutput"
}⚪account
Credits left and this month's usage for the key in use.
Esquema de entrada
{
"type": "object",
"properties": {},
"title": "accountArguments"
}Esquema de salida
{
"type": "object",
"additionalProperties": true,
"title": "accountDictOutput"
}Comunidad
Evidencia