scrapewright

Give it a URL, get structured rows. A model writes the parser once; replays are free.

我该使用它吗

质量与安全性

B
描述质量
90%
模式完整度
68%
命名质量
80%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

发现(1)

  • LOWTool 'account' description lacks action verb在 account 中

基于对工具定义和协议合规性的自动分析。

上下文开销

~567token 数(工具定义)
~844 B典型响应大小
对注意力的影响极小(占 128k 上下文窗口的 0.44%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "scrapewright": {
      "url": "https://scrapewright.app/mcp"
    }
  }
}

远程端点

https://scrapewright.app/mcpstreamable-http

它能做什么

工具清单

工具(5)

🟢 只读🟡 写入🔴 删除⚪ 未知
⚪detect_site(url)

Report what platform a site runs on and which strategy to use. Cheap; call it before a large job.

输入模式

{
  "type": "object",
  "properties": {
    "url": {
      "title": "Url",
      "type": "string"
    }
  },
  "required": [
    "url"
  ],
  "title": "detect_siteArguments"
}

输出模式

{
  "type": "object",
  "additionalProperties": true,
  "title": "detect_siteDictOutput"
}
🟢extract_page(url, fields, js)

Extract structured data from ONE page. ``fields`` declares your own schema, e.g. ["title", "salary:number", "tags:list"]; omit it for the product schema. First call on a new site compiles a recipe (300 credits); later calls replay it for 1 credit per row.

输入模式

{
  "type": "object",
  "properties": {
    "url": {
      "title": "Url",
      "type": "string"
    },
    "fields": {
      "anyOf": [
        {
          "items": {
            "type": "string"
          },
          "type": "array"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Fields"
    },
    "js": {
      "default": false,
      "title": "Js",
      "type": "boolean"
    }
  },
  "required": [
    "url"
  ],
  "title": "extract_pageArguments"
}

输出模式

{
  "type": "object",
  "additionalProperties": true,
  "title": "extract_pageDictOutput"
}
🟢crawl_site(listing_url, fields, max_items, js, scroll)

Walk a site from one listing URL and extract every item. Waits up to four minutes; a longer crawl returns a job_id to pass to crawl_status.

输入模式

{
  "type": "object",
  "properties": {
    "listing_url": {
      "title": "Listing Url",
      "type": "string"
    },
    "fields": {
      "anyOf": [
        {
          "items": {
            "type": "string"
          },
          "type": "array"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "title": "Fields"
    },
    "max_items": {
      "default": 25,
      "title": "Max Items",
      "type": "integer"
    },
    "js": {
      "default": false,
      "title": "Js",
      "type": "boolean"
    },
    "scroll": {
      "default": 0,
      "title": "Scroll",
      "type": "integer"
    }
  },
  "required": [
    "listing_url"
  ],
  "title": "crawl_siteArguments"
}

输出模式

{
  "type": "object",
  "additionalProperties": true,
  "title": "crawl_siteDictOutput"
}
🟢crawl_status(job_id)

Fetch a crawl that outlived its call.

输入模式

{
  "type": "object",
  "properties": {
    "job_id": {
      "title": "Job Id",
      "type": "string"
    }
  },
  "required": [
    "job_id"
  ],
  "title": "crawl_statusArguments"
}

输出模式

{
  "type": "object",
  "additionalProperties": true,
  "title": "crawl_statusDictOutput"
}
⚪account

Credits left and this month's usage for the key in use.

输入模式

{
  "type": "object",
  "properties": {},
  "title": "accountArguments"
}

输出模式

{
  "type": "object",
  "additionalProperties": true,
  "title": "accountDictOutput"
}

社区

评价此服务器

证据

最近观测

已验证未记录版本5 个工具
已验证未记录版本5 个工具