Boolsai Grep

Parallel regex across all Boolsai scans — discover new vendor patterns, niche signals. 4 tools.

我该使用它吗

质量与安全性

A
描述质量
100%
模式完整度
76%
命名质量
90%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~759token 数(工具定义)
~1.5 KB典型响应大小
对注意力有中等影响(占 128k 上下文窗口的 0.59%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "grep": {
      "url": "https://grep.boolsai.ai/mcp"
    }
  }
}

远程端点

https://grep.boolsai.ai/mcpstreamable-http

它能做什么

工具清单

工具(4)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟡grep_pattern(pattern, limit, max_scan, tld, since, ...)

Run a JavaScript regex across every scan we've ever taken. Returns matching site URLs (and optional snippets). Best for ad-hoc discovery of patterns NOT already extracted into id_index. Pass narrowing filters (vendor, tld, since) to keep scan volume manageable. Default returns URLs only — set with_snippets=true if you also want the matched JSON context.

输入模式

{
  "type": "object",
  "properties": {
    "pattern": {
      "type": "string",
      "description": "JavaScript regex (case-insensitive by default)"
    },
    "limit": {
      "type": "integer",
      "description": "Max matching sites to return (default 200, max 2000)",
      "default": 200
    },
    "max_scan": {
      "type": "integer",
      "description": "Max R2 objects to read for this query (default 10000, max 200000)",
      "default": 10000
    },
    "tld": {
      "type": "string",
      "description": "Restrict to this TLD, e.g. 'com.au' or 'co.uk'"
    },
    "since": {
      "type": "string",
      "description": "Only scans >= this date, format YYYY-MM-DD"
    },
    "vendor": {
      "type": "string",
      "description": "Pre-narrow via id_index to sites known to use this vendor (e.g. 'shopify')"
    },
    "domain_like": {
      "type": "string",
      "description": "Substring to match in domain, e.g. 'patagonia'"
    },
    "case_sensitive": {
      "type": "boolean",
      "description": "Default false",
      "default": false
    },
    "with_snippets": {
      "type": "boolean",
      "description": "Include matched JSON context. Default false (URLs only).",
      "default": false
    },
    "max_shards": {
      "type": "integer",
      "description": "Cap how many shards to query (newest-first). Default 100, raise to scan deeper history. Each shard ≈ 800 scans.",
      "default": 100
    },
    "all": {
      "type": "boolean",
      "description": "Scan ALL shards — entire corpus. May exceed Worker CPU budget for complex regex; use only when you've narrowed via vendor/tld.",
      "default": false
    }
  },
  "required": [
    "pattern"
  ]
}
🟢count_pattern(pattern, max_scan, tld, since, vendor, ...)

Same as grep_pattern but returns only the count of matching sites. Faster — no per-match serialization. Useful for sizing 'how many sites use X' before doing a full grep.

输入模式

{
  "type": "object",
  "properties": {
    "pattern": {
      "type": "string"
    },
    "max_scan": {
      "type": "integer",
      "default": 10000
    },
    "tld": {
      "type": "string"
    },
    "since": {
      "type": "string"
    },
    "vendor": {
      "type": "string"
    },
    "domain_like": {
      "type": "string"
    },
    "case_sensitive": {
      "type": "boolean"
    }
  },
  "required": [
    "pattern"
  ]
}
🟢sites_with_signal(signal_type, signal_value, limit)

Return sites that have a known signal_type=signal_value pair in id_index. MUCH faster than grep_pattern — uses pre-indexed D1 lookup. Use this whenever the signal_type is in the list (see instructions). Examples: sites_with_signal(signal_type='vendor', signal_value='mparticle_workspace') ← won't work, that's an ID type, use sites_with_signal(signal_type='mparticle_workspace') WITHOUT signal_value to list all values + their domains.

输入模式

{
  "type": "object",
  "properties": {
    "signal_type": {
      "type": "string",
      "description": "One of the known signal_type values (see instructions)"
    },
    "signal_value": {
      "type": "string",
      "description": "Optional exact value to filter to. If omitted, returns all domains for any value of this signal_type."
    },
    "limit": {
      "type": "integer",
      "default": 200
    }
  },
  "required": [
    "signal_type"
  ]
}
🟢list_signal_types

Returns the full catalog of signal_types currently present in id_index, with counts of unique values and unique domains for each.

输入模式

{
  "type": "object",
  "properties": {}
}

对比

同类对比

将该服务器与同一类别中的其他服务器进行对比。

工具:0
工具:0
工具:0
工具:0
工具:0
可用率:100.0%
延迟:277ms

社区

评价此服务器

证据

最近观测

已验证未记录版本4 个工具
已验证未记录版本4 个工具