Boolsai Grep
Parallel regex across all Boolsai scans — discover new vendor patterns, niche signals. 4 tools.
사용해야 할까요
품질 및 안전성
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"grep": {
"url": "https://grep.boolsai.ai/mcp"
}
}
}원격 엔드포인트
https://grep.boolsai.ai/mcpstreamable-http할 수 있는 일
도구 목록
도구 (4)
🟡grep_pattern(pattern, limit, max_scan, tld, since, ...)
Run a JavaScript regex across every scan we've ever taken. Returns matching site URLs (and optional snippets). Best for ad-hoc discovery of patterns NOT already extracted into id_index. Pass narrowing filters (vendor, tld, since) to keep scan volume manageable. Default returns URLs only — set with_snippets=true if you also want the matched JSON context.
입력 스키마
{
"type": "object",
"properties": {
"pattern": {
"type": "string",
"description": "JavaScript regex (case-insensitive by default)"
},
"limit": {
"type": "integer",
"description": "Max matching sites to return (default 200, max 2000)",
"default": 200
},
"max_scan": {
"type": "integer",
"description": "Max R2 objects to read for this query (default 10000, max 200000)",
"default": 10000
},
"tld": {
"type": "string",
"description": "Restrict to this TLD, e.g. 'com.au' or 'co.uk'"
},
"since": {
"type": "string",
"description": "Only scans >= this date, format YYYY-MM-DD"
},
"vendor": {
"type": "string",
"description": "Pre-narrow via id_index to sites known to use this vendor (e.g. 'shopify')"
},
"domain_like": {
"type": "string",
"description": "Substring to match in domain, e.g. 'patagonia'"
},
"case_sensitive": {
"type": "boolean",
"description": "Default false",
"default": false
},
"with_snippets": {
"type": "boolean",
"description": "Include matched JSON context. Default false (URLs only).",
"default": false
},
"max_shards": {
"type": "integer",
"description": "Cap how many shards to query (newest-first). Default 100, raise to scan deeper history. Each shard ≈ 800 scans.",
"default": 100
},
"all": {
"type": "boolean",
"description": "Scan ALL shards — entire corpus. May exceed Worker CPU budget for complex regex; use only when you've narrowed via vendor/tld.",
"default": false
}
},
"required": [
"pattern"
]
}🟢count_pattern(pattern, max_scan, tld, since, vendor, ...)
Same as grep_pattern but returns only the count of matching sites. Faster — no per-match serialization. Useful for sizing 'how many sites use X' before doing a full grep.
입력 스키마
{
"type": "object",
"properties": {
"pattern": {
"type": "string"
},
"max_scan": {
"type": "integer",
"default": 10000
},
"tld": {
"type": "string"
},
"since": {
"type": "string"
},
"vendor": {
"type": "string"
},
"domain_like": {
"type": "string"
},
"case_sensitive": {
"type": "boolean"
}
},
"required": [
"pattern"
]
}🟢sites_with_signal(signal_type, signal_value, limit)
Return sites that have a known signal_type=signal_value pair in id_index. MUCH faster than grep_pattern — uses pre-indexed D1 lookup. Use this whenever the signal_type is in the list (see instructions). Examples: sites_with_signal(signal_type='vendor', signal_value='mparticle_workspace') ← won't work, that's an ID type, use sites_with_signal(signal_type='mparticle_workspace') WITHOUT signal_value to list all values + their domains.
입력 스키마
{
"type": "object",
"properties": {
"signal_type": {
"type": "string",
"description": "One of the known signal_type values (see instructions)"
},
"signal_value": {
"type": "string",
"description": "Optional exact value to filter to. If omitted, returns all domains for any value of this signal_type."
},
"limit": {
"type": "integer",
"default": 200
}
},
"required": [
"signal_type"
]
}🟢list_signal_types
Returns the full catalog of signal_types currently present in id_index, with counts of unique values and unique domains for each.
입력 스키마
{
"type": "object",
"properties": {}
}비교
동종 서버 비교
같은 카테고리의 다른 서버와 이 서버를 비교하세요.
커뮤니티
증거