Boolsai Grep

Parallel regex across all Boolsai scans — discover new vendor patterns, niche signals. 4 tools.

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
76%
Calidad de los nombres
90%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~759Tokens (definiciones de herramientas)
~1.5 KBTamaño de respuesta típico
Impacto moderado en la atención (0.59% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "grep": {
      "url": "https://grep.boolsai.ai/mcp"
    }
  }
}

Puntos de conexión remotos

https://grep.boolsai.ai/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (4)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟡grep_pattern(pattern, limit, max_scan, tld, since, ...)

Run a JavaScript regex across every scan we've ever taken. Returns matching site URLs (and optional snippets). Best for ad-hoc discovery of patterns NOT already extracted into id_index. Pass narrowing filters (vendor, tld, since) to keep scan volume manageable. Default returns URLs only — set with_snippets=true if you also want the matched JSON context.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "pattern": {
      "type": "string",
      "description": "JavaScript regex (case-insensitive by default)"
    },
    "limit": {
      "type": "integer",
      "description": "Max matching sites to return (default 200, max 2000)",
      "default": 200
    },
    "max_scan": {
      "type": "integer",
      "description": "Max R2 objects to read for this query (default 10000, max 200000)",
      "default": 10000
    },
    "tld": {
      "type": "string",
      "description": "Restrict to this TLD, e.g. 'com.au' or 'co.uk'"
    },
    "since": {
      "type": "string",
      "description": "Only scans >= this date, format YYYY-MM-DD"
    },
    "vendor": {
      "type": "string",
      "description": "Pre-narrow via id_index to sites known to use this vendor (e.g. 'shopify')"
    },
    "domain_like": {
      "type": "string",
      "description": "Substring to match in domain, e.g. 'patagonia'"
    },
    "case_sensitive": {
      "type": "boolean",
      "description": "Default false",
      "default": false
    },
    "with_snippets": {
      "type": "boolean",
      "description": "Include matched JSON context. Default false (URLs only).",
      "default": false
    },
    "max_shards": {
      "type": "integer",
      "description": "Cap how many shards to query (newest-first). Default 100, raise to scan deeper history. Each shard ≈ 800 scans.",
      "default": 100
    },
    "all": {
      "type": "boolean",
      "description": "Scan ALL shards — entire corpus. May exceed Worker CPU budget for complex regex; use only when you've narrowed via vendor/tld.",
      "default": false
    }
  },
  "required": [
    "pattern"
  ]
}
🟢count_pattern(pattern, max_scan, tld, since, vendor, ...)

Same as grep_pattern but returns only the count of matching sites. Faster — no per-match serialization. Useful for sizing 'how many sites use X' before doing a full grep.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "pattern": {
      "type": "string"
    },
    "max_scan": {
      "type": "integer",
      "default": 10000
    },
    "tld": {
      "type": "string"
    },
    "since": {
      "type": "string"
    },
    "vendor": {
      "type": "string"
    },
    "domain_like": {
      "type": "string"
    },
    "case_sensitive": {
      "type": "boolean"
    }
  },
  "required": [
    "pattern"
  ]
}
🟢sites_with_signal(signal_type, signal_value, limit)

Return sites that have a known signal_type=signal_value pair in id_index. MUCH faster than grep_pattern — uses pre-indexed D1 lookup. Use this whenever the signal_type is in the list (see instructions). Examples: sites_with_signal(signal_type='vendor', signal_value='mparticle_workspace') ← won't work, that's an ID type, use sites_with_signal(signal_type='mparticle_workspace') WITHOUT signal_value to list all values + their domains.

Esquema de entrada

{
  "type": "object",
  "properties": {
    "signal_type": {
      "type": "string",
      "description": "One of the known signal_type values (see instructions)"
    },
    "signal_value": {
      "type": "string",
      "description": "Optional exact value to filter to. If omitted, returns all domains for any value of this signal_type."
    },
    "limit": {
      "type": "integer",
      "default": 200
    }
  },
  "required": [
    "signal_type"
  ]
}
🟢list_signal_types

Returns the full catalog of signal_types currently present in id_index, with counts of unique values and unique domains for each.

Esquema de entrada

{
  "type": "object",
  "properties": {}
}

Comparar

Comparación con similares

Compara este servidor con otros de la misma categoría.

Herramientas:0
Herramientas:0
Herramientas:0
Herramientas:0
Herramientas:0
Disponibilidad:100.0%
Latencia:277ms

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada4 herramientas
verificadoversión no registrada4 herramientas