Siteiz

Check which AI crawlers a site allows and see pages the way AI crawlers do. Free, read-only.

Should I use this

Quality & Safety

A
Description quality
100%
Schema completeness
95%
Naming quality
90%
Poisoning risk
100%
Permission match
100%
Protocol compliance
100%

Based on automated analysis of tool definitions and protocol compliance.

Context Cost

~1,197Tokens (tool definitions)
~2.2 KBTypical response size
Moderate attention impact (0.94% of 128k context)

This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.

Install

One-Click Install

Add this to your `claude_desktop_config.json` file:

{
  "mcpServers": {
    "siteiz": {
      "url": "https://siteiz.com/mcp"
    }
  }
}

Remote endpoints

https://siteiz.com/mcpstreamable-http

What it can do

Tool inventory

Tools (4)

🟢 Read-only🟡 Write🔴 Delete⚪ Unknown
🟢check_ai_crawlers(url)

Reads a website's robots.txt and reports, for 11 AI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, CCBot, Bytespider), whether each may reach the given page. Separates AI search crawlers (decide if the site appears in ChatGPT, Claude and Perplexity answers) from training crawlers. Also reports llms.txt and declared sitemaps.

Input Schema

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "Website or page address, e.g. example.com or https://example.com/pricing"
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "url": {
      "type": [
        "string",
        "null"
      ]
    },
    "robotsTxt": {
      "type": "string",
      "enum": [
        "ok",
        "missing",
        "unreachable",
        "invalid",
        "pasted"
      ]
    },
    "httpStatus": {
      "type": [
        "integer",
        "null"
      ]
    },
    "agents": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "token": {
            "type": "string"
          },
          "verdict": {
            "type": "string",
            "enum": [
              "allowed",
              "default",
              "blocked"
            ]
          },
          "operator": {
            "type": "string"
          },
          "role": {
            "type": "string",
            "enum": [
              "search",
              "user",
              "training",
              "control"
            ]
          }
        },
        "required": [
          "token",
          "verdict",
          "operator",
          "role"
        ]
      }
    },
    "llmsTxt": {
      "type": [
        "boolean",
        "null"
      ]
    },
    "sitemaps": {
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "contentSignals": {
      "type": "array",
      "items": {
        "type": "string"
      }
    }
  },
  "required": [
    "url",
    "robotsTxt",
    "agents"
  ]
}
🟢view_page_as_ai_crawler(url)

Fetches one page the way most AI crawlers do (plain HTTP, no JavaScript) and reports what they actually receive: title, meta description, canonical, H1, heading count, structured data types, visible word count, token reading cost, and a text preview. Use it to check whether content is visible to AI without JavaScript.

Input Schema

{
  "type": "object",
  "properties": {
    "url": {
      "type": "string",
      "description": "Website or page address, e.g. example.com or https://example.com/pricing"
    }
  },
  "required": [
    "url"
  ],
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "finalUrl": {
      "type": "string"
    },
    "status": {
      "type": "integer"
    },
    "bytes": {
      "type": "integer"
    },
    "title": {
      "type": [
        "string",
        "null"
      ]
    },
    "metaDescription": {
      "type": [
        "string",
        "null"
      ]
    },
    "canonical": {
      "type": [
        "string",
        "null"
      ]
    },
    "h1": {
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "headingCount": {
      "type": "integer"
    },
    "structuredData": {
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "words": {
      "type": "integer"
    },
    "htmlTokens": {
      "type": "integer"
    },
    "textTokens": {
      "type": "integer"
    },
    "fullScan": {
      "type": "string"
    }
  },
  "required": [
    "finalUrl",
    "status",
    "words",
    "fullScan"
  ]
}
🟢explain_ai_crawler(name)

Explains what an AI crawler user agent does (operator, purpose, and what blocking it in robots.txt changes). Pass a name such as GPTBot or OAI-SearchBot, or omit it to list all crawlers Siteiz tracks.

Input Schema

{
  "type": "object",
  "properties": {
    "name": {
      "type": "string",
      "description": "Crawler user agent, e.g. GPTBot. Omit to list all."
    }
  },
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "crawlers": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "token": {
            "type": "string"
          },
          "operator": {
            "type": "string"
          },
          "product": {
            "type": "string"
          },
          "role": {
            "type": "string"
          },
          "roleLabel": {
            "type": "string"
          },
          "does": {
            "type": "string"
          }
        },
        "required": [
          "token",
          "operator",
          "role",
          "does"
        ]
      }
    }
  },
  "required": [
    "crawlers"
  ]
}
🟢get_ai_visibility_report(company)

Returns a published, dated Siteiz AI visibility report for a well-known company's homepage (score, grade, pillar scores, top issues). Omit the company to list all published reports.

Input Schema

{
  "type": "object",
  "properties": {
    "company": {
      "type": "string",
      "description": "Company name, e.g. Notion. Omit to list all reports."
    }
  },
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "reports": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "company": {
            "type": "string"
          },
          "score": {
            "type": "integer"
          },
          "grade": {
            "type": "string"
          },
          "scannedAt": {
            "type": "string"
          },
          "url": {
            "type": "string"
          }
        }
      }
    },
    "company": {
      "type": "string"
    },
    "url": {
      "type": "string"
    },
    "scannedAt": {
      "type": "string"
    },
    "score": {
      "type": "integer"
    },
    "grade": {
      "type": "string"
    },
    "pillars": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": {
            "type": "string"
          },
          "label": {
            "type": "string"
          },
          "score": {
            "type": "integer"
          },
          "notAssessed": {
            "type": "boolean"
          }
        }
      }
    },
    "report": {
      "type": "string"
    }
  },
  "description": "Either `reports` (a list, when no company matched) or the single report fields."
}

Community

Rate this Server

Evidence

Recent observations

verifiedversion not recorded4 tools