Crawl Census

Ask before you fetch: will this domain serve your crawler, refuse it, or charge it?

사용해야 할까요

품질 및 안전성

B
설명 품질
97%
스키마 완전성
80%
이름 품질
83%
오염 위험
80%
권한 일치
100%
프로토콜 준수
100%

발견 사항 (2)

  • HIGHTool poisoning patterns detected
  • MEDIUMTool description contains URL to non-standard domaincrawl_preflight에서

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~963토큰 (도구 정의)
~482 B일반적인 응답 크기
중간 정도의 주의 영향 (128k 컨텍스트의 0.75%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "crawl-census": {
      "url": "https://crawlcensus.com/mcp"
    }
  }
}

원격 엔드포인트

https://crawlcensus.com/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (7)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢scan_site(domain)

Run a live AI-accessibility audit of a domain: robots.txt policy for every tracked AI crawler, live user-agent probes, JavaScript-free readability, structured data and llms.txt. Returns a score out of 100 with per-check detail.

입력 스키마

{
  "type": "object",
  "properties": {
    "domain": {
      "type": "string",
      "description": "Bare hostname, for example example.com"
    }
  },
  "required": [
    "domain"
  ]
}
⚪site_report(domain)

Return the most recent stored audit for a domain without triggering a new scan. Faster and free of load on the target site.

입력 스키마

{
  "type": "object",
  "properties": {
    "domain": {
      "type": "string",
      "description": "Bare hostname"
    }
  },
  "required": [
    "domain"
  ]
}
⚪census_stats

Corpus-level statistics: how many measured domains block each AI crawler, mean access score, llms.txt adoption.

입력 스키마

{
  "type": "object",
  "properties": {}
}
🟢crawl_preflight(agent, domains)

Decide whether a crawler may fetch a list of domains before spending requests on them. Works for any crawler token, not only the ones this census tracks: an unrecognised agent is resolved from each domain's stored robots.txt rather than refused. For each domain returns one of: allow (robots permits it and a live request carrying that agent's user agent was served), disallow (robots.txt forbids it), refuse (robots permits it but the edge refused the agent anyway, so the allowance is not real), pay (the origin answered HTTP 402 Payment Required, meaning it will serve this agent on commercial terms), or unknown. The full definition of each, including what it obliges a crawler to do, is published at https://crawlcensus.com/api/v1/verdicts. Built for crawler operators rather than site owners: it prevents wasted fetches against doors that are shut, and flags content an operator is trying to sell rather than withhold.

입력 스키마

{
  "type": "object",
  "properties": {
    "agent": {
      "type": "string",
      "description": "Crawler token, e.g. gptbot, claudebot, perplexitybot, oai-searchbot, ccbot."
    },
    "domains": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "Domains to check. Up to 25 per call anonymously; send an Authorization: Bearer key for more. An over-large batch is refused outright rather than partly answered."
    }
  },
  "required": [
    "agent",
    "domains"
  ]
}
⚪agent_profile(agent)

What this census measures and publishes about one AI crawler: how often it is disallowed in robots.txt, how often live requests carrying its user agent are refused at the network edge whatever robots.txt says, whether its operator documents it as honouring robots.txt, and where to correct any of that. Intended for the operator of the agent as much as for anyone studying it, so it includes the correction channel and the public page a claim can be disputed against.

입력 스키마

{
  "type": "object",
  "properties": {
    "agent": {
      "type": "string",
      "description": "Crawler token, e.g. gptbot, claudebot, ccbot, google-extended."
    }
  },
  "required": [
    "agent"
  ]
}
⚪census_facts

Every headline finding from the census as discrete, dated records rather than prose. Each carries its value, unit, denominator, measurement date, the page it comes from and a ready-made citation line, plus the caveats that apply to all of them. Use this when answering a question about how open the web is to AI crawlers: lifting a percentage out of a rendered page loses the denominator and the date, which is what makes the number wrong when it is repeated.

입력 스키마

{
  "type": "object",
  "properties": {}
}
🟡submit_domains(domains)

Queue domains the census has not measured yet so a later crawl_preflight can answer them. This closes the loop crawl_preflight starts: anything it returns as unknown with measurable true is worth submitting, and the reply names any that were already fresh or that this census will never measure, so a caller looping over its own unknowns converges instead of resubmitting the same set. Queueing is a database write rather than a fetch, so the allowance is far higher than scan_site and submitted domains are measured ahead of the ranked backlog.

입력 스키마

{
  "type": "object",
  "properties": {
    "domains": {
      "type": "array",
      "items": {
        "type": "string"
      },
      "description": "Hostnames to queue. Up to 50 per call anonymously; an over-large batch is refused outright rather than partly queued."
    }
  },
  "required": [
    "domains"
  ]
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 7개
검증됨버전이 기록되지 않음도구 7개