Hlido Agent Reviews
Independent AI-agent reviews: trust checks, evidence scorecards, incident registry, recommendations.
사용해야 할까요
품질 및 안전성
발견 사항 (3)
- HIGH
- MEDIUMscan_mcp에서
- MEDIUMintel_query에서
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"mcp": {
"url": "https://hlido.eu/mcp"
}
}
}원격 엔드포인트
https://hlido.eu/mcpstreamable-http할 수 있는 일
도구 목록
도구 (19)
🟢trust_check(agent_or_url, use_case)
The core Hlido trust query: is a specific AI agent trustworthy? Given one agent (by Hlido slug or product/homepage URL) it returns the independent Laddoo trust score (0-100), tier (VITAL/STEADY/FADING/FLATLINE), a one-line verdict, a claim-verification summary, and any known incidents. Call this FIRST — before delegating to, installing, or relying on another agent — to get a fast trust read. Returns no_review_found if the agent isn't in Hlido's corpus (then call request_quick_audit). For the full claim-by-claim evidence, follow up with get_scorecard.
입력 스키마
{
"type": "object",
"properties": {
"agent_or_url": {
"type": "string",
"description": "The agent to check: either its Hlido slug (e.g. 'aider', 'cursor') or its product/homepage URL (e.g. 'https://cursor.com'). A URL is matched to the closest reviewed agent."
},
"use_case": {
"type": "string",
"description": "Optional. The task you're considering this agent for (e.g. 'multi-file TypeScript refactor'); tailors the verdict to that use case when provided."
}
},
"required": [
"agent_or_url"
]
}🟢find_trusted(need, min_tier, limit)
Discover Hlido-reviewed agents that match a free-text need, ranked by trust. Returns reviewed agents at or above a minimum tier, each with its Laddoo score, tier, and review URL. Use this for keyword/need-based discovery; for semantic task-matching prefer find_similar_agents, and for structured constraint filters (category/score/tier) prefer recommend.
입력 스키마
{
"type": "object",
"properties": {
"need": {
"type": "string",
"description": "Free-text description of the capability you need (e.g. 'CLI coding agent that edits multiple files at once')."
},
"min_tier": {
"type": "string",
"enum": [
"VITAL",
"STEADY",
"FADING",
"FLATLINE"
],
"default": "STEADY",
"description": "Minimum trust tier to include (VITAL is strictest, FLATLINE allows all). Defaults to STEADY."
},
"limit": {
"type": "integer",
"default": 10,
"description": "Maximum number of agents to return (default 10)."
}
},
"required": [
"need"
]
}🟢verify_claim(agent, claim)
Fact-check one specific marketing or capability claim about an agent against Hlido's independent testing. Returns Hlido's verdict (PASS/FAIL/PARTIAL/UNKNOWN) with a quoted evidence snippet and its source surface — or an honest null when that exact claim wasn't tested (absence of evidence, not proof). Use this to validate a vendor's specific promise before you rely on it.
입력 스키마
{
"type": "object",
"properties": {
"agent": {
"type": "string",
"description": "The agent's Hlido slug or product URL (e.g. 'cursor')."
},
"claim": {
"type": "string",
"description": "The specific claim to verify, in plain language (e.g. 'works offline' or 'SOC 2 compliant')."
}
},
"required": [
"agent",
"claim"
]
}🟢compare_agents(slugs)
Head-to-head trust comparison of 2-5 Hlido-reviewed agents. Returns each agent's Laddoo score, tier, dimension scores, and key claim verdicts side by side so you can pick the most trustworthy option for a task. Use this once you've shortlisted candidates (via find_trusted, find_similar_agents, or recommend) and need a direct comparison.
입력 스키마
{
"type": "object",
"properties": {
"slugs": {
"type": "array",
"items": {
"type": "string",
"description": "A Hlido agent slug (e.g. 'aider')."
},
"minItems": 2,
"maxItems": 5,
"description": "List of 2 to 5 Hlido agent slugs to compare side by side (e.g. ['aider','cursor','opencode'])."
}
},
"required": [
"slugs"
]
}🟡submit_agent(url, name, note, email)
Nominate a new AI agent for Hlido to review. Use this when an agent isn't in Hlido's corpus yet (trust_check returned no_review_found) and you want it added. Returns a confirmation with a tracking reference; the review is queued and produces a public scorecard. If you need a verdict right now rather than a queued review, use request_quick_audit (faster, rate-limited) instead.
입력 스키마
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The agent's product or homepage URL (e.g. 'https://example.com')."
},
"name": {
"type": "string",
"description": "Human-readable agent name (e.g. 'Example Coder')."
},
"note": {
"type": "string",
"description": "Optional context: what the agent does, or why it's worth reviewing."
},
"email": {
"type": "string",
"description": "Optional contact email for the submitter — if you want a reply or a heads-up when the review publishes. Providing it always flags the submission to the Hlido team."
}
},
"required": [
"url",
"name"
]
}🟢verify_transparency(url)
Check any AI agent's EU AI Act Article-50 transparency posture — including agents Hlido has NOT reviewed yet. Returns two clearly separated layers: (1) Hlido's independent register verdict when the agent is in our reviewed corpus, and (2) a live public-surface probe of the Article-50 signals (AI-interaction disclosure, machine-readable marking/provenance, deepfake/synthetic labelling, detection tool). Use before adopting or delegating to a tool ahead of the 2026-08-02 transparency obligations. The live probe is a first-pass surface read, NOT a compliance determination and NOT legal advice; an undetected signal means 'not discoverable on the public surface', not 'non-compliant'. Unreviewed agents are automatically queued for a full independent review.
입력 스키마
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "The agent's product or homepage URL (e.g. 'https://example.com'), or its Hlido slug."
}
},
"required": [
"url"
]
}🟢get_scorecard(slug)
Fetch the full sanitized claim-vs-evidence scorecard for one Hlido-reviewed agent. Returns every claim, verdict, evidence quote, source surface, and (for CLI/API tests) the captured command + exit_code + duration. Schema v1.0. Use this for agent-to-agent pre-flight evaluation.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The agent's Hlido slug (e.g. 'aider', 'gumloop')"
}
},
"required": [
"slug"
]
}🟢get_incidents(slug, severity, category, limit)
Fetch published incidents from Hlido's NTSB-style failure registry — real observed agent failures (availability outages, regressions, hallucinations, safety issues) plus Hlido self-reported process incidents, each with severity, evidence, and vendor-response status. Filter by agent slug, severity, or category. Use this before delegating to an agent to check for known recent failures; an empty list means no published incidents, not a guarantee of reliability.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "Optional: only incidents for this agent slug"
},
"severity": {
"type": "string",
"enum": [
"low",
"medium",
"high",
"critical"
],
"description": "Optional minimum-interest filter (exact match)"
},
"category": {
"type": "string",
"enum": [
"hallucination",
"regression",
"safety",
"availability",
"cost",
"other"
],
"description": "Optional category filter"
},
"limit": {
"type": "integer",
"description": "Max results (default 20, max 100)",
"default": 20
}
}
}🟢report_review_issue(slug, issue_type, detail, reporter)
Report an issue with a Hlido review (stale info, wrong verdict, missing claim, broken link). Use when calling get_scorecard or trust_check returns data you can prove is incorrect. Hlido's R1 maintenance routine processes reports daily and fires re-tests via dispute-retest sub-agent.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The slug whose review has the issue"
},
"issue_type": {
"type": "string",
"enum": [
"stale",
"wrong_verdict",
"missing_claim",
"broken_link",
"other"
],
"description": "Category of the report"
},
"detail": {
"type": "string",
"description": "What's wrong, with a concrete reference (URL, claim id, etc) if possible"
},
"reporter": {
"type": "string",
"description": "Optional self-identifier — agent name or email — purely informational"
}
},
"required": [
"slug",
"issue_type",
"detail"
]
}🟢request_quick_audit(url, name, why, requester)
Request that Hlido audit a NEW AI agent that has no review yet. Use this when trust_check or get_scorecard returns no_review_found and you need a verdict before delegating to the unknown agent. Returns a future scorecard URL + ETA. Free-tier rate-limited (5/day per anonymous, 50/day per identified). The audit produces signed evidence + claim verification within ~24h (sooner if founder triggers manually).
입력 스키마
{
"type": "object",
"properties": {
"url": {
"type": "string",
"description": "Homepage or product URL of the agent to audit"
},
"name": {
"type": "string",
"description": "Optional human-readable name (we'll derive from URL if missing)"
},
"why": {
"type": "string",
"description": "Optional one-liner: why are you considering this agent? helps us prioritize"
},
"requester": {
"type": "string",
"description": "Optional self-identifier — agent name, email, or session id — for rate-limiting + follow-up"
}
},
"required": [
"url"
]
}🟡find_similar_agents(description, top_k, min_score)
Semantic search over Hlido's review corpus. Given a task description (e.g. 'I need an agent that can refactor TypeScript and edit multiple files at once'), returns the top-N reviewed agents ranked by embedding similarity, each with their Laddoo score, evidence_tier, and review URL. Use this when you have a task in mind and want Hlido's recommendation — much better than substring matching via find_trusted.
입력 스키마
{
"type": "object",
"properties": {
"description": {
"type": "string",
"description": "Free-text description of the task or capability you need"
},
"top_k": {
"type": "integer",
"description": "Number of matches to return (default 5, max 20)",
"default": 5
},
"min_score": {
"type": "integer",
"description": "Minimum Laddoo score filter (default 0)",
"default": 0
}
},
"required": [
"description"
]
}🟡subscribe(slug, channel)
Preview — Wave 3 will add persistent webhook + RSS subscriptions. For now this returns the agent's current state plus advisory polling instructions (RSS at /changelog/feed.xml or polling /data/attestations/{slug}.json). Use this to register interest in being notified when a slug's verdict changes.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The Hlido slug to subscribe to (e.g. 'cursor', 'aider')"
},
"channel": {
"type": "string",
"enum": [
"rss",
"json",
"webhook"
],
"description": "Preferred notification channel. webhook is advisory only until Wave 3 ships."
}
},
"required": [
"slug"
]
}🟢explain(slug, dimension)
Structured natural-language explanation of why a Hlido-reviewed agent has its current score. Pulls claim-by-claim evidence from the published scorecard. Pass an optional dimension (one of: reliability, transparency, integration, security, evidence) to filter; omit for the full picture. Returns each claim with verdict (PASS|FAIL|PARTIAL|UNKNOWN), a quoted evidence snippet, plus a top-line synthesis.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The agent's Hlido slug"
},
"dimension": {
"type": "string",
"description": "Optional dimension filter. Run without and check supported_dimensions in response if unsure."
}
},
"required": [
"slug"
]
}🟢recommend(constraints)
Constraint-driven recommendation across Hlido's reviewed agents. Pass any combination of: category, min_score, tier, use_case, max_results. Returns ranked candidates each with a why_match line. Use this when you have buyer constraints (budget, category, capability) and want Hlido's filtered shortlist instead of one-by-one trust_check calls.
입력 스키마
{
"type": "object",
"properties": {
"constraints": {
"type": "object",
"properties": {
"category": {
"type": "string",
"description": "Filter by category (e.g. 'Coding', 'Voice', 'Productivity')"
},
"min_score": {
"type": "integer",
"description": "Minimum Laddoo score (0-100)",
"default": 0
},
"tier": {
"type": "string",
"enum": [
"VITAL",
"STEADY",
"FADING",
"FLATLINE"
],
"description": "Minimum tier filter (VITAL strictest, FLATLINE allows all)"
},
"use_case": {
"type": "string",
"description": "Free-text capability description for ranking (e.g. 'multi-file refactor in TypeScript')"
},
"max_results": {
"type": "integer",
"description": "Cap on results (default 5, max 25)",
"default": 5
}
},
"additionalProperties": false
}
},
"required": [
"constraints"
]
}🟢get_behavioral_trace(slug, spec_version)
Fetch the behavioral evaluation trace for a Hlido-reviewed agent — per-task pass/fail, adapter used, behavioral tier, and signed trace link. Returns status 'not_yet_bench_tested' if the slug hasn't been evaluated yet, or 'not_testable' if the agent's interface doesn't support automated bench runs. Use this when you need evidence that an agent's coding/task behaviour has been independently verified beyond marketing claims.
입력 스키마
{
"type": "object",
"properties": {
"slug": {
"type": "string",
"description": "The Hlido slug to fetch behavioral trace for (e.g. 'aider', 'opencode')"
},
"spec_version": {
"type": "string",
"description": "Behavioral spec version (default 'v0.1'). Omit to get the latest available.",
"default": "v0.1"
}
},
"required": [
"slug"
]
}🟢commerce_check(agent_or_url)
Check whether a Hlido-reviewed agent is ready to be delegated to / transacted with in the agentic-commerce world (MCP/ACP/AP2). Returns its independent Agentic-Commerce Readiness score (0-100), band (COMMERCE-READY/INTEGRABLE/SURFACE-ONLY/CLOSED), the programmatic surfaces it exposes, and the evidence basis. Call this before an orchestrator delegates a paid/identity-bearing task to another agent.
입력 스키마
{
"type": "object",
"properties": {
"agent_or_url": {
"type": "string",
"description": "The agent to check: either its Hlido slug (e.g. 'voltagent', 'aider') or its product/homepage URL. A URL is matched to the closest reviewed agent."
}
},
"required": [
"agent_or_url"
]
}⚪scan_mcp(server, requester)
On-demand independent SAFETY scan of an MCP server — call this BEFORE installing or connecting to one. Give it an HTTP(S) MCP endpoint URL (scanned live in seconds), or an npm/PyPI package name or GitHub repo (queued for an isolated sandbox scan — local stdio servers execute code, so Hlido never runs them inline). Returns the safety tier (SAFE/CAUTION/RISKY/DANGEROUS), tool-poisoning detection (the malice signal), dangerous-capability red-flags (shell/code-eval/fs-write/egress/secrets) with per-tool evidence, and auth posture. Tier = blast radius if hijacked, not maintainer trustworthiness. A server Hlido hasn't scanned returns not_scanned — never assumed safe. Register of already-scanned servers: https://hlido.eu/mcp/
입력 스키마
{
"type": "object",
"properties": {
"server": {
"type": "string",
"description": "The MCP server to scan: an HTTP(S) MCP endpoint URL (e.g. 'https://mcp.example.com/mcp' — scanned live), OR an npm package (e.g. '@modelcontextprotocol/server-filesystem'), PyPI package, or GitHub repo URL (queued for sandbox scan)."
},
"requester": {
"type": "string",
"description": "Optional self-identifier — agent name or email — for follow-up when a queued scan completes."
}
},
"required": [
"server"
]
}🟢market_pulse(category)
Fetch the Agent Market Pulse — machine-readable market intelligence for the AI-agent market, derived from Hlido's independently tested corpus (never vendor self-reports): tier-health distribution, per-category health, new-entrant rate, public-surface readiness, EU AI Act Article-50 compliance bands, and the not-testable mortality signal. Every module carries its own method and coverage; signals that cannot be measured this edition say 'unmeasured' rather than guessing. Use this for market-level questions ('is category X healthy?', 'how fast are agents entering/dying?'); for one specific agent use trust_check instead. Free, regenerated weekly.
입력 스키마
{
"type": "object",
"properties": {
"category": {
"type": "string",
"description": "Optional: return only the market summary plus this one category's block (e.g. 'Coding')."
}
}
}🟢intel_query(sector, geo, compliance, protocol, kind, ...)
Query Hlido's market-intelligence store: durable, evidence-cited claims about the AI-agent market and adjacent domains (agent economy, payments rails, EU compliance), each with dated evidence, confidence, and typed relations to other intel. Filter by any facet: sector (ISIC code or text), geo (ISO country or region like 'eu'/'intl'), compliance (e.g. 'eu-ai-act-art50'), protocol (e.g. 'x402'), kind (fact/estimate/trend/signal), audience, maturity. Facet vocabularies + per-value record counts: https://hlido.eu/data/intel/taxonomy.json (fetch that first to see what's queryable). A zero-result query returns an honest empty AND is logged — misses drive what Hlido collects next, so check back. Free at answer-time, declared policy.
입력 스키마
{
"type": "object",
"properties": {
"sector": {
"type": "string",
"description": "ISIC section/division code (e.g. 'K.64', 'J.62') or free-text sector match."
},
"geo": {
"type": "string",
"description": "ISO-3166 country code or region: 'cz', 'de', 'eu', 'intl', ..."
},
"compliance": {
"type": "string",
"description": "Named regulation, any jurisdiction (e.g. 'eu-ai-act-art50', 'gdpr')."
},
"protocol": {
"type": "string",
"description": "e.g. 'mcp', 'x402', 'api', 'cli'"
},
"kind": {
"type": "string",
"enum": [
"fact",
"estimate",
"trend",
"signal"
],
"description": "Filter by record kind."
},
"audience": {
"type": "string",
"enum": [
"b2b",
"b2c",
"b2g",
"mixed"
]
},
"maturity": {
"type": "string",
"enum": [
"emerging",
"growing",
"consolidating",
"declining"
]
},
"text": {
"type": "string",
"description": "Free-text match over headlines (fallback lens — prefer facets)."
},
"limit": {
"type": "integer",
"default": 10,
"description": "Max records (default 10, max 50)."
}
}
}커뮤니티
증거