protein-mcp-server
MCP Server for 3D protein structural data retrieval & analysis from RCSB PDB, PDBe, and UniProt.
사용해야 할까요
품질 및 안전성
도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.
컨텍스트 비용
이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.
설치
원클릭 설치
`claude_desktop_config.json` 파일에 다음을 추가하세요:
{
"mcpServers": {
"protein-mcp-server": {
"command": "bun",
"args": [
"protein-mcp-server"
]
}
}
}실행 가능한 패키지
1.0.3streamable-http원격 엔드포인트
https://protein.caseyjhand.com/mcpstreamable-http할 수 있는 일
도구 목록
도구 (7)
🟢protein_search_structures(query, sequence, organism, method, max_resolution, ...)
Search experimental (PDB) and predicted (computed-model) protein structures by free text, protein sequence (triggers an mmseqs2 similarity search), and/or organism, method, and resolution filters. Returns ranked hits; the experimental page is enriched with title, method, resolution, and organism. Chain hit IDs into protein_get_structure. Optionally returns a facet breakdown (counts by method / organism / release year / …) alongside the hits in the same response. A facet on a dimension you are already filtering (e.g. the organism facet while organism is set) lists unfiltered alternatives by design — it does not constrain by its own active filter, so you can see sibling values to pivot to. Numeric histogram buckets (resolution, molecular weight) carry explicit rangeFrom/rangeTo bounds so a boundary label is unambiguous.
입력 스키마
{
"type": "object",
"properties": {
"query": {
"description": "Free-text query (protein name, gene, keyword, PDB title terms).",
"type": "string"
},
"sequence": {
"description": "One-letter amino-acid sequence; triggers an RCSB mmseqs2 sequence-similarity search.",
"type": "string"
},
"organism": {
"description": "Filter by source organism scientific name (e.g. \"Homo sapiens\").",
"type": "string"
},
"method": {
"description": "Filter by experimental method (e.g. \"X-RAY DIFFRACTION\", \"ELECTRON MICROSCOPY\").",
"type": "string"
},
"max_resolution": {
"description": "Maximum resolution in Å (lower is sharper); applies to experimental structures.",
"type": "number",
"exclusiveMinimum": 0
},
"min_identity": {
"description": "Minimum sequence identity (0–1) for a sequence search. Requires sequence — supplying it without one is rejected, since the threshold filters only a sequence search. Default 0.",
"type": "number",
"minimum": 0,
"maximum": 1
},
"max_evalue": {
"description": "Maximum E-value for a sequence search. Requires sequence — supplying it without one is rejected, since the threshold filters only a sequence search. Default 1.",
"type": "number",
"exclusiveMinimum": 0
},
"content_type": {
"default": "all",
"description": "Which structure universe to search: experimental (PDB), predicted (computed models), or all.",
"type": "string",
"enum": [
"experimental",
"predicted",
"all"
]
},
"facets": {
"description": "Optional dimensions to summarize as a facet breakdown alongside the hits. Each dimension at most once; repeating one is rejected.",
"type": "array",
"items": {
"type": "string",
"enum": [
"method",
"organism",
"polymer_type",
"resolution",
"release_year",
"molecular_weight"
]
}
},
"limit": {
"default": 25,
"description": "Maximum hits to return (1–100).",
"type": "integer",
"minimum": 1,
"maximum": 100
},
"start": {
"default": 0,
"description": "Zero-based result offset. Combine with limit to retrieve later pages.",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"hits": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Structure identifier (PDB entry ID or computed-model ID)."
},
"entityId": {
"description": "Matched polymer-entity ID for sequence hits (e.g. 4HHB_1, AF_AFP69905F1_1); id remains the chainable entry ID.",
"type": "string"
},
"source": {
"type": "string",
"enum": [
"experimental",
"predicted"
],
"description": "Which universe the hit came from."
},
"score": {
"description": "RCSB relevance score.",
"type": "number"
},
"uniprotAccession": {
"description": "UniProt accession parsed from a computed-model ID, when available.",
"type": "string"
},
"title": {
"description": "Structure title (enriched experimental hits).",
"type": "string"
},
"method": {
"description": "Experimental method(s) (enriched experimental hits).",
"type": "string"
},
"resolution": {
"description": "Resolution in Å (enriched experimental hits).",
"type": "number"
},
"organism": {
"description": "Primary source organism (enriched experimental hits).",
"type": "string"
}
},
"required": [
"id",
"source"
],
"additionalProperties": false,
"description": "A ranked structure hit with optional enrichment metadata."
},
"description": "Ranked structure hits."
},
"facets": {
"description": "Optional facet breakdown when requested — one flat dimension per requested facet. Cross-tabs (a dimension nested inside another) are protein_analyze_collection territory.",
"type": "array",
"items": {
"type": "object",
"properties": {
"dimension": {
"type": "string",
"description": "Friendly dimension name (e.g. method, organism, release_year)."
},
"buckets": {
"type": "array",
"items": {
"type": "object",
"properties": {
"label": {
"type": "string",
"description": "Bucket value — category, numeric bin start, or period."
},
"count": {
"type": "number",
"description": "Number of entries in the bucket."
},
"rangeFrom": {
"description": "Inclusive lower bound of a numeric histogram bin (= label). Present only for numeric facets (resolution, molecular_weight); absent for term/period facets.",
"type": "number"
},
"rangeTo": {
"description": "Exclusive upper bound of a numeric histogram bin (rangeFrom + bin interval), so the bin covers [rangeFrom, rangeTo). Present only for numeric facets.",
"type": "number"
}
},
"required": [
"label",
"count"
],
"additionalProperties": false,
"description": "A leaf aggregation bucket: a value and its entry count."
},
"description": "Aggregation buckets, count-descending for terms."
},
"truncated": {
"description": "True when buckets were capped by the per-dimension bucket limit.",
"type": "boolean"
},
"missingValueCount": {
"type": "number",
"description": "Matches in scope that carry no value for this attribute, so they count toward the response total but fall in no bucket (e.g. computed models have no experimental method; solution-NMR entries have no diffraction resolution). Independent of truncation — measured before the bucket cap is applied. 0 means no shortfall was detectable: a multi-valued attribute such as organism can place one match in several buckets, which offsets the shortfall rather than adding to it."
}
},
"required": [
"dimension",
"buckets",
"missingValueCount"
],
"additionalProperties": false,
"description": "A facet dimension and its aggregation buckets."
}
},
"totalCount": {
"type": "number",
"description": "Total matches upstream before pagination."
},
"start": {
"type": "number",
"description": "Zero-based offset of the returned page."
},
"nextStart": {
"description": "Offset for the next page; absent on the final or past-end page.",
"type": "number"
},
"effectiveQuery": {
"description": "Echoed text query for follow-up calls.",
"type": "string"
},
"notice": {
"description": "Advisory note (empty results, predicted-search caveats, truncation, facet dimensions whose buckets cover materially less than totalCount). Carries every applicable advisory in one string.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `no_criteria`: No query, sequence, organism, method, or maximum resolution was provided — nothing to search on. `sequence_modifier_without_sequence`: min_identity or max_evalue was supplied with no sequence, so the threshold would never reach a sequence search. `duplicate_dimension`: facets lists the same dimension twice, which would return that breakdown twice. Other values are possible when a failure originates below the handler.",
"examples": [
"no_criteria",
"sequence_modifier_without_sequence",
"duplicate_dimension"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"hits",
"totalCount",
"start"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_get_structure(ids, source, include_coords, sections)
Fetch structures with metadata and coordinate-file URLs. source "experimental" takes PDB entry IDs, and also resolves the computed-model IDs protein_search_structures returns (AF_*/MA_*), which come back marked source "predicted" with their modelling provider; "predicted" takes UniProt accessions (AlphaFold, with pLDDT/PAE confidence); "best_available" takes UniProt accessions and returns the top federated model — the highest-resolution experimental structure if one exists (optimizing resolution, not biological representativeness, so it can return an engineered mutant over the wild-type entry), else the best prediction. Records fetched under source "experimental", computed models included, also carry polymer entities with both chain namespaces (labelAsymIds for protein_compare_structures, authAsymIds for protein_get_annotations), bound ligands, molecular weight, and release date. Resolves up to the configured batch cap per call with per-ID partial success — missed IDs are listed in failed[], and IDs beyond the cap are reported in the notice. Set include_coords to inline coordinate content; if the inlined bytes exceed the response budget the content is withheld and overflow lists each structure's size — re-call with sections:[ids] for specific structures, or for a single oversized file download it from that record's coordinateUrls.
입력 스키마
{
"type": "object",
"properties": {
"ids": {
"minItems": 1,
"type": "array",
"items": {
"type": "string",
"minLength": 1
},
"description": "PDB entry IDs or computed-model IDs such as AF_AFP69905F1 (source experimental), or UniProt accessions (predicted / best_available)."
},
"source": {
"default": "experimental",
"description": "Where to fetch: experimental (PDB), predicted (AlphaFold), or best_available (federated pick).",
"type": "string",
"enum": [
"experimental",
"predicted",
"best_available"
]
},
"include_coords": {
"default": false,
"description": "Inline coordinate-file content — mmCIF, or PDB format when no mmCIF URL is available; coordinateFormat names which. Off by default — URLs are always returned.",
"type": "boolean"
},
"sections": {
"description": "Structure IDs to inline coordinates for, from a prior overflow outline.",
"type": "array",
"items": {
"type": "string"
}
}
},
"required": [
"ids"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"structures": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Structure identifier (PDB entry ID or UniProt accession)."
},
"source": {
"type": "string",
"enum": [
"experimental",
"predicted"
],
"description": "Whether the structure is experimental or predicted."
},
"pdbId": {
"description": "Chosen PDB entry ID when a best_available query resolved to an experimental structure (id stays the queried UniProt accession), so the structure can be cited without parsing the coordinate URL.",
"type": "string"
},
"title": {
"description": "Structure / protein title.",
"type": "string"
},
"method": {
"description": "Experimental method(s).",
"type": "string"
},
"resolution": {
"description": "Resolution in Å (experimental).",
"type": "number"
},
"organism": {
"description": "Source organism.",
"type": "string"
},
"molecularWeight": {
"description": "Deposited structure molecular weight in kDa. Present only when the call used source experimental; omitted when the record does not report one.",
"type": "number"
},
"releaseDate": {
"description": "Initial release date (ISO 8601). Present only when the call used source experimental; omitted when the record does not report one.",
"type": "string"
},
"polymerEntities": {
"description": "Modeled polymer entities, each with both chain namespaces — authAsymIds for protein_get_annotations.chain, labelAsymIds for protein_compare_structures.chain. Present when the call used source experimental (computed-model IDs included); omitted under source predicted or best_available and for entries with none.",
"type": "array",
"items": {
"type": "object",
"properties": {
"entityId": {
"type": "string",
"description": "Polymer entity ID (e.g. 4HHB_1)."
},
"description": {
"description": "Entity description.",
"type": "string"
},
"organism": {
"description": "Source organism.",
"type": "string"
},
"authAsymIds": {
"description": "Author-assigned chain IDs (auth_asym_id, e.g. [\"A\",\"C\"]) — the namespace protein_get_annotations.chain takes. Omitted when upstream does not report them.",
"type": "array",
"items": {
"type": "string"
}
},
"labelAsymIds": {
"description": "mmCIF label_asym_id chain IDs (e.g. [\"I\",\"OB\"]) — the namespace protein_compare_structures.chain takes. Not derivable from authAsymIds: for 6QNR_9 the label chains are I/OB against author chains 82/8E. Omitted when upstream does not report them.",
"type": "array",
"items": {
"type": "string"
}
},
"sequenceLength": {
"description": "Residue count of the sample sequence.",
"type": "number"
}
},
"required": [
"entityId"
],
"additionalProperties": false,
"description": "A modeled polymer entity (chain group), with both chain namespaces."
}
},
"ligands": {
"description": "Bound non-polymer components. Present when the call used source experimental; omitted under source predicted or best_available, when the entry binds none, or when the record carries no ligand data.",
"type": "array",
"items": {
"type": "object",
"properties": {
"compId": {
"type": "string",
"description": "Chemical component ID (e.g. HEM)."
},
"name": {
"description": "Chemical name.",
"type": "string"
},
"formula": {
"description": "Molecular formula.",
"type": "string"
}
},
"required": [
"compId"
],
"additionalProperties": false,
"description": "A bound ligand (non-polymer chemical component)."
}
},
"provider": {
"description": "Model provider (predicted / best_available).",
"type": "string"
},
"confidence": {
"description": "Model confidence on its native scale; confidenceType names the metric and range (e.g. pLDDT 0–100, QMEANDisCo 0–1). Present for predicted models that report a score.",
"type": "number"
},
"confidenceType": {
"description": "Name of the confidence metric carried in confidence (e.g. \"pLDDT\", \"QMEANDisCo\"). Providers score on different scales; this names the one in use.",
"type": "string"
},
"meanPlddt": {
"description": "Mean pLDDT confidence (0–100); present only for pLDDT-scored models. For other metrics read confidence + confidenceType.",
"type": "number"
},
"confidenceBuckets": {
"description": "pLDDT confidence-band fractions (predicted).",
"type": "object",
"properties": {
"veryLow": {
"type": "number",
"description": "Fraction of residues with pLDDT < 50."
},
"low": {
"type": "number",
"description": "Fraction with pLDDT 50–70."
},
"confident": {
"type": "number",
"description": "Fraction with pLDDT 70–90."
},
"veryHigh": {
"type": "number",
"description": "Fraction with pLDDT > 90."
}
},
"required": [
"veryLow",
"low",
"confident",
"veryHigh"
],
"additionalProperties": false
},
"paeDocUrl": {
"description": "Predicted Aligned Error documentation URL (predicted).",
"type": "string"
},
"coordinateUrls": {
"type": "object",
"properties": {
"cif": {
"description": "mmCIF coordinate file URL.",
"type": "string"
},
"pdb": {
"description": "PDB-format coordinate file URL.",
"type": "string"
},
"bcif": {
"description": "Binary CIF coordinate file URL.",
"type": "string"
}
},
"additionalProperties": false,
"description": "Coordinate file download URLs. A format with no published file is omitted — e.g. pdb for large entries archived as mmCIF only."
},
"coordinateFormat": {
"description": "Format of inlined coordinates, when present.",
"type": "string",
"enum": [
"cif",
"pdb",
"bcif"
]
},
"coordinates": {
"description": "Inlined coordinate-file content (only when include_coords).",
"type": "string"
}
},
"required": [
"id",
"source",
"coordinateUrls"
],
"additionalProperties": false,
"description": "A resolved structure with metadata and coordinate-file URLs."
},
"description": "Resolved structures (metadata always present)."
},
"failed": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Requested ID that failed."
},
"reason": {
"type": "string",
"description": "Why it failed."
}
},
"required": [
"id",
"reason"
],
"additionalProperties": false,
"description": "A requested ID that did not resolve, with the reason."
},
"description": "IDs that could not be resolved (partial success)."
},
"attribution": {
"type": "array",
"items": {
"type": "object",
"properties": {
"source": {
"type": "string",
"description": "Contributing data-source display name (e.g. \"RCSB PDB\", \"AlphaFold DB\", \"SWISS-MODEL\", \"UniProt\"). Open-ended — best_available structures are federated through 3D-Beacons providers."
},
"license": {
"type": "string",
"description": "License the source data is released under (e.g. \"CC BY 4.0\")."
},
"citation": {
"type": "string",
"description": "Primary-literature citation to credit the source."
},
"homepage": {
"type": "string",
"description": "Source homepage (absolute URL)."
}
},
"required": [
"source",
"license",
"citation",
"homepage"
],
"additionalProperties": false,
"description": "Upstream data-source attribution: license, citation, and homepage."
},
"description": "Upstream data-source licenses and citations for every source present in structures[] — RCSB PDB for experimental records, the modelling provider (AlphaFold DB, ModelArchive, SWISS-MODEL, …) for predicted ones. Always present — the attribution obligation travels with the data."
},
"overflow": {
"description": "Present only when inlined coordinates across the batch exceeded the response budget.",
"type": "object",
"properties": {
"sections": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Structure ID whose coordinates were withheld."
},
"bytes": {
"type": "number",
"description": "Serialized size of the withheld coordinate content."
}
},
"required": [
"id",
"bytes"
],
"additionalProperties": false,
"description": "A withheld structure and its coordinate byte size."
},
"description": "Per-structure coordinate sizes available for targeted re-call."
},
"notice": {
"type": "string",
"description": "How to retrieve specific coordinates via the sections parameter."
}
},
"required": [
"sections",
"notice"
],
"additionalProperties": false
},
"requested": {
"type": "number",
"description": "Number of IDs in the original request, before the batch cap was applied."
},
"processed": {
"type": "number",
"description": "Number of IDs actually processed after the batch cap. Lower than requested means the excess IDs were ignored and never looked up — re-submit them in a follow-up call."
},
"resolved": {
"type": "number",
"description": "Number of processed IDs that resolved to a structure."
},
"notice": {
"description": "Every applicable advisory joined into one string: batch cap, partial failures, a computed model whose provider coordinate lookup failed (only its BinaryCIF URL listed), coordinate-budget overflow, and failed coordinate inlining.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `mixed_id_types`: The batch mixes PDB IDs and UniProt accessions under a single source that cannot serve both. `all_failed`: No requested ID resolved to a structure. Other values are possible when a failure originates below the handler.",
"examples": [
"mixed_id_types",
"all_failed"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"structures",
"failed",
"attribution",
"requested",
"processed",
"resolved"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_find_similar(by, sequence, pdb_id, uniprot, ticket_id, ...)
Find structurally or evolutionarily related proteins. by:"sequence" runs an RCSB mmseqs2 sequence-similarity search (synchronous) over a sequence — supplied directly, or pulled from a PDB ID or UniProt accession. by:"structure" runs a Foldseek fold-similarity search (asynchronous) against experimental and predicted databases; if the job is still computing when the poll budget elapses, the response reports status "computing" with a ticket — re-call with ticket_id set to that value to resume the same job instead of resubmitting, and a complete response carries the same ticket so a finished job can be re-paged with a different start. Foldseek searches each chain of a multichain structure as its own query: a response covers one query (query, default 0), reports queryCount, and ranks that query's hits by score across all searched databases. Each mode reads only its own controls: sequence, max_evalue and min_identity belong to by:"sequence"; ticket_id, databases and query belong to by:"structure"; pdb_id, uniprot, start and limit are shared. A field the selected mode cannot consume is rejected rather than silently ignored. Output names the engine and database each hit came from.
입력 스키마
{
"type": "object",
"properties": {
"by": {
"type": "string",
"enum": [
"sequence",
"structure"
],
"description": "Similarity axis: sequence (mmseqs2) or structure (Foldseek)."
},
"sequence": {
"description": "One-letter amino-acid sequence to search from (by:sequence).",
"type": "string"
},
"pdb_id": {
"description": "PDB entry ID to derive the query from.",
"type": "string"
},
"uniprot": {
"description": "UniProt accession to derive the query from.",
"type": "string"
},
"ticket_id": {
"description": "Foldseek ticket ID from a prior by:structure response whose status was \"computing\". When set, polls that existing job instead of submitting a new search (by:structure only) — pdb_id/uniprot/databases are ignored.",
"type": "string"
},
"databases": {
"description": "Foldseek target databases (by:structure). Default pdb100 + afdb50. e.g. afdb-swissprot, BFVD.",
"type": "array",
"items": {
"type": "string"
}
},
"query": {
"description": "Zero-based Foldseek query index (by:structure). Default 0. Foldseek splits a multichain structure into one query per chain, in file order; queryCount in the response says how many the job holds. Pass the same value with ticket_id to resume or re-page that query.",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
},
"max_evalue": {
"description": "Maximum E-value (by:sequence). Default 1.",
"type": "number",
"exclusiveMinimum": 0
},
"min_identity": {
"description": "Minimum sequence identity 0–1 (by:sequence). Default 0.",
"type": "number",
"minimum": 0,
"maximum": 1
},
"limit": {
"default": 25,
"description": "Maximum hits to return (1–100).",
"type": "integer",
"minimum": 1,
"maximum": 100
},
"start": {
"default": 0,
"description": "Zero-based result offset. Pages a by:sequence search, and a completed by:structure job when re-called with its ticket_id.",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"by"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"by": {
"type": "string",
"enum": [
"sequence",
"structure"
],
"description": "Echoed similarity axis."
},
"engine": {
"type": "string",
"description": "The engine that answered (e.g. \"RCSB mmseqs2\", \"Foldseek\")."
},
"status": {
"type": "string",
"enum": [
"complete",
"computing"
],
"description": "complete with hits, or computing (async — re-call with ticket_id to resume)."
},
"ticketId": {
"description": "Async job ticket ID (by:structure). Present while status is computing, and also on a complete response — re-call with ticket_id plus a different start/limit to page the same finished job instead of resubmitting.",
"type": "string"
},
"query": {
"description": "Zero-based Foldseek query index the hits belong to (by:structure).",
"type": "number"
},
"queryCount": {
"description": "Number of queries the Foldseek job holds, one per chain of the submitted structure (by:structure, complete). Valid query values run from 0 to queryCount - 1.",
"type": "number"
},
"hits": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "Hit identifier that chains directly into protein_get_structure: a bare PDB entry ID (e.g. 1A00), or a UniProt accession for predicted hits."
},
"entityId": {
"description": "Matched polymer-entity ID (e.g. 1A00_1) for sequence hits; the chainable entry ID is in `id`.",
"type": "string"
},
"source": {
"type": "string",
"enum": [
"experimental",
"predicted"
],
"description": "Whether the hit is an experimental or predicted structure."
},
"score": {
"description": "Relevance / alignment score.",
"type": "number"
},
"evalue": {
"description": "Alignment E-value (structure hits).",
"type": "number"
},
"identity": {
"description": "Sequence identity 0–1 over the alignment (structure hits).",
"type": "number"
},
"database": {
"description": "Source database the hit came from (structure hits).",
"type": "string"
},
"title": {
"description": "Structure title (enriched sequence hits).",
"type": "string"
},
"organism": {
"description": "Source organism (enriched sequence hits).",
"type": "string"
},
"uniprotAccession": {
"description": "UniProt accession (predicted structure hits).",
"type": "string"
}
},
"required": [
"id",
"source"
],
"additionalProperties": false,
"description": "A similar protein, with alignment scores when available."
},
"description": "Similar proteins, best first. Structure hits are ranked by score across all searched databases (hits without a score last, ties by database then target)."
},
"totalCount": {
"description": "Total upstream matches before pagination.",
"type": "number"
},
"start": {
"description": "Zero-based offset of the returned result page.",
"type": "number"
},
"nextStart": {
"description": "Offset for the next result page; absent on the final or past-end page.",
"type": "number"
},
"notice": {
"description": "Advisory note: a structure job still computing (with how to resume it), no matches, a start past the end of the results, or other queries in a multichain structure job.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `missing_query`: None of sequence, pdb_id, or uniprot was provided. `mode_mismatched_field`: An input field only the other by mode consumes was supplied: sequence, max_evalue or min_identity under by:\"structure\"; ticket_id, query or a non-empty databases list under by:\"sequence\". `no_sequence`: A sequence could not be resolved from the given PDB ID or UniProt accession. `query_out_of_range`: The requested query index is not below the number of queries the completed Foldseek job holds. `search_failed`: The Foldseek search service rejected or failed the structure job. `ticket_not_found`: The supplied ticket_id was rejected by Foldseek as an invalid or expired ticket. Other values are possible when a failure originates below the handler.",
"examples": [
"missing_query",
"mode_mismatched_field",
"no_sequence",
"query_out_of_range",
"search_failed",
"ticket_not_found"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"by",
"engine",
"status",
"hits"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_track_ligands(mode, query, comp_id, pdb_id, limit, ...)
Ligand discovery and binding-site analysis across the PDB. mode "find_ligand" resolves a name or formula to chemical component IDs with metadata (formula, weight, SMILES), ranked by deposition frequency — most-deposited component first, so the top hit is the most common match for the name, not necessarily an exact name-string match. The ranking covers a bounded candidate pool; totalCount and candidatesConsidered report how many components matched and how many were ranked. mode "structures_with_ligand" returns PDB entries containing a ligand (by exact component ID — get the ID from find_ligand first), highest-resolution first, each with its resolution in Å. mode "binding_site" returns the protein residues lining a ligand's pocket in a given structure, with contact distances, each numbered in both the mmCIF label namespace (asymId, seqId) and the author namespace (authAsymId, authSeqId) used by deposited coordinates and most literature. Binding sites are experimental-only (computed from deposited coordinates; predicted models carry no bound ligands).
입력 스키마
{
"type": "object",
"properties": {
"mode": {
"type": "string",
"enum": [
"find_ligand",
"structures_with_ligand",
"binding_site"
],
"description": "Operation: resolve a ligand, find structures containing it, or analyze its binding site."
},
"query": {
"description": "Ligand name or formula (mode find_ligand).",
"type": "string"
},
"comp_id": {
"description": "Exact chemical component ID (modes structures_with_ligand and binding_site).",
"type": "string"
},
"pdb_id": {
"description": "PDB entry ID (mode binding_site).",
"type": "string"
},
"limit": {
"default": 25,
"description": "Maximum results to return (1–100).",
"type": "integer",
"minimum": 1,
"maximum": 100
},
"start": {
"default": 0,
"description": "Zero-based result offset (modes structures_with_ligand and binding_site; binding_site pages ligand instances).",
"type": "integer",
"minimum": 0,
"maximum": 9007199254740991
}
},
"required": [
"mode"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"mode": {
"type": "string",
"enum": [
"find_ligand",
"structures_with_ligand",
"binding_site"
],
"description": "Echoed mode."
},
"ligands": {
"description": "Resolved chemical components (find_ligand).",
"type": "array",
"items": {
"type": "object",
"properties": {
"compId": {
"type": "string",
"description": "Chemical component ID (e.g. STI, HEM)."
},
"name": {
"description": "Chemical name.",
"type": "string"
},
"formula": {
"description": "Molecular formula.",
"type": "string"
},
"formulaWeight": {
"description": "Formula weight in Da.",
"type": "number"
},
"smiles": {
"description": "Isomeric SMILES.",
"type": "string"
},
"inchikey": {
"description": "InChIKey.",
"type": "string"
},
"type": {
"description": "Component type (e.g. non-polymer).",
"type": "string"
},
"depositionCount": {
"type": "number",
"description": "Number of PDB entries containing this component (deposition frequency). Candidates are ranked by this value, most-deposited first."
}
},
"required": [
"compId",
"depositionCount"
],
"additionalProperties": false,
"description": "A resolved chemical component (ligand) and its identifiers."
}
},
"structures": {
"description": "PDB entries containing the ligand, highest-resolution first (structures_with_ligand).",
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "PDB entry ID containing the ligand."
},
"resolution": {
"description": "Best reported resolution in Å (absent for methods without one, e.g. NMR).",
"type": "number"
}
},
"required": [
"id"
],
"additionalProperties": false,
"description": "A PDB entry containing the ligand, with its resolution."
}
},
"bindingSites": {
"description": "Binding-site residues (binding_site).",
"type": "array",
"items": {
"type": "object",
"properties": {
"ligandCompId": {
"type": "string",
"description": "Bound ligand chemical component ID."
},
"ligandAsymId": {
"description": "Author chain ID (auth_asym_id) of the ligand instance — the author namespace, unlike residues[].asymId.",
"type": "string"
},
"ligandAuthSeqId": {
"description": "Author residue number (auth_seq_id) of the ligand instance.",
"type": "number"
},
"residues": {
"type": "array",
"items": {
"type": "object",
"properties": {
"residueCompId": {
"type": "string",
"description": "Interacting residue type (e.g. ASP)."
},
"asymId": {
"type": "string",
"description": "mmCIF label_asym_id of the residue's chain (label namespace, paired with seqId). Can differ from authAsymId."
},
"seqId": {
"description": "mmCIF label_seq_id: position in the entity sequence (label namespace). Not the deposited residue number; see authSeqId.",
"type": "number"
},
"authAsymId": {
"description": "Author chain ID (auth_asym_id) of the residue's chain, as deposited and as most structure viewers label it.",
"type": "string"
},
"authSeqId": {
"description": "Author residue number (auth_seq_id), as deposited and as most literature and viewers number it. Related to seqId by no fixed offset; insertion codes are not reported.",
"type": "number"
},
"distance": {
"description": "Contact distance to the ligand in Å.",
"type": "number"
}
},
"required": [
"residueCompId",
"asymId"
],
"additionalProperties": false,
"description": "A pocket residue in contact with the ligand, in both numbering namespaces."
},
"description": "Protein residues lining the pocket, nearest first."
}
},
"required": [
"ligandCompId",
"residues"
],
"additionalProperties": false,
"description": "A ligand instance and the protein residues lining its pocket."
}
},
"totalCount": {
"description": "Total upstream matches before any local narrowing: PDB entries containing the component (structures_with_ligand), chemical components matching the query (find_ligand), or ligand instances with pocket contacts (binding_site).",
"type": "number"
},
"candidatesConsidered": {
"description": "Chemical components pulled into the deposition-frequency ranking (find_ligand). Below totalCount when the candidate pool was truncated; the ranking then covers only these candidates.",
"type": "number"
},
"start": {
"description": "Zero-based offset of the structures_with_ligand or binding_site result page.",
"type": "number"
},
"nextStart": {
"description": "Offset for the next structures_with_ligand or binding_site page; absent on the final or past-end page.",
"type": "number"
},
"resolvedCompId": {
"description": "The chemical component ID used to query.",
"type": "string"
},
"notice": {
"description": "Advisory note: a find_ligand candidate pool smaller than the upstream match count, a structures_with_ligand page that is empty because no entry contains the component or because start is past the end, or a binding_site page that leaves instances for a later start or starts past the last instance.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `missing_param`: The mode-specific required input is absent: query (find_ligand), comp_id (structures_with_ligand), or pdb_id (binding_site). `not_found`: No chemical component matched the name/formula, or the structure has no instance of the ligand. Other values are possible when a failure originates below the handler.",
"examples": [
"missing_param",
"not_found"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"mode"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_compare_structures(structures, reference, method, timeout_s, resume)
Structurally align multiple structures (up to the configured batch cap) via the RCSB Structural Comparison service (TM-align / jFATCAT). reference:"first" aligns every structure to the first; reference:"all_pairs" computes the full pairwise matrix. Each pair is an independent async alignment job with per-pair partial success — a pair still computing when the budget elapses returns status "computing" with its job UUID, and a failed pair degrades its row without sinking the others. Re-call with a matching entry in resume[] to poll a computing pair's UUID instead of resubmitting; a resumed pair reports a and b in the order its job was submitted, and a resume under a different method is rejected. Returns TM-score, RMSD, and aligned-residue count per pair, plus each structure's modeled-residue count and alignment coverage. TM-score is length-normalized and can shift sharply between structures that differ only by a terminal residue or two — the greedy superposition can settle into a worse local optimum — so read tmScore alongside rmsd, alignedResidues, modeledResidues and coverage, the columns that make such cases diagnosable.
입력 스키마
{
"type": "object",
"properties": {
"structures": {
"minItems": 2,
"maxItems": 25,
"type": "array",
"items": {
"type": "object",
"properties": {
"pdb_id": {
"type": "string",
"minLength": 1,
"description": "PDB entry ID."
},
"chain": {
"description": "mmCIF label_asym_id restricting the alignment to a single chain. Read it from polymerEntities[].labelAsymIds on protein_get_structure (source experimental) or the pdb://{entry_id} resource. Author chain IDs are a different namespace — polymerEntities[].authAsymIds, what protein_get_annotations.chain takes — and are not interchangeable with this one. Case-sensitive.",
"type": "string"
}
},
"required": [
"pdb_id"
],
"description": "A structure to align, by PDB entry ID with optional chain."
},
"description": "The structures to compare, up to the configured batch cap (excess is dropped with a notice). A structure repeated here is compared once."
},
"reference": {
"default": "first",
"description": "Align all to the first structure, or compute the full pairwise matrix.",
"type": "string",
"enum": [
"first",
"all_pairs"
]
},
"method": {
"default": "tm-align",
"description": "Alignment algorithm: tm-align, fatcat-rigid, or fatcat-flexible.",
"type": "string",
"enum": [
"tm-align",
"fatcat-rigid",
"fatcat-flexible"
]
},
"timeout_s": {
"description": "Poll budget per pair in seconds before returning \"computing\". Defaults to the server setting.",
"type": "integer",
"minimum": 5,
"maximum": 120
},
"resume": {
"description": "Resume tickets from a prior call: for each pair whose labels match an entry here, poll the existing UUID instead of submitting a new alignment job. Copy a, b, and uuid verbatim from a prior response's pairs[]; keep structures, reference, and method unchanged. The order of structures may change — a resumed pair keeps the orientation its job was submitted in.",
"type": "array",
"items": {
"type": "object",
"properties": {
"a": {
"type": "string",
"minLength": 1,
"description": "First structure label (entry or entry.chain) of a pair from a prior response."
},
"b": {
"type": "string",
"minLength": 1,
"description": "Second structure label (entry or entry.chain) of a pair from a prior response."
},
"uuid": {
"type": "string",
"minLength": 1,
"description": "Alignment job UUID returned for that pair by a prior call."
}
},
"required": [
"a",
"b",
"uuid"
],
"description": "A prior pair to resume by UUID instead of resubmitting."
}
}
},
"required": [
"structures"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"method": {
"type": "string",
"description": "Alignment method used."
},
"reference": {
"type": "string",
"enum": [
"first",
"all_pairs"
],
"description": "Comparison mode used."
},
"pairs": {
"type": "array",
"items": {
"type": "object",
"properties": {
"a": {
"type": "string",
"description": "First structure of the pair (entry[.chain]), as the alignment job was submitted — for a resumed pair that can differ from this call's structures[] order."
},
"b": {
"type": "string",
"description": "Second structure of the pair (entry[.chain])."
},
"status": {
"type": "string",
"enum": [
"complete",
"computing",
"failed"
],
"description": "Outcome for this pair."
},
"tmScore": {
"description": "TM-score (0–1; higher is more similar). Length-normalized by structure a's modeled length, so the same pair aligned b-first can score very differently (0.17 vs 0.40 for a 141- vs a 46-residue chain). It can also be sensitive to terminal length differences between the two structures — a one-residue overhang can flip the greedy superposition into a worse local optimum, dropping the score sharply. Cross-check rmsd and alignedResidues to spot such cases.",
"type": "number"
},
"rmsd": {
"description": "RMSD in Å over aligned residues.",
"type": "number"
},
"alignedResidues": {
"description": "Number of aligned residue pairs.",
"type": "number"
},
"modeledResidues": {
"description": "Modeled residue count per structure, ordered [a, b] to match this pair. A large gap between the two is the terminal-length asymmetry that can depress tmScore.",
"minItems": 2,
"maxItems": 2,
"type": "array",
"items": {
"type": "number"
}
},
"coverage": {
"description": "Alignment coverage per structure as a 0–100 percentage of that structure's own modeled-residue count — not of the full sequence and not of the shorter structure — ordered [a, b] to match this pair. Read alongside modeledResidues: equal aligned counts give the shorter structure the higher coverage.",
"minItems": 2,
"maxItems": 2,
"type": "array",
"items": {
"type": "number"
}
},
"uuid": {
"description": "Alignment job UUID (present for computing/complete pairs).",
"type": "string"
},
"error": {
"description": "Failure detail (failed pairs).",
"type": "string"
}
},
"required": [
"a",
"b",
"status"
],
"additionalProperties": false,
"description": "Alignment outcome for one structure pair."
},
"description": "One row per aligned pair."
},
"pairsTotal": {
"type": "number",
"description": "Number of pairs compared."
},
"computing": {
"type": "number",
"description": "Number of pairs still computing."
},
"notice": {
"description": "Advisory note: pairs still computing or failed (with how to resume them), structures beyond the batch cap that were ignored, and repeated structures compared once.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `resume_pair_unmatched`: A resume entry's a/b labels don't match any pair generated from structures + reference. `no_distinct_pair`: Every entry in structures[] denotes the same structure, leaving no pair to align. `resume_method_mismatch`: A resumed alignment job completed under a different method than this call's method input. `resume_job_mismatch`: A resume entry's uuid belongs to an alignment job for a different structure pair than its a/b labels. Other values are possible when a failure originates below the handler.",
"examples": [
"resume_pair_unmatched",
"no_distinct_pair",
"resume_method_mismatch",
"resume_job_mismatch"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"method",
"reference",
"pairs",
"pairsTotal",
"computing"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_analyze_collection(group_by, query, organism, method, max_resolution, ...)
Profile the PDB into distributions and trends over an optional scoping query: counts by method, organism, or polymer composition; resolution and molecular-weight histograms; release-year timelines; and multidimensional cross-tabs (e.g. method × release_year). Aggregation runs at RCSB, so the response carries compact counts per bucket rather than the matching entries. Pass one group_by dimension for a single breakdown, or two distinct dimensions for a cross-tab (the first nests the second). bucket_limit caps each dimension level separately rather than the response, so a cross-tab returns up to that many nested buckets under each of its capped parent buckets; bucketsReturned reports the realized total.
입력 스키마
{
"type": "object",
"properties": {
"group_by": {
"minItems": 1,
"maxItems": 2,
"type": "array",
"items": {
"type": "string",
"enum": [
"method",
"organism",
"polymer_type",
"resolution",
"release_year",
"molecular_weight"
]
},
"description": "1 dimension for a breakdown, or 2 distinct dimensions for a cross-tab (the first nests the second). Repeating a dimension is rejected."
},
"query": {
"description": "Optional free-text scope (e.g. \"kinase\"); omit to profile the whole PDB.",
"type": "string"
},
"organism": {
"description": "Optional source-organism scope.",
"type": "string"
},
"method": {
"description": "Optional experimental-method scope.",
"type": "string"
},
"max_resolution": {
"description": "Optional maximum-resolution scope (Å).",
"type": "number",
"exclusiveMinimum": 0
},
"content_type": {
"default": "experimental",
"description": "Which structure universe to profile. Default experimental. Computed models carry no experimental metadata, so method and resolution return nothing under \"predicted\".",
"type": "string",
"enum": [
"experimental",
"predicted",
"all"
]
},
"interval": {
"description": "Bin width for a histogram dimension (a number, for resolution or molecular_weight) or period for a date histogram (\"year\", for release_year). Applies to whichever requested group_by dimension can consume that value type — primary or nested child — so a cross-tab like [\"method\",\"resolution\"] bins its nested resolution child. When both requested dimensions can consume it the primary takes it and the child keeps its default. When neither can, the call is rejected rather than silently ignoring the override.",
"anyOf": [
{
"type": "number",
"exclusiveMinimum": 0,
"description": "Numeric bin width for a value histogram (e.g. resolution Å)."
},
{
"type": "string",
"enum": [
"year"
],
"description": "Period granularity for a date histogram. Only \"year\"."
}
]
},
"bucket_limit": {
"description": "Max buckets per dimension level, not per response. A cross-tab applies the cap separately to the parent dimension and to the nested child inside each parent bucket, so up to bucket_limit × (1 + bucket_limit) buckets can come back — 2550 at the default 50. The realized count comes back as bucketsReturned. Defaults to the configured server cap.",
"type": "integer",
"minimum": 1,
"maximum": 500
}
},
"required": [
"group_by"
],
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"total": {
"type": "number",
"description": "Total entries in the scoped collection."
},
"facets": {
"type": "array",
"items": {
"type": "object",
"properties": {
"dimension": {
"type": "string",
"description": "Friendly dimension name (e.g. method, organism, release_year)."
},
"buckets": {
"type": "array",
"items": {
"type": "object",
"properties": {
"label": {
"type": "string",
"description": "Bucket value — category, numeric bin start, or period."
},
"count": {
"type": "number",
"description": "Number of entries in the bucket."
},
"rangeFrom": {
"description": "Inclusive lower bound of a numeric histogram bin (= label). Present only for numeric facets (resolution, molecular_weight); absent for term/period facets.",
"type": "number"
},
"rangeTo": {
"description": "Exclusive upper bound of a numeric histogram bin (rangeFrom + bin interval), so the bin covers [rangeFrom, rangeTo). Present only for numeric facets.",
"type": "number"
},
"child": {
"description": "Nested dimension breakdown for a cross-tab. Present whenever a second dimension was requested — with an empty bucket list when nothing in this bucket carries a value for it — so its absence means no cross-tab was asked for. At most one: a bucket is never cross-tabbed by more than one dimension.",
"type": "object",
"properties": {
"dimension": {
"type": "string",
"description": "Nested dimension name."
},
"buckets": {
"type": "array",
"items": {
"type": "object",
"properties": {
"label": {
"type": "string",
"description": "Bucket value — category, numeric bin start, or period."
},
"count": {
"type": "number",
"description": "Number of entries in the bucket."
},
"rangeFrom": {
"description": "Inclusive lower bound of a numeric histogram bin (= label). Present only for numeric facets (resolution, molecular_weight); absent for term/period facets.",
"type": "number"
},
"rangeTo": {
"description": "Exclusive upper bound of a numeric histogram bin (rangeFrom + bin interval), so the bin covers [rangeFrom, rangeTo). Present only for numeric facets.",
"type": "number"
}
},
"required": [
"label",
"count"
],
"additionalProperties": false,
"description": "A leaf aggregation bucket: a value and its entry count."
},
"description": "Nested buckets within the parent bucket."
},
"truncated": {
"description": "True when this nested bucket list was capped by the per-dimension bucket limit.",
"type": "boolean"
},
"missingValueCount": {
"type": "number",
"description": "Matches inside the parent bucket that carry no value for this nested attribute, so they fall in no nested bucket. 0 when the nested buckets account for the whole parent bucket."
}
},
"required": [
"dimension",
"buckets",
"missingValueCount"
],
"additionalProperties": false
}
},
"required": [
"label",
"count"
],
"additionalProperties": false,
"description": "A top-level aggregation bucket, optionally cross-tabbed by a nested dimension."
},
"description": "Aggregation buckets, count-descending for terms."
},
"truncated": {
"description": "True when buckets were capped by the per-dimension bucket limit.",
"type": "boolean"
},
"missingValueCount": {
"type": "number",
"description": "Matches in scope that carry no value for this attribute, so they count toward the response total but fall in no bucket (e.g. computed models have no experimental method; solution-NMR entries have no diffraction resolution). Independent of truncation — measured before the bucket cap is applied. 0 means no shortfall was detectable: a multi-valued attribute such as organism can place one match in several buckets, which offsets the shortfall rather than adding to it."
}
},
"required": [
"dimension",
"buckets",
"missingValueCount"
],
"additionalProperties": false,
"description": "A facet dimension and its aggregation buckets, each optionally cross-tabbed."
},
"description": "The requested breakdown(s)."
},
"scope": {
"description": "Echoed scope description for follow-up calls.",
"type": "string"
},
"notice": {
"description": "Advisory note (per-position bucket truncation, a scope that matched nothing, dimensions with no data under the requested content_type, dimensions whose buckets cover materially less than the total). Carries every applicable advisory in one string.",
"type": "string"
},
"truncated": {
"description": "True when at least one dimension position — the top-level dimension or a nested cross-tab child — had more buckets than the applied cap. Which positions, and by how much, is named in notice.",
"type": "boolean"
},
"bucketsReturned": {
"type": "number",
"description": "Buckets in this response, summed over every dimension level: the top-level buckets plus, for a cross-tab, the nested child buckets under each of them. Since bucket_limit caps each level separately, this is the size those caps actually produced — always present, cross-tab or not."
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `interval_not_applicable`: interval was supplied but neither requested group_by dimension aggregates by a histogram that accepts a value of that type. `duplicate_dimension`: group_by lists the same dimension twice, which would cross a dimension with itself. Other values are possible when a failure originates below the handler.",
"examples": [
"interval_not_applicable",
"duplicate_dimension"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"total",
"facets",
"bucketsReturned"
]
},
{
"required": [
"error"
]
}
]
}🟢protein_get_annotations(uniprot, pdb_id, chain, include, limit)
Sequence and functional annotation for a protein: UniProt features (domains, binding sites, PTMs), natural variants, and InterPro domain/family memberships (Pfam, PROSITE, …) with GO terms. Provide a UniProt accession directly, or a PDB ID — it is resolved to its UniProt accession via the structure's sequence cross-reference. A multi-chain PDB entry can map to several accessions; the default pick is deterministic (lowest author chain ID) and the alternatives are listed in "ambiguity" — pass "chain" to select a specific one. Use "include" to scope which annotation classes are fetched. Every response carries an "attribution" block with the upstream data licenses and citations.
입력 스키마
{
"type": "object",
"properties": {
"uniprot": {
"description": "UniProt accession (e.g. P69905). Takes precedence over pdb_id.",
"type": "string"
},
"pdb_id": {
"description": "PDB entry ID; resolved to a UniProt accession via cross-reference.",
"type": "string"
},
"chain": {
"description": "Author chain ID (auth_asym_id, e.g. \"A\") that disambiguates a multi-chain PDB entry to a specific UniProt accession. Case-sensitive — must match the author chain ID exactly (large structures can carry distinct \"A\" and \"a\" chains). Only applies with pdb_id; ignored when uniprot is supplied directly. See polymerEntities[].authAsymIds in the pdb://{entry_id} resource for an entry's author chain IDs.",
"type": "string"
},
"include": {
"default": "all",
"description": "Which annotation classes to fetch: features, domains (InterPro), variants, or all.",
"type": "string",
"enum": [
"features",
"domains",
"variants",
"all"
]
},
"limit": {
"default": 50,
"description": "Per-class cap (1–200): features, natural variants, and InterPro domains are each independently truncated to at most this many records. The default keeps a typical annotation set intact while bounding a densely-annotated protein (a well-studied protein can carry 150+ natural variants); a truncated class is disclosed in the response notice — raise it to retrieve more.",
"type": "integer",
"minimum": 1,
"maximum": 200
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false
}출력 스키마
{
"type": "object",
"properties": {
"accession": {
"type": "string",
"description": "UniProt accession the annotations describe."
},
"proteinName": {
"description": "Recommended protein name.",
"type": "string"
},
"geneNames": {
"type": "array",
"items": {
"type": "string"
},
"description": "Gene names."
},
"organism": {
"description": "Source organism scientific name.",
"type": "string"
},
"function": {
"description": "UniProt function summary.",
"type": "string"
},
"sequenceLength": {
"description": "Sequence length in residues.",
"type": "number"
},
"features": {
"description": "Structural/functional features.",
"type": "array",
"items": {
"type": "object",
"properties": {
"type": {
"type": "string",
"description": "Feature type (e.g. Domain, Binding site, Modified residue)."
},
"description": {
"description": "Feature description.",
"type": "string"
},
"start": {
"description": "Start residue (1-based).",
"type": "number"
},
"end": {
"description": "End residue (1-based).",
"type": "number"
}
},
"required": [
"type"
],
"additionalProperties": false,
"description": "A sequence feature or natural variant over a residue range."
}
},
"variants": {
"description": "Natural sequence variants.",
"type": "array",
"items": {
"type": "object",
"properties": {
"type": {
"type": "string",
"description": "Feature type (e.g. Domain, Binding site, Modified residue)."
},
"description": {
"description": "Feature description.",
"type": "string"
},
"start": {
"description": "Start residue (1-based).",
"type": "number"
},
"end": {
"description": "End residue (1-based).",
"type": "number"
}
},
"required": [
"type"
],
"additionalProperties": false,
"description": "A sequence feature or natural variant over a residue range."
}
},
"domains": {
"description": "InterPro domain/family memberships.",
"type": "array",
"items": {
"type": "object",
"properties": {
"accession": {
"type": "string",
"description": "InterPro accession (e.g. IPR000001)."
},
"name": {
"type": "string",
"description": "Domain/family name."
},
"type": {
"type": "string",
"description": "Entry type (e.g. domain, family, homologous_superfamily)."
},
"memberDatabases": {
"type": "array",
"items": {
"type": "string"
},
"description": "Contributing member databases (e.g. pfam, profile)."
},
"goTerms": {
"type": "array",
"items": {
"type": "object",
"properties": {
"id": {
"type": "string",
"description": "GO term ID (e.g. GO:0005515)."
},
"name": {
"type": "string",
"description": "GO term name."
},
"category": {
"description": "GO aspect (molecular_function, biological_process, …).",
"type": "string"
}
},
"required": [
"id",
"name"
],
"additionalProperties": false,
"description": "An associated Gene Ontology term."
},
"description": "Associated GO terms."
}
},
"required": [
"accession",
"name",
"type",
"memberDatabases",
"goTerms"
],
"additionalProperties": false,
"description": "An InterPro domain/family membership with GO terms."
}
},
"ambiguity": {
"description": "Alternative UniProt mappings when a PDB ID resolved ambiguously (no chain given).",
"type": "object",
"properties": {
"accessions": {
"type": "array",
"items": {
"type": "object",
"properties": {
"chain": {
"type": "array",
"items": {
"type": "string"
},
"description": "Author chain IDs (auth_asym_id) this entity covers (e.g. [\"A\", \"C\"])."
},
"accession": {
"type": "string",
"description": "UniProt accession the chains map to."
},
"proteinName": {
"description": "Polymer entity description.",
"type": "string"
}
},
"required": [
"chain",
"accession"
],
"additionalProperties": false,
"description": "One chain-group → UniProt accession mapping within the entry."
},
"description": "Every distinct UniProt mapping for the entry, lowest-chain first."
},
"notice": {
"type": "string",
"description": "Why multiple mappings exist and how to select one via the chain input."
}
},
"required": [
"accessions",
"notice"
],
"additionalProperties": false
},
"attribution": {
"type": "array",
"items": {
"type": "object",
"properties": {
"source": {
"type": "string",
"description": "Contributing data-source display name (e.g. \"RCSB PDB\", \"AlphaFold DB\", \"SWISS-MODEL\", \"UniProt\"). Open-ended — best_available structures are federated through 3D-Beacons providers."
},
"license": {
"type": "string",
"description": "License the source data is released under (e.g. \"CC BY 4.0\")."
},
"citation": {
"type": "string",
"description": "Primary-literature citation to credit the source."
},
"homepage": {
"type": "string",
"description": "Source homepage (absolute URL)."
}
},
"required": [
"source",
"license",
"citation",
"homepage"
],
"additionalProperties": false,
"description": "Upstream data-source attribution: license, citation, and homepage."
},
"description": "Upstream data-source licenses and citations for every source that contributed to this response. Always present — the attribution obligation travels with the data."
},
"resolvedFrom": {
"description": "PDB ID the accession was resolved from, when applicable.",
"type": "string"
},
"truncated": {
"description": "True when at least one annotation class hit the per-class limit and was capped; the notice names each capped class with its pre-cap count.",
"type": "boolean"
},
"notice": {
"description": "Advisory note: a per-class \"showing N of M\" line for each capped class and a no-data note for any requested-but-empty class.",
"type": "string"
},
"error": {
"description": "Present when the call failed. Absent on success.",
"type": "object",
"properties": {
"code": {
"type": "integer",
"minimum": -9007199254740991,
"maximum": 9007199254740991,
"description": "JSON-RPC error code for this failure."
},
"message": {
"type": "string",
"description": "Human-readable description of what went wrong."
},
"data": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "Machine-readable failure mode. Declared by this tool: `missing_identifier`: Neither uniprot nor pdb_id was supplied, so there is no protein to look up. `invalid_accession`: The supplied uniprot value is not a syntactically valid UniProt accession. `no_uniprot_mapping`: A supplied PDB ID resolved to no usable UniProt cross-reference (e.g. a nucleic-acid-only entry). `chain_not_found`: A supplied chain matches no UniProt-mapped polymer entity in the resolved PDB entry. Other values are possible when a failure originates below the handler.",
"examples": [
"missing_identifier",
"invalid_accession",
"no_uniprot_mapping",
"chain_not_found"
]
},
"recovery": {
"description": "Actionable next step for the caller.",
"type": "object",
"properties": {
"hint": {
"type": "string"
}
},
"required": [
"hint"
],
"additionalProperties": {}
},
"retryable": {
"description": "Whether retrying may succeed.",
"type": "boolean"
}
},
"additionalProperties": {}
}
},
"required": [
"code",
"message"
],
"additionalProperties": {}
}
},
"$schema": "https://json-schema.org/draft/2020-12/schema",
"additionalProperties": false,
"anyOf": [
{
"not": {
"required": [
"error"
]
},
"required": [
"accession",
"geneNames",
"attribution"
]
},
{
"required": [
"error"
]
}
]
}커뮤니티
증거