ensembl-mcp-server

Look up genes, sequences, variants, homologs, and cross-database xrefs from Ensembl REST.

Should I use this

Quality & Safety

A
Description quality
100%
Schema completeness
96%
Naming quality
80%
Poisoning risk
100%
Permission match
100%
Protocol compliance
100%

Based on automated analysis of tool definitions and protocol compliance.

Context Cost

~9,930Tokens (tool definitions)
~15.5 KBTypical response size
Significant attention impact (7.76% of 128k context)

This is the approximate number of tokens consumed each time the server's tools are loaded into a model's context. Higher counts reduce the attention available for other tasks.

Install

One-Click Install

Add this to your `claude_desktop_config.json` file:

{
  "mcpServers": {
    "ensembl-mcp-server": {
      "command": "bun",
      "args": [
        "@cyanheads/ensembl-mcp-server"
      ]
    }
  }
}

Runnable packages

npm@cyanheads/ensembl-mcp-server0.5.0streamable-http

Remote endpoints

https://ensembl.caseyjhand.com/mcpstreamable-http

What it can do

Tool inventory

Tools (7)

🟢 Read-only🟡 Write🔴 Delete⚪ Unknown
🟢ensembl_list_species(division, nameContains)

List species supported by Ensembl with display name, common name, assembly, taxon ID, and division. Required discovery step — species names like homo_sapiens are opaque to non-biologists and are the input format every other Ensembl tool expects. Filter by division to select one; use nameContains to find a species by partial name match. With no division, returns the endpoint default division — the vertebrates (~356 species on the default GRCh38 endpoint); pass a division to list that division.

Input Schema

{
  "type": "object",
  "properties": {
    "division": {
      "description": "Filter to a specific Ensembl division. EnsemblVertebrates includes human, mouse, zebrafish, and other vertebrates. EnsemblPlants covers crop and model plant genomes. EnsemblFungi, EnsemblMetazoa, EnsemblProtists cover non-vertebrate model organisms. Omit to return the endpoint default division (vertebrates).",
      "type": "string",
      "enum": [
        "EnsemblVertebrates",
        "EnsemblPlants",
        "EnsemblFungi",
        "EnsemblMetazoa",
        "EnsemblProtists"
      ]
    },
    "nameContains": {
      "description": "Case-insensitive substring filter applied locally after fetching. Matches against species name, display name, and common name. Example: \"sapiens\" matches homo_sapiens; \"mouse\" matches mus_musculus.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "species": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "name": {
            "type": "string",
            "description": "Ensembl internal species name in lowercase_underscore format (e.g. homo_sapiens, mus_musculus). This is the value to pass as the species parameter in all other Ensembl tools."
          },
          "displayName": {
            "description": "Human-readable scientific name (e.g. Homo sapiens).",
            "type": "string"
          },
          "commonName": {
            "description": "Common name (e.g. Human, Mouse).",
            "type": "string"
          },
          "taxonId": {
            "description": "NCBI taxonomy ID for this species.",
            "type": "string"
          },
          "assembly": {
            "description": "Current genome assembly name (e.g. GRCh38).",
            "type": "string"
          },
          "division": {
            "description": "Ensembl division this species belongs to (e.g. EnsemblVertebrates, EnsemblPlants).",
            "type": "string"
          }
        },
        "required": [
          "name"
        ],
        "additionalProperties": false,
        "description": "A single Ensembl species entry."
      },
      "description": "Species matching the filter criteria, sorted by internal name."
    },
    "totalCount": {
      "type": "number",
      "description": "Total number of matching species after local filtering."
    },
    "notice": {
      "description": "Guidance when the filter matches no species.",
      "type": "string"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode."
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "species",
        "totalCount"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_lookup_gene(symbol, id, species, ids, symbols, ...)

Resolve a gene by symbol + species (or by stable ID) to its Ensembl ID, genomic location (chr:start-end:strand), biotype, description, and transcript list. Entry point for most workflows — the stable ID and coordinates returned here are inputs to other tools. Accepts both symbol lookup (BRCA2 + homo_sapiens) and direct ID lookup (ENSG00000139618). Supports batch lookup of up to 20 IDs or symbols in one call via the ids or symbols field. Provide exactly one of symbol, id, ids, or symbols. For symbol lookups species defaults to homo_sapiens (override for other organisms); for ID lookups species is not needed. Use ensembl_list_species to discover valid species names.

Input Schema

{
  "type": "object",
  "properties": {
    "symbol": {
      "description": "Gene symbol to look up (e.g. BRCA2, TP53, EGFR). Species defaults to homo_sapiens; set species for other organisms. Case-insensitive in most species.",
      "type": "string"
    },
    "id": {
      "description": "Ensembl stable gene ID (e.g. ENSG00000139618 or ENSG00000139618.7 with version). Species is not required for ID lookup.",
      "type": "string"
    },
    "species": {
      "description": "Species in Ensembl internal format: lowercase scientific name with underscores (e.g. homo_sapiens, mus_musculus, danio_rerio). Optional for symbol lookups — defaults to homo_sapiens; set it for other organisms. Use ensembl_list_species to discover valid values.",
      "type": "string"
    },
    "ids": {
      "description": "Batch lookup: up to 20 Ensembl stable IDs (ENSG…, ENST…). Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols.",
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1,
        "description": "An Ensembl stable gene or transcript ID to resolve in this batch."
      }
    },
    "symbols": {
      "description": "Batch lookup: up to 20 gene symbols. Species defaults to homo_sapiens; set species for other organisms. Returns a succeeded/failed split. Provide exactly one of symbol, id, ids, or symbols.",
      "maxItems": 20,
      "type": "array",
      "items": {
        "type": "string",
        "minLength": 1,
        "description": "A gene symbol to resolve in this batch (e.g. BRCA2, TP53)."
      }
    },
    "expand_transcripts": {
      "default": false,
      "description": "When true, include the full transcript list in the response. Each transcript has its ID, biotype, canonical flag, and coordinates. Default is false to keep responses compact.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "gene": {
      "description": "Single gene record. Present for symbol or id lookups.",
      "type": "object",
      "properties": {
        "id": {
          "type": "string",
          "description": "Ensembl gene stable ID (ENSG…). Use this as input to other Ensembl tools."
        },
        "species": {
          "description": "Species in Ensembl internal format (e.g. homo_sapiens). Echoed from lookup.",
          "type": "string"
        },
        "displayName": {
          "description": "Gene symbol or display name (e.g. BRCA2, TP53).",
          "type": "string"
        },
        "description": {
          "description": "Brief gene description from Ensembl.",
          "type": "string"
        },
        "biotype": {
          "description": "Gene biotype (e.g. protein_coding, lncRNA, pseudogene).",
          "type": "string"
        },
        "chromosome": {
          "description": "Chromosome or sequence region name.",
          "type": "string"
        },
        "start": {
          "description": "Gene start position on the chromosome (1-based).",
          "type": "number"
        },
        "end": {
          "description": "Gene end position on the chromosome (1-based).",
          "type": "number"
        },
        "strand": {
          "description": "Strand: 1 for forward, -1 for reverse.",
          "type": "number"
        },
        "assemblyName": {
          "description": "Genome assembly name (e.g. GRCh38). All coordinates are relative to this assembly.",
          "type": "string"
        },
        "transcripts": {
          "description": "Transcript list. Present only when expand_transcripts is true.",
          "type": "array",
          "items": {
            "type": "object",
            "properties": {
              "id": {
                "type": "string",
                "description": "Ensembl transcript stable ID (ENST…)."
              },
              "displayName": {
                "description": "Transcript display name.",
                "type": "string"
              },
              "biotype": {
                "description": "Transcript biotype (e.g. protein_coding, lncRNA, retained_intron).",
                "type": "string"
              },
              "isCanonical": {
                "type": "boolean",
                "description": "True when this is the canonical transcript for the gene."
              },
              "start": {
                "description": "Transcript start position on the chromosome (1-based).",
                "type": "number"
              },
              "end": {
                "description": "Transcript end position on the chromosome (1-based).",
                "type": "number"
              },
              "strand": {
                "description": "Strand: 1 for forward, -1 for reverse.",
                "type": "number"
              },
              "lengthInBp": {
                "description": "Transcript length in base pairs.",
                "type": "number"
              }
            },
            "required": [
              "id",
              "isCanonical"
            ],
            "additionalProperties": false,
            "description": "A single transcript summary entry."
          }
        }
      },
      "required": [
        "id"
      ],
      "additionalProperties": false
    },
    "batch": {
      "description": "Batch results. Present for ids or symbols lookups.",
      "type": "object",
      "properties": {
        "succeeded": {
          "type": "array",
          "items": {
            "type": "object",
            "properties": {},
            "additionalProperties": {},
            "description": "A resolved gene record with the same shape as the gene output field."
          },
          "description": "Gene records for IDs/symbols that resolved successfully. Same shape as gene."
        },
        "failed": {
          "type": "array",
          "items": {
            "type": "object",
            "properties": {},
            "additionalProperties": {},
            "description": "A failed lookup entry with query (the submitted ID or symbol) and error (reason string) fields."
          },
          "description": "IDs/symbols that could not be resolved, with per-item query and error fields."
        }
      },
      "required": [
        "succeeded",
        "failed"
      ],
      "additionalProperties": false
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The gene symbol or stable ID was not found in Ensembl. `invalid_species`: The species string was not recognized by Ensembl. `no_input`: Neither symbol, id, ids, nor symbols was provided. `conflicting_input`: More than one of symbol, id, ids, or symbols was provided. Other values are possible when a failure originates below the handler.",
              "examples": [
                "not_found",
                "invalid_species",
                "no_input",
                "conflicting_input"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      }
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_get_sequence(id, type, species, expand_5prime, expand_3prime, ...)

Fetch the DNA, cDNA, CDS, or protein sequence for a gene, transcript, protein, or genomic region. Returns a window of the sequence — the first 10,000 characters by default — with its stable ID, molecule type, and full length. When more follows the window, truncated is true and nextOffset is the offset to request next; walking nextOffset reconstructs the whole sequence, and max_length 0 returns everything from offset to the end. The type parameter selects which sequence is fetched: genomic (default, includes introns), cdna (spliced transcript), cds (coding sequence only), protein. For region mode, set id to a region — either species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end with species set (e.g. id 13:32315086-32400268, species homo_sapiens), spanning at most 10,000,000 bases; regions return genomic DNA only. Protein sequences require a transcript or protein stable ID (ENST…/ENSP…), not a gene ID — use ensembl_lookup_gene with expand_transcripts=true to get the canonical transcript ID first.

Input Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "description": "Ensembl stable ID (ENSG…, ENST…, ENSP…) or a genomic region for region mode. Region accepts species:chr:start-end (e.g. homo_sapiens:13:32315086-32400268) or a bare chr:start-end (e.g. 13:32315086-32400268) when the species field is set. A region needs start at or below end, within the sequence region, and spans at most 10,000,000 bases."
    },
    "type": {
      "default": "genomic",
      "description": "Sequence type to retrieve. genomic: full genomic DNA including introns (default). cdna: spliced transcript sequence (requires ENST… ID). cds: coding sequence only, no UTRs (requires ENST… ID with coding transcript). protein: amino acid sequence (requires ENST… or ENSP… ID). Region ids are genomic-only — request cdna, cds, or protein from a transcript or protein stable ID.",
      "type": "string",
      "enum": [
        "genomic",
        "cdna",
        "cds",
        "protein"
      ]
    },
    "species": {
      "description": "Species in Ensembl internal format (e.g. homo_sapiens). Required for a bare chr:start-end region; optional for the species:chr:start-end form (the embedded species is used when the field is omitted). Optional for stable ID lookups — Ensembl infers species from the ID prefix.",
      "type": "string"
    },
    "expand_5prime": {
      "default": 0,
      "description": "Number of base pairs to extend upstream (5' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "expand_3prime": {
      "default": 0,
      "description": "Number of base pairs to extend downstream (3' direction) of the requested feature. Default 0. Only applies to genomic sequences and region queries.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "offset": {
      "default": 0,
      "description": "0-based character offset where the returned window starts, counted in the resolved sequence (including any expand_5prime/expand_3prime flank). Default 0. Pass nextOffset from a truncated response to fetch the following window; an offset at or past the end returns an empty window.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "max_length": {
      "default": 10000,
      "description": "Maximum number of characters in the returned window. Default 10000. Set to 0 to return everything from offset to the end, uncapped.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "description": "The stable ID or region used for the lookup."
    },
    "type": {
      "type": "string",
      "description": "Sequence type returned (genomic, cdna, cds, or protein)."
    },
    "seq": {
      "type": "string",
      "description": "The requested window of the sequence: at most max_length characters starting at offset. DNA sequences use IUPAC nucleotide codes (ACGT + ambiguity codes); protein sequences use single-letter amino acid codes. Empty when offset is at or past the end."
    },
    "length": {
      "type": "number",
      "description": "Full sequence length in characters, not the window size — nucleotides for genomic/cdna/cds, amino-acid residues for protein. Includes any expand_5prime/expand_3prime flank."
    },
    "offset": {
      "type": "number",
      "description": "0-based character offset where this window starts."
    },
    "truncated": {
      "type": "boolean",
      "description": "True when more sequence follows this window; request nextOffset to continue."
    },
    "nextOffset": {
      "description": "Offset of the first character after this window — pass it as offset to fetch the next window. Present only when truncated.",
      "type": "number"
    },
    "description": {
      "description": "Sequence description from Ensembl, if provided.",
      "type": "string"
    },
    "notice": {
      "description": "Guidance about the window: how to continue when truncated, or why it is empty when the offset is past the end.",
      "type": "string"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The stable ID or region was not found in Ensembl. `type_mismatch`: A non-genomic type (cdna, cds, or protein) was requested for a region id or a gene ID. `missing_species`: A bare chr:start-end region was given without a species. `invalid_region`: A region id has its start after its end, starts past the end of its sequence region, spans more than the 10,000,000-base maximum, or names a sequence region the species lacks. Other values are possible when a failure originates below the handler.",
              "examples": [
                "not_found",
                "type_mismatch",
                "missing_species",
                "invalid_region"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "id",
        "type",
        "seq",
        "length",
        "offset",
        "truncated"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_query_region(species, region, feature, biotype, max_results)

Find genomic features overlapping a chromosomal region: genes, transcripts, variants, regulatory elements, or exons. Returns each feature with its stable ID, type, location, biotype, and name, plus the genome assembly the coordinates are on. Useful for "what's in this locus?" and for seeding follow-up lookups. Region format is chr:start-end (e.g. 13:32315086-32400268 for the BRCA2 locus), spanning at most 5,000,000 bases. Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. The feature parameter defaults to gene only — requesting variation in an 85 kb region matches 44,000+ entries. Explicitly include variation, regulatory, transcript, or exon only when needed. The response returns up to max_results features (default 100) while totalCount always reports the full count; set max_results to 0 for every feature, or query a smaller region to see a different slice. Exon rows carry the parent transcript ID, so the same exon appears once per transcript it belongs to.

Input Schema

{
  "type": "object",
  "properties": {
    "species": {
      "type": "string",
      "minLength": 1,
      "description": "Species in Ensembl internal format (e.g. homo_sapiens, mus_musculus). Use ensembl_list_species to discover valid values."
    },
    "region": {
      "type": "string",
      "minLength": 1,
      "description": "Genomic region in chr:start-end format (e.g. 13:32315086-32400268). Ensembl serves at most 5,000,000 bases per region; split a larger area into smaller windows. Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (13, not chr13); a chr-prefixed name like chr13 is also accepted. For large regions (>100 kb), limit to gene feature type to avoid overwhelming results."
    },
    "feature": {
      "default": [
        "gene"
      ],
      "description": "Feature types to retrieve — at least one. Default is gene only. Requesting variation in a large region can match tens of thousands of features. Include variation only for targeted small regions (single gene loci or smaller).",
      "minItems": 1,
      "type": "array",
      "items": {
        "type": "string",
        "enum": [
          "gene",
          "transcript",
          "variation",
          "regulatory",
          "exon"
        ],
        "description": "A feature type to retrieve: gene, transcript, variation, regulatory, or exon."
      }
    },
    "biotype": {
      "description": "Optional biotype filter (e.g. protein_coding, lncRNA, SNV). Applied server-side by Ensembl. Not all feature types support biotype filtering.",
      "type": "string"
    },
    "max_results": {
      "default": 100,
      "description": "Maximum number of features to return. A gene-length region can hold tens of thousands of variation features; the default keeps the response compact. Set to 0 to return every feature uncapped. totalCount always reports the true number found before this cap.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "species",
    "region"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "features": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "id": {
            "description": "Ensembl stable ID for this feature (e.g. ENSG…, rs…).",
            "type": "string"
          },
          "name": {
            "description": "External name or symbol for this feature.",
            "type": "string"
          },
          "featureType": {
            "type": "string",
            "description": "Feature type: gene, transcript, variation, regulatory, or exon."
          },
          "biotype": {
            "description": "Biotype of the feature (e.g. protein_coding, lncRNA, SNV).",
            "type": "string"
          },
          "chromosome": {
            "type": "string",
            "description": "Chromosome or sequence region name."
          },
          "start": {
            "type": "number",
            "description": "Start position on the chromosome (1-based)."
          },
          "end": {
            "type": "number",
            "description": "End position on the chromosome (1-based)."
          },
          "strand": {
            "description": "Strand: 1 for forward, -1 for reverse.",
            "type": "number"
          },
          "description": {
            "description": "Feature description when provided.",
            "type": "string"
          },
          "consequenceType": {
            "description": "Most severe consequence type for variation features.",
            "type": "string"
          },
          "clinicalSignificance": {
            "description": "Clinical significance terms for variation features (e.g. pathogenic, benign).",
            "type": "array",
            "items": {
              "type": "string",
              "description": "A clinical significance term for this variant."
            }
          },
          "parentId": {
            "description": "Parent transcript ID (ENST…) for exon features. An exon is reported once per parent transcript it belongs to, so the same exon ID can appear on multiple rows that differ only by this field — not duplicates.",
            "type": "string"
          },
          "rank": {
            "description": "Position (1-based) of an exon within its parent transcript.",
            "type": "number"
          }
        },
        "required": [
          "featureType",
          "chromosome",
          "start",
          "end"
        ],
        "additionalProperties": false,
        "description": "A single genomic feature overlapping the queried region."
      },
      "description": "Genomic features found in the requested region, capped to max_results. totalCount reports the full count found before the cap."
    },
    "totalCount": {
      "type": "number",
      "description": "Total number of features found in the region before the max_results cap. Exceeds the returned features count when the list was capped."
    },
    "region": {
      "type": "string",
      "description": "The region queried, as provided."
    },
    "species": {
      "type": "string",
      "description": "The species queried."
    },
    "assemblyName": {
      "description": "Genome assembly the coordinates are on (e.g. GRCh38). Omitted only when it could not be resolved, in which case the notice says so.",
      "type": "string"
    },
    "notice": {
      "description": "Guidance about the result set: empty, large, capped, or missing assembly.",
      "type": "string"
    },
    "truncated": {
      "description": "True when the feature list was capped at max_results.",
      "type": "boolean"
    },
    "shown": {
      "description": "Number of features returned after the max_results cap.",
      "type": "number"
    },
    "cap": {
      "description": "The max_results limit applied to the feature list.",
      "type": "number"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `invalid_region`: The region string could not be parsed, contains invalid coordinates, or spans more than the 5,000,000-base maximum. `invalid_species`: The species string was not recognized by Ensembl. Other values are possible when a failure originates below the handler.",
              "examples": [
                "invalid_region",
                "invalid_species"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "features",
        "totalCount",
        "region",
        "species"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_predict_variant(variant, species, max_transcript_consequences, max_pubmed_ids_per_variant, include_all_colocated_pubmed)

Predict the functional consequences of a sequence variant using the Ensembl Variant Effect Predictor (VEP). Accepts three input formats: HGVS notation (transcript-relative, e.g. ENST00000380152.8:c.2T>A, or genomic, e.g. 13:g.32316462T>A); region+allele (chr:start:end:strand/allele, e.g. 1:65568:65568:1/T); and a dbSNP rsID (e.g. rs334). Returns the most severe consequence term, affected transcripts and genes, impact level (HIGH/MODERATE/LOW/MODIFIER), and any colocated known variants with clinical significance. HGVS input: provide the full notation including transcript version for best results. Region+allele input: Ensembl normalizes chromosome names and canonical vertebrate output omits the chr prefix (a chr-prefixed name is also accepted). By default the response caps transcript consequences (max_transcript_consequences) and per-variant PubMed IDs (max_pubmed_ids_per_variant) to keep large VEP results compact — well-studied variants like rs334 otherwise carry 60+ consequences and 100+ citations. Truthful totals are always reported; set a cap to 0 (or include_all_colocated_pubmed=true) to retrieve the full set.

Input Schema

{
  "type": "object",
  "properties": {
    "variant": {
      "type": "string",
      "minLength": 1,
      "description": "Variant in one of three formats: (1) HGVS notation — transcript-relative: ENST00000380152.8:c.2T>A; genomic: 13:g.32316462T>A; (2) Region+allele: chr:start:end:strand/allele — e.g. 1:65568:65568:1/T (strand is 1 for forward or -1 for reverse); (3) dbSNP rsID — e.g. rs334. Ensembl normalizes chromosome names; canonical vertebrate output omits the \"chr\" prefix, though a chr-prefixed name is also accepted."
    },
    "species": {
      "default": "homo_sapiens",
      "description": "Species in Ensembl internal format. Default is homo_sapiens. For non-human variants, set the appropriate species (e.g. mus_musculus for mouse). Use ensembl_list_species to discover valid values.",
      "type": "string",
      "minLength": 1
    },
    "max_transcript_consequences": {
      "default": 10,
      "description": "Maximum transcript consequences to return per VEP record. High-impact variants can affect 60+ transcripts; the default keeps the response focused on the top consequences. Set to 0 to return every transcript consequence uncapped. transcriptConsequencesTotal on each record always reports the true pre-cap count.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "max_pubmed_ids_per_variant": {
      "default": 10,
      "description": "Maximum PubMed IDs to return per colocated known variant. Well-studied variants (e.g. rs334) cite 100+ papers; the default trims each list. Set to 0 to return every PubMed ID uncapped. pubmedTotal on each colocated variant reports the true pre-cap count. Ignored when include_all_colocated_pubmed is true.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "include_all_colocated_pubmed": {
      "default": false,
      "description": "When true, return every PubMed ID for each colocated variant, overriding max_pubmed_ids_per_variant. Default false to keep responses compact.",
      "type": "boolean"
    }
  },
  "required": [
    "variant"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "results": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "input": {
            "description": "The input variant notation as submitted to VEP.",
            "type": "string"
          },
          "chromosome": {
            "description": "Chromosome the variant is on.",
            "type": "string"
          },
          "start": {
            "description": "Variant start position (1-based).",
            "type": "number"
          },
          "end": {
            "description": "Variant end position (1-based).",
            "type": "number"
          },
          "assemblyName": {
            "description": "Genome assembly name (e.g. GRCh38).",
            "type": "string"
          },
          "mostSevereConsequence": {
            "description": "Most severe Sequence Ontology consequence term across all transcripts (e.g. stop_gained, missense_variant, synonymous_variant).",
            "type": "string"
          },
          "transcriptConsequences": {
            "type": "array",
            "items": {
              "type": "object",
              "properties": {
                "transcriptId": {
                  "description": "Ensembl transcript ID (ENST…) affected by this variant.",
                  "type": "string"
                },
                "geneId": {
                  "description": "Ensembl gene ID (ENSG…) harboring the affected transcript.",
                  "type": "string"
                },
                "geneSymbol": {
                  "description": "Gene symbol (e.g. BRCA2, TP53).",
                  "type": "string"
                },
                "consequenceTerms": {
                  "type": "array",
                  "items": {
                    "type": "string",
                    "description": "A Sequence Ontology consequence term (e.g. missense_variant, stop_gained)."
                  },
                  "description": "Sequence Ontology consequence terms for this transcript."
                },
                "impact": {
                  "description": "Impact level: HIGH (frameshift, stop_gained), MODERATE (missense), LOW (synonymous), or MODIFIER.",
                  "type": "string"
                },
                "biotype": {
                  "description": "Transcript biotype (e.g. protein_coding).",
                  "type": "string"
                },
                "hgvsc": {
                  "description": "HGVS notation at the cDNA level (e.g. c.2T>A).",
                  "type": "string"
                },
                "hgvsp": {
                  "description": "HGVS notation at the protein level (e.g. p.Met1Thr).",
                  "type": "string"
                },
                "aminoAcids": {
                  "description": "Reference/alternate amino acids separated by \"/\" (e.g. M/T).",
                  "type": "string"
                },
                "sift": {
                  "description": "SIFT pathogenicity prediction for missense variants. Omitted when not applicable.",
                  "type": "object",
                  "properties": {
                    "prediction": {
                      "type": "string",
                      "description": "SIFT prediction: deleterious or tolerated."
                    },
                    "score": {
                      "type": "number",
                      "description": "SIFT score (0-1). Lower scores indicate more deleterious variants."
                    }
                  },
                  "required": [
                    "prediction",
                    "score"
                  ],
                  "additionalProperties": false
                },
                "polyphen": {
                  "description": "PolyPhen pathogenicity prediction for missense variants. Omitted when not applicable.",
                  "type": "object",
                  "properties": {
                    "prediction": {
                      "type": "string",
                      "description": "PolyPhen prediction: probably_damaging, possibly_damaging, or benign."
                    },
                    "score": {
                      "type": "number",
                      "description": "PolyPhen score (0-1). Higher scores indicate more damaging variants."
                    }
                  },
                  "required": [
                    "prediction",
                    "score"
                  ],
                  "additionalProperties": false
                }
              },
              "required": [
                "consequenceTerms"
              ],
              "additionalProperties": false,
              "description": "Consequence details for one affected transcript, including impact, HGVS notation, and pathogenicity scores."
            },
            "description": "Per-transcript consequence details, capped to max_transcript_consequences. High-impact variants may affect many transcripts; focus on canonical transcripts (isCanonical from ensembl_lookup_gene) for primary effect. transcriptConsequencesTotal reports the full count."
          },
          "transcriptConsequencesTotal": {
            "description": "Total transcript consequences available before the max_transcript_consequences cap. Equals the returned transcriptConsequences length when not capped.",
            "type": "number"
          },
          "colocatedVariants": {
            "type": "array",
            "items": {
              "type": "object",
              "properties": {
                "id": {
                  "description": "Known variant ID (e.g. rs1234567 for dbSNP entries).",
                  "type": "string"
                },
                "alleleString": {
                  "description": "Allele string showing reference/alternate (e.g. A/T).",
                  "type": "string"
                },
                "clinicalSignificance": {
                  "description": "Clinical significance terms from ClinVar (e.g. pathogenic, benign).",
                  "type": "array",
                  "items": {
                    "type": "string",
                    "description": "A clinical significance term for this colocated known variant."
                  }
                },
                "pubmed": {
                  "description": "PubMed IDs for literature associated with this variant, capped to max_pubmed_ids_per_variant. pubmedTotal reports the full count when the list was capped.",
                  "type": "array",
                  "items": {
                    "type": "number",
                    "description": "A PubMed ID for literature citing this variant."
                  }
                },
                "pubmedTotal": {
                  "description": "Total PubMed IDs available for this variant before the max_pubmed_ids_per_variant cap. Equals the returned pubmed length when the list was not capped.",
                  "type": "number"
                }
              },
              "additionalProperties": false,
              "description": "A known variant at the same genomic position from public databases (dbSNP, ClinVar)."
            },
            "description": "Known variants at the same position from public databases (dbSNP, ClinVar). Empty when the variant is novel."
          }
        },
        "required": [
          "transcriptConsequences",
          "colocatedVariants"
        ],
        "additionalProperties": false,
        "description": "VEP consequence record for one genomic position, with transcript consequences and colocated known variants."
      },
      "description": "VEP consequence records — typically one per input variant. Multiple records appear when a single notation matches multiple genomic positions."
    },
    "totalCount": {
      "type": "number",
      "description": "Number of VEP consequence records returned."
    },
    "notice": {
      "description": "Guidance when no results are returned or when caps omitted detail.",
      "type": "string"
    },
    "truncated": {
      "description": "True when transcript consequences were capped at max_transcript_consequences.",
      "type": "boolean"
    },
    "shown": {
      "description": "Total transcript consequences returned across all records after the cap.",
      "type": "number"
    },
    "cap": {
      "description": "The max_transcript_consequences limit applied.",
      "type": "number"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `invalid_notation`: The variant notation is malformed or cannot be parsed by VEP. `not_found`: The variant location falls outside any known transcript or assembly region. Other values are possible when a failure originates below the handler.",
              "examples": [
                "invalid_notation",
                "not_found"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "results",
        "totalCount"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_get_homology(symbol, id, species, target_species, type, ...)

Find orthologs and/or paralogs of a gene across species. Returns each homolog's stable ID, species, homology type (ortholog_one2one, ortholog_one2many, paralog_many2many, etc.), perc_id (percent identity), perc_pos (percent positives), and taxonomy level. Essential for cross-species research — for example, "what is the mouse equivalent of human TP53?" or "how conserved is BRCA2 across mammals?". Provide either symbol + species or a stable gene ID. Target species can be filtered to a single species or left open to return all available homologs.

Input Schema

{
  "type": "object",
  "properties": {
    "symbol": {
      "description": "Gene symbol in the source species (e.g. BRCA2, TP53). Species defaults to homo_sapiens; set species for other organisms. Cannot be combined with id.",
      "type": "string"
    },
    "id": {
      "description": "Ensembl stable gene ID (e.g. ENSG00000139618). Use ensembl_lookup_gene to get the stable ID from a symbol. Cannot be combined with symbol.",
      "type": "string"
    },
    "species": {
      "default": "homo_sapiens",
      "description": "Source species (the species the query gene belongs to) in Ensembl internal format. Default is homo_sapiens. Use ensembl_list_species to discover valid values.",
      "type": "string",
      "minLength": 1
    },
    "target_species": {
      "description": "Filter to homologs in a single target species (e.g. mus_musculus for mouse). Omit to return homologs across all available species. Use ensembl_list_species to discover valid values.",
      "type": "string"
    },
    "type": {
      "default": "orthologues",
      "description": "Type of homologs to return. orthologues: genes related by speciation (cross-species equivalents). paralogues: genes related by duplication (within or across species). all: both orthologs and paralogs.",
      "type": "string",
      "enum": [
        "orthologues",
        "paralogues",
        "all"
      ]
    },
    "max_results": {
      "default": 25,
      "description": "Maximum number of homologs to return. Broad orthology queries (e.g. BRCA2 across all species) can return 150+ homologs; the default keeps responses focused. Set to 0 to return every homolog uncapped. totalCount always reports the true number available before this cap.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "homologs": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "targetId": {
            "type": "string",
            "description": "Ensembl stable ID of the homologous gene in the target species."
          },
          "targetSpecies": {
            "description": "Target species in Ensembl internal format (e.g. mus_musculus).",
            "type": "string"
          },
          "type": {
            "description": "Homology type: ortholog_one2one, ortholog_one2many, ortholog_many2many, paralog_many2many, within_species_paralog, or similar.",
            "type": "string"
          },
          "percId": {
            "description": "Percent identity between the query and target gene sequences (0-100). Higher values indicate more conserved sequences.",
            "type": "number"
          },
          "percPos": {
            "description": "Percent positive (similar) positions in the alignment (0-100). Includes conservative substitutions as well as identical residues.",
            "type": "number"
          },
          "taxonomyLevel": {
            "description": "Last common ancestor taxonomic level for this homology relationship (e.g. Amniota, Vertebrata, Bilateria).",
            "type": "string"
          }
        },
        "required": [
          "targetId"
        ],
        "additionalProperties": false,
        "description": "A single homologous gene with its stable ID, species, homology type, and sequence identity metrics."
      },
      "description": "Homologous genes found for the query gene, capped to max_results. totalCount reports the full count available before the cap."
    },
    "totalCount": {
      "type": "number",
      "description": "Total number of homologs available before the max_results cap. Exceeds the returned homologs count when the list was capped."
    },
    "queryId": {
      "type": "string",
      "description": "The resolved Ensembl gene ID used for the homology query."
    },
    "querySpecies": {
      "type": "string",
      "description": "The source species used for the query."
    },
    "queryType": {
      "type": "string",
      "description": "The homology type queried (orthologues, paralogues, or all)."
    },
    "notice": {
      "description": "Guidance when no homologs are found or the list was capped.",
      "type": "string"
    },
    "truncated": {
      "description": "True when the homolog list was capped at max_results.",
      "type": "boolean"
    },
    "shown": {
      "description": "Number of homologs returned after the max_results cap.",
      "type": "number"
    },
    "cap": {
      "description": "The max_results limit applied to the homolog list.",
      "type": "number"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The gene symbol or stable ID was not found in Ensembl. `no_input`: Neither symbol nor id was provided. `conflicting_input`: Both symbol and id were provided. Other values are possible when a failure originates below the handler.",
              "examples": [
                "not_found",
                "no_input",
                "conflicting_input"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "homologs",
        "totalCount",
        "queryId",
        "querySpecies",
        "queryType"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}
🟢ensembl_get_xrefs(id, dbname)

Retrieve cross-database references for a gene or feature — HGNC, UniProt, EntrezGene, OMIM, RefSeq, Reactome, and others. Returns each xref with its database name, primary ID, display ID, and description. The dbname filter narrows to specific databases; omit to return all xrefs. IDs returned here chain to protein (pubchem via UniProt), literature (pubmed via PubMed IDs), disease (OMIM via MIM_GENE), and pathway (Reactome) resources. Requires an Ensembl stable ID — use ensembl_lookup_gene to get the ENSG… ID first. Common dbname values: HGNC, Uniprot_gn, EntrezGene, MIM_GENE, RefSeq_mRNA, RefSeq_peptide, Reactome, GO (Gene Ontology), ChEMBL.

Input Schema

{
  "type": "object",
  "properties": {
    "id": {
      "type": "string",
      "minLength": 1,
      "description": "Ensembl stable gene ID (ENSG…) or transcript ID (ENST…). Use ensembl_lookup_gene to get the stable ID from a gene symbol. xrefs/id returns the full cross-reference set (56+ entries for well-annotated genes like BRCA2)."
    },
    "dbname": {
      "description": "Filter to a specific external database by its Ensembl internal name. Examples: HGNC (HGNC gene ID), Uniprot_gn (UniProt gene name), EntrezGene (NCBI Gene ID), MIM_GENE (OMIM disease gene), RefSeq_mRNA (NCBI RefSeq transcript), Reactome (pathway IDs), GO (Gene Ontology terms). Omit to return all available xrefs.",
      "type": "string"
    }
  },
  "required": [
    "id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false
}

Output Schema

{
  "type": "object",
  "properties": {
    "xrefs": {
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "dbname": {
            "description": "Database name in Ensembl internal format (e.g. HGNC, Uniprot_gn, EntrezGene, MIM_GENE, RefSeq_mRNA).",
            "type": "string"
          },
          "dbDisplayName": {
            "description": "Human-readable database display name.",
            "type": "string"
          },
          "primaryId": {
            "description": "Primary identifier in the external database (e.g. HGNC:1101 for BRCA2 in HGNC, P51587 for BRCA2 in UniProt).",
            "type": "string"
          },
          "displayId": {
            "description": "Display identifier — often the same as primaryId but may be a formatted accession.",
            "type": "string"
          },
          "description": {
            "description": "Description of the cross-reference entry.",
            "type": "string"
          }
        },
        "additionalProperties": false,
        "description": "A single cross-database reference entry with database name, primary ID, and description."
      },
      "description": "Cross-database references for the queried Ensembl ID."
    },
    "totalCount": {
      "type": "number",
      "description": "Total number of cross-references returned."
    },
    "queriedId": {
      "type": "string",
      "description": "The Ensembl stable ID that was queried."
    },
    "notice": {
      "description": "Guidance when no cross-references are found.",
      "type": "string"
    },
    "error": {
      "description": "Present when the call failed. Absent on success.",
      "type": "object",
      "properties": {
        "code": {
          "type": "integer",
          "minimum": -9007199254740991,
          "maximum": 9007199254740991,
          "description": "JSON-RPC error code for this failure."
        },
        "message": {
          "type": "string",
          "description": "Human-readable description of what went wrong."
        },
        "data": {
          "type": "object",
          "properties": {
            "reason": {
              "type": "string",
              "description": "Machine-readable failure mode. Declared by this tool: `not_found`: The Ensembl stable ID was not found or has no cross-references. Other values are possible when a failure originates below the handler.",
              "examples": [
                "not_found"
              ]
            },
            "recovery": {
              "description": "Actionable next step for the caller.",
              "type": "object",
              "properties": {
                "hint": {
                  "type": "string"
                }
              },
              "required": [
                "hint"
              ],
              "additionalProperties": {}
            },
            "retryable": {
              "description": "Whether retrying may succeed.",
              "type": "boolean"
            }
          },
          "additionalProperties": {}
        }
      },
      "required": [
        "code",
        "message"
      ],
      "additionalProperties": {}
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "additionalProperties": false,
  "anyOf": [
    {
      "not": {
        "required": [
          "error"
        ]
      },
      "required": [
        "xrefs",
        "totalCount",
        "queriedId"
      ]
    },
    {
      "required": [
        "error"
      ]
    }
  ]
}

Community

Rate this Server

Evidence

Recent observations

verifiedversion not recorded7 tools
verifiedversion not recorded7 tools