Skip to main content

Connector operation reference

Look up exact operation names, required fields, defaults and example calls. To choose a data source, start with the database catalog. Expand the Connector family you intend to call; availability and credentials must be configured separately.

Where the example calls run

The host object is supplied by Open-Science's agent execution environment. The JavaScript below is an agent-side call fragment, not a standalone Node.js program and not a method on the public Task SDK client. Ask the agent to load the relevant Connector instructions and use the matching operation. A framework may expose a Python bridge instead of this JavaScript form.

First enable the Connector in Settings → Connectors, configure any required credentials, and grant access to the selected Specialist if applicable. The call still follows the conversation's permission policy. Public Node.js integrations can manage Connector settings with the Task SDK, but cannot obtain this host by importing that client.

Read a result before chaining calls

Example Pass returned PubMed IDs to a metadata lookup

For example, ask: Use PubMed to search for PRISMA reporting guidance; return the total match count and five PMIDs. The operation search_articles returns a total and a page of identifiers. Feed those returned PMIDs to get_article_metadata to obtain titles, authors and DOI links. An empty page, a truncated result and an authentication error need different handling.

Returned informationUse it for
Total match count and returned rowsDistinguish a small page from the complete result set
truncated, records_truncated or family-specific completeness flagsDecide whether to page, narrow the query or retrieve the rest
not_found, missing, not_processedIdentify unresolved inputs and retry only appropriate items
DOI, accession, source URL and release/buildRetain the identity and source needed for subsequent queries
Full-text status or license noteDecide whether text was retrieved and can be reused

Return field names differ by operation. The descriptions and downloadable schemas below specify each contract; the table is not a universal JSON response. Use the right-hand family list to jump, then expand that family's parameters. Searching for an operation name also opens its containing group.

Operation inputs

Expand one Connector at a time. Required fields are marked required; limits/defaults shown here come from the application schema. Consult the complete downloadable registry for nested JSON schemas, full return descriptions and agent-side call examples. Do not pass a generic id when a tool expects accessions, cids, rs_id or another namespace-specific field.

Chemistry

Show operations and parameters

pubchem_search_compounds

Resolve a chemical identifier (name, SMILES, InChIKey, or CID) to PubChem CIDs, optionally with core computed properties for the top hits.

FieldTypeRequirement and constraints
querystringrequired
namespacestringoptional; default: "name"; enum: ["name", "smiles", "inchikey", "cid"]
max_cidsintegeroptional; default: 25; minimum: 1; maximum: 100
with_propertiesbooleanoptional; default: true
const result = await host.mcp("chemistry", "pubchem_search_compounds", {"query": "aspirin", "max_cids": 25})

pubchem_get_compounds

Full computed-property records for a batch of PubChem CIDs, with optional capped synonym lists.

FieldTypeRequirement and constraints
cidsarray of integerrequired; minItems: 1; maxItems: 50
include_synonymsbooleanoptional; default: false
max_synonymsintegeroptional; default: 30
const result = await host.mcp("chemistry", "pubchem_get_compounds", {"cids": [2244, 2519], "include_synonyms": false})

2D Tanimoto similarity search over all of PubChem for a query SMILES (synchronous fastsimilarity_2d route, no job polling).

FieldTypeRequirement and constraints
smilesstringrequired
thresholdintegeroptional; default: 90; minimum: 1; maximum: 100
max_recordsintegeroptional; default: 50; minimum: 1; maximum: 200
with_propertiesbooleanoptional; default: false
const result = await host.mcp("chemistry", "pubchem_similarity_search", {"smiles": "CC(=O)OC1=CC=CC=C1C(=O)O", "threshold": 90})

pubchem_get_bioassay_summary

Bioassay activity summary for one PubChem compound — which assays tested it, against which targets, with what outcome and potency.

FieldTypeRequirement and constraints
cidintegerrequired
active_onlybooleanoptional; default: false
max_rowsintegeroptional; default: 100; minimum: 1; maximum: 1000
const result = await host.mcp("chemistry", "pubchem_get_bioassay_summary", {"cid": 2244, "active_only": true})

pubchem_get_safety

GHS safety classification for one PubChem compound (PUG-View 'GHS Classification' heading), aggregated across reporting sources.

FieldTypeRequirement and constraints
cidintegerrequired
const result = await host.mcp("chemistry", "pubchem_get_safety", {"cid": 702})

Full-text search over ChEBI entities (names, synonyms, formulae, InChIKeys).

FieldTypeRequirement and constraints
termstringrequired
max_resultsintegeroptional; default: 20; minimum: 1; maximum: 100
pageintegeroptional; default: 1; minimum: 1
const result = await host.mcp("chemistry", "chebi_search", {"term": "caffeine", "max_results": 20})

chebi_get_entity

Full ChEBI entity record: names, structure, chemical data, roles and cross-references.

FieldTypeRequirement and constraints
chebi_idstringrequired
max_synonymsintegeroptional; default: 30
max_xrefsintegeroptional; default: 50
const result = await host.mcp("chemistry", "chebi_get_entity", {"chebi_id": "CHEBI:27732"})

chebi_get_ontology

Ontology relations of a ChEBI entity — what it IS (outgoing: is a / has role / conjugate acid...) and what points AT it (incoming: children/derivatives).

FieldTypeRequirement and constraints
chebi_idstringrequired
relation_typestringoptional
max_relationsintegeroptional; default: 100
const result = await host.mcp("chemistry", "chebi_get_ontology", {"chebi_id": "CHEBI:27732", "relation_type": "has role"})

rhea_search_reactions

Search Rhea master reactions by equation text, participant ChEBI id, or EC number (query type auto-detected).

FieldTypeRequirement and constraints
querystringrequired
limitintegeroptional; default: 50; minimum: 1; maximum: 500
const result = await host.mcp("chemistry", "rhea_search_reactions", {"query": "caffeine", "limit": 50})

rhea_get_reaction

Full record for one Rhea reaction: equation, participants with ChEBI ids and stoichiometry, EC links, direction family and literature.

FieldTypeRequirement and constraints
rhea_idstringrequired
const result = await host.mcp("chemistry", "rhea_get_reaction", {"rhea_id": "10280"})

bindingdb_ligands_by_target

Measured binding affinities (Ki/Kd/IC50/EC50) of all BindingDB ligands against one protein target, by UniProt accession.

FieldTypeRequirement and constraints
uniprotstringrequired
affinity_cutoff_nmnumberoptional; default: 10000
max_rowsintegeroptional; default: 100; minimum: 1; maximum: 1000
const result = await host.mcp("chemistry", "bindingdb_ligands_by_target", {"uniprot": "P00533", "affinity_cutoff_nm": 100})

bindingdb_targets_by_compound

Protein targets with measured affinities for compounds 2D-similar to a query SMILES — "what does this molecule (or its close analogs) bind?".

FieldTypeRequirement and constraints
smilesstringrequired
similaritynumberoptional; default: 0.85; minimum: 0.5; maximum: 1
max_rowsintegeroptional; default: 100; minimum: 1; maximum: 1000
const result = await host.mcp("chemistry", "bindingdb_targets_by_compound", {"smiles": "CC(=O)OC1=CC=CC=C1C(=O)O", "similarity": 0.85})

Literature Graph

Show operations and parameters

openalex_search_works

Search OpenAlex scholarly works (all disciplines, ~250M records) with year/type/OA/venue filters. Args: query (free-text over title+abstract+fulltext; optional if a filter is set), year_from, year_to (inclusive years), work_type (article/review/preprint/book-chapter/dataset/dissertation), open_access_only, venue (S-id, openalex.org URL, ISSN, or a plain name resolved to the top sources hit — surfaced in venue_resolved; pass an exact ID to skip resolution), sort (relevance default / cited_by_count / publication_date), max_records (default 50, hard ceiling 500; pages of 200), include_abstracts (reconstructed from the inverted index, but ONLY for verified-open licenses — cc-by/cc-by-sa/cc0/public-domain; others get abstract=null + abstract_policy note + abstract_license; adds bulk). Returns {query, filters, sort, api_total, n_records_returned, records_truncated, records}; each record is the lean work shape (openalex_id, doi, pmid, title, publication_year/date, type, language, is_retracted, authors[...], source{...}, biblio, cited_by_count, fwci, referenced_works_count, open_access{...}, best_oa_pdf_url, primary_topic, keywords).

FieldTypeRequirement and constraints
querystringoptional
year_fromintegeroptional
year_tointegeroptional
work_typestringoptional
open_access_onlybooleanoptional
venuestringoptional
sortstringoptional; default: "relevance"; enum: ["relevance", "cited_by_count", "publication_date"]
max_recordsintegeroptional; default: 50
include_abstractsbooleanoptional; default: false
const result = await host.mcp("literature", "openalex_search_works", {"query": "CRISPR base editing", "year_from": 2020, "open_access_only": true, "sort": "cited_by_count", "max_records": 25})

openalex_get_work

Fetch one OpenAlex work in full — metadata, abstract (reconstructed from the inverted index, license-gated as in openalex_search_works), OA locations, referenced_works (outgoing W-ids — hydrate with openalex_references) and counts_by_year. Args: work_id (W-id, openalex.org URL, bare DOI, or doi.org URL). DOI lookups resolve via the claimant filter; when several works share one DOI the most-cited is selected and doi_claimants + doi_resolution_note are included. Raises not-found for unknown IDs/DOIs.

FieldTypeRequirement and constraints
work_idstringrequired
const result = await host.mcp("literature", "openalex_get_work", {"work_id": "W2741809807"})

openalex_citations

List works that CITE a given work (incoming citations) via OpenAlex's citation graph. Args: work_id (W-id/URL/DOI — DOIs cost one extra resolution request), sort (cited_by_count default / publication_date / relevance), max_records (default 50, ceiling 500), include_abstracts. Returns {work_id, api_total (the true citing-work count), n_records_returned, records_truncated, records} (lean work records).

FieldTypeRequirement and constraints
work_idstringrequired
sortstringoptional; default: "cited_by_count"; enum: ["cited_by_count", "publication_date", "relevance"]
max_recordsintegeroptional; default: 50
include_abstractsbooleanoptional; default: false
const result = await host.mcp("literature", "openalex_citations", {"work_id": "W2741809807", "sort": "cited_by_count", "max_records": 50})

openalex_references

List the works a given work CITES (outgoing references), hydrated to full metadata in reference-list order. Args: work_id (W-id/URL/DOI), max_records (default 100, ceiling 500; hydration batched 50/request). Returns {work_id, n_references, n_records_returned, records_truncated, references_not_hydrated (IDs OpenAlex has no record for — never silently dropped), reference_ids (ALL outgoing W-ids), records}.

FieldTypeRequirement and constraints
work_idstringrequired
max_recordsintegeroptional; default: 100
const result = await host.mcp("literature", "openalex_references", {"work_id": "W2741809807", "max_records": 100})

openalex_search_authors

Search OpenAlex author profiles by name. Args: query (matches display name + alternatives; expect homonyms — check affiliations/topics/ORCID), max_records (default 25, ceiling 500). Returns {query, api_total, n_records_returned, records_truncated, records}; each record {author_id, name, orcid, works_count, cited_by_count, h_index, i10_index, affiliations[{institution, years}], last_known_institutions, top_topics}. Use author_id with openalex_get_author.

FieldTypeRequirement and constraints
querystringrequired
max_recordsintegeroptional; default: 25
const result = await host.mcp("literature", "openalex_search_authors", {"query": "Jennifer Doudna", "max_records": 25})

openalex_get_author

Fetch one OpenAlex author profile plus their top-cited works. Args: author_id (A-id, openalex.org URL, or ORCID; CAVEAT: OpenAlex's ORCID pointer can resolve to a sparse duplicate — prefer the A-id from openalex_search_authors), works_sample (default 10, max 200; 0 skips the extra request). Returns the author record plus counts_by_year, top_works_total (true total works count) and top_works (lean work records by citations).

FieldTypeRequirement and constraints
author_idstringrequired
works_sampleintegeroptional; default: 10
const result = await host.mcp("literature", "openalex_get_author", {"author_id": "A5023888391", "works_sample": 10})

openalex_venue_info

Look up journals/repositories ('sources') in OpenAlex — OA status, DOAJ listing, APC, citation metrics. Args: venue (exact S-id, openalex.org URL, or ISSN for a single record; anything else is a name search), max_records (default 10, ceiling 500; name-search only). Returns: exact -> one source record + counts_by_year; name search -> {query, api_total, n_records_returned, records_truncated, records}. Source record: {source_id, display_name, type, issn_l, issn, host_organization, country_code, homepage_url, is_oa, is_in_doaj, is_core, apc_usd, works_count, cited_by_count, h_index, two_year_mean_citedness, first/last_publication_year, top_topics}.

FieldTypeRequirement and constraints
venuestringrequired
max_recordsintegeroptional; default: 10
const result = await host.mcp("literature", "openalex_venue_info", {"venue": "Nature", "max_records": 10})

Search arXiv preprints (physics, math, CS, stats, q-bio, ...) via the official Atom API. Args: query (arXiv query string; plain terms search all fields, field prefixes ti:/au:/abs: and booleans AND/OR/ANDNOT work; optional if category or a date range is set), category (arXiv code AND-ed in, e.g. q-bio.GN, cs.LG, stat.ML), date_from / date_to (submission date YYYY-MM-DD, inclusive), start (0-based paging offset; the API paces ~3s between requests — page politely), max_results (default 25, max 100 per call), sort_by (relevance default / submittedDate / lastUpdatedDate), sort_order (descending default / ascending). Returns {search_query (the exact query sent), api_total (arXiv's total match count), start_index, n_records_returned, records_truncated, sort_by, sort_order, records}; each record {arxiv_id, version, id_versioned, title, abstract, authors, published, updated, primary_category, categories, doi, journal_ref, comment, abs_url, pdf_url}. doi/journal_ref appear only after journal publication. Malformed queries raise an error (arXiv's HTTP-200 error feed is detected, never returned as data).

FieldTypeRequirement and constraints
querystringoptional
categorystringoptional
date_fromstringoptional
date_tostringoptional
startintegeroptional; default: 0
max_resultsintegeroptional; default: 25
sort_bystringoptional; default: "relevance"; enum: ["relevance", "submittedDate", "lastUpdatedDate"]
sort_orderstringoptional; default: "descending"; enum: ["descending", "ascending"]
const result = await host.mcp("literature", "arxiv_search", {"query": "ti:transformer", "category": "cs.LG", "max_results": 10})

arxiv_get_papers

Batch-fetch arXiv paper metadata (incl. abstracts) by ID — one paced request for up to 100 papers. Args: arxiv_ids (up to 100 IDs in any common form — 2103.14030, versioned 2103.14030v2, old-style q-bio/0601001, arXiv:-prefixed, or abs/pdf URLs; unversioned IDs resolve to the latest version). Returns {n_requested, n_found, duplicates (inputs that resolved to an already-returned paper), not_found (unknown AND malformed IDs — arXiv silently skips unknowns and rejects whole batches over malformed ones; this tool does neither), records} — records in requested order, same shape as arxiv_search records. Withdrawn papers still return metadata (check comment for withdrawal notes).

FieldTypeRequirement and constraints
arxiv_idsarray of stringrequired
const result = await host.mcp("literature", "arxiv_get_papers", {"arxiv_ids": ["2103.14030", "1706.03762v5"]})

crossref_get_work

Retrieve publisher-deposited metadata for a Crossref DOI. A bare DOI, doi: prefix or doi.org URL is accepted. No API key is required. If the DOI belongs to another registration agency, use the matching service; a Crossref 404 does not prove the DOI is invalid. Check the returned DOI, title and source_url.

FieldTypeRequirement and constraints
doistringrequired; minLength: 1; maxLength: 2048
const result = await host.mcp("literature", "crossref_get_work", {"doi": "10.1038/nature12968"})

crossref_get_updates

Read deposited correction, retraction and other update relationships. updated_by points to notices updating this work; update_to points to works updated by this DOI. Preserve relationship direction and source labels. Empty arrays do not establish reliability or prove that no retraction exists.

FieldTypeRequirement and constraints
doistringrequired; minLength: 1; maxLength: 2048
const result = await host.mcp("literature", "crossref_get_updates", {"doi": "10.1038/nature12968"})

datacite_search_records

Search public DataCite dataset/software DOI metadata. Supply query, related_doi or both; query uses DataCite query syntax. Keep the same filters and page_size when following next_page. Page-number retrieval is limited to the first 10,000 records: narrow the query if necessary. Check related_identifiers, rights and landing URLs; metadata does not guarantee downloadable data or reuse permission.

FieldTypeRequirement and constraints
querystringoptional; minLength: 1; maxLength: 2000
related_doistringoptional; minLength: 1; maxLength: 2048
resource_typestringoptional; default: "dataset"; enum: ["dataset", "software"]
page_sizeintegeroptional; default: 20; minimum: 1; maximum: 100
pageintegeroptional; default: 1; minimum: 1; maximum: 10000
const result = await host.mcp("literature", "datacite_search_records", {"query": "climate", "resource_type": "dataset", "page_size": 5})

datacite_get_record

Retrieve one public DataCite DOI record, including titles, creators, resource type, rights, related identifiers and available version. Accepts a bare DOI, doi: prefix or doi.org URL. Check the identifier and relationship direction before using a linked dataset or software package.

FieldTypeRequirement and constraints
doistringrequired; minLength: 1; maxLength: 2048
const result = await host.mcp("literature", "datacite_get_record", {"doi": "10.14454/qdd3-ps68"})

PubMed

Show operations and parameters

search_articles

Search PubMed (biomedical & life-sciences literature via NCBI esearch) for articles matching a query. Returns the total match count plus a page of PMIDs. Supports PubMed field tags ([Title], [Author], [Journal], [MeSH Terms], ...), Boolean operators, date filtering and sort. PubMed does not index physics / CS / math / pure-chemistry papers.

FieldTypeRequirement and constraints
querystringrequired
max_resultsintegeroptional; default: 20
retstartintegeroptional; default: 0
sortstringoptional; enum: ["relevance", "pub_date", "author", "journal_name", "title"]
date_fromstringoptional
date_tostringoptional
datetypestringoptional; default: "pdat"; enum: ["pdat", "edat", "mdat"]
const result = await host.mcp("pubmed", "search_articles", {"query": "CRISPR gene editing", "max_results": 10})

get_article_metadata

Retrieve detailed article metadata from PubMed by PMID (bulk, via efetch): identifiers (pmid/pmc/doi), title, abstract, journal, authors with affiliations, publication date, MeSH terms, article types, language and citation. On every use, cite PubMed and include the returned article DOIs (identifiers.doi) as links.

FieldTypeRequirement and constraints
pmids['string', 'array']required
const result = await host.mcp("pubmed", "get_article_metadata", {"pmids": ["35486828", "33264437"]})

Find related PubMed content for one or more source PMIDs via NCBI elink. pubmed_pubmed (default) returns similar articles ranked by word-weighted similarity of titles/abstracts/MeSH (NOT citations); pubmed_pmc returns full-text PMC links; pubmed_gene/pubmed_protein/pubmed_nucleotide return linked sequence/gene records.

FieldTypeRequirement and constraints
pmids['string', 'array']required
link_typestringoptional; default: "pubmed_pubmed"; enum: ["pubmed_pubmed", "pubmed_pmc", "pubmed_nucleotide", "pubmed_protein", "pubmed_gene"]
max_resultsintegeroptional
const result = await host.mcp("pubmed", "find_related_articles", {"pmids": ["35486828"], "link_type": "pubmed_pubmed"})

lookup_article_by_citation

Resolve bibliographic citations to PMIDs via NCBI ecitmatch. Each citation supplies some of {journal, year, volume, first_page, author, key}; provide 2-3+ fields for reliable matching. Use when you have a reference list and need PMIDs.

FieldTypeRequirement and constraints
citationsarray of objectrequired
const result = await host.mcp("pubmed", "lookup_article_by_citation", {"citations": [{"journal": "Science", "year": 1987, "volume": "235", "first_page": "182", "author": "Palmenberg AC"}]})

convert_article_ids

Convert between PMID, PMCID and DOI via the NCBI/PMC ID Converter. Homogeneous input ids per call (set id_type to match). Commonly used to check whether a PMID has a PMCID (i.e. full text in PMC) before calling get_full_text_article.

FieldTypeRequirement and constraints
ids['string', 'array']required
id_typestringoptional; default: "pmid"; enum: ["pmid", "pmcid", "doi"]
const result = await host.mcp("pubmed", "convert_article_ids", {"ids": ["PMC9046468"], "id_type": "pmcid"})

get_full_text_article

Retrieve open-access full text from PubMed Central via Europe PMC by PMC id ("PMC12345" or "12345"). Returns structured section text plus the license; when full text is unavailable the reason is reported explicitly (fulltext_status). Only OA-subset articles have retrievable full text. On every use, cite PubMed and include the returned article DOIs as links.

FieldTypeRequirement and constraints
pmc_ids['string', 'array']required
const result = await host.mcp("pubmed", "get_full_text_article", {"pmc_ids": ["PMC9046468"]})

Report copyright and license status per PMID by combining PubMed CopyrightInformation, the PMC ID Converter (PMID -> PMCID/DOI), and the PMC <permissions> block (license type, ALI license URL, copyright statement/year). Use to check open-access reuse rights before reproducing content.

FieldTypeRequirement and constraints
pmids['string', 'array']required
const result = await host.mcp("pubmed", "get_copyright_status", {"pmids": ["35891187", "34375400"]})

Genes & Ontologies

Show operations and parameters

query_genes

Resolve gene identifiers/symbols via mygene.info (batched, up to 1000 terms/request). Use this to map gene symbols to Ensembl gene IDs, Entrez IDs, names, and any other mygene.info field — or the reverse (set scopes to the namespace of your input terms, e.g. "entrezgene", "ensembl.gene", "symbol,alias"). Args: terms (query terms, e.g. ["TP53","BRCA1"]; terms containing commas are not supported); scopes (comma-separated identifier namespaces to match terms against); fields (comma-separated mygene fields to return, or "all"); species (common name "human"/"mouse" or NCBI taxid). Returns {n_input, n_records, not_found, records}. A term matching several genes yields several records (each carries its query). Records are deterministically ordered (input order, then _id).

FieldTypeRequirement and constraints
termsarray of stringrequired
scopesstringoptional
fieldsstringoptional; default: "symbol,name,taxid,entrezgene,ensembl.gene"
speciesstringoptional
const result = await host.mcp("genes", "query_genes", {"terms": ["TP53", "BRCA1"], "scopes": "symbol,alias", "fields": "symbol,name,entrezgene,ensembl.gene", "species": "human"})

list_ontologies

List ontologies in the EBI Ontology Lookup Service (OLS4). With ontology_ids (e.g. ["efo","cl","chebi","go","mondo"]): fetch structured metadata records for just those ontologies; unknown IDs are reported in not_found. Without: the complete OLS4 catalogue (~250 ontologies, paginated fully and count-verified). Returns: {records:[{ontology_id, title, version, status, num_terms, ...}], not_found:[...]} for an ID list, or {records:[...], total_elements, complete} for the full catalogue.

FieldTypeRequirement and constraints
ontology_idsarray of stringoptional
const result = await host.mcp("genes", "list_ontologies", {"ontology_ids": ["efo", "go", "mondo"]})

search_ontology_terms

Search ontology terms by label/synonym across one or more OLS4 ontologies. Typical uses: find an EFO ID for a disease name (ontologies=["efo"]), Cell Ontology terms for a cell type (["cl"]), ChEBI terms for a chemical (["chebi"]), GO terms by name (["go"]) — or search all ontologies at once. Args: query (term label, synonym, or identifier); ontologies (lowercase IDs to restrict to; None searches every ontology); exact (whole-string match); include_obsolete (default False); max_results (ranked by OLS relevance). Returns {query, total_found, n_returned, truncated, terms:[{curie, iri, label, short_form, ontology, description, type, is_defining_ontology}]}.

FieldTypeRequirement and constraints
querystringrequired
ontologiesarray of stringoptional
exactbooleanoptional; default: false
include_obsoletebooleanoptional; default: false
max_resultsintegeroptional; default: 20
const result = await host.mcp("genes", "search_ontology_terms", {"query": "asthma", "ontologies": ["efo"], "max_results": 20})

get_ontology_term

Fetch one ontology term's details, or its complete related-term set. With relation=None: full term record (label, synonyms, description, obsolete flag, direct parents). With a relation: the COMPLETE, fully paginated set of related terms — e.g. relation="hierarchicalChildren" for direct children incl. part_of etc., "descendants"/"hierarchicalDescendants" for the whole subtree, "ancestors"/"hierarchicalAncestors", "parents", "children". Retrieval is count-verified against the API's own total. Args: ontology (lowercase, e.g. "efo","go","cl","chebi"); term_id (CURIE "EFO:0000305"/"GO:0006281" or full IRI); relation (None or one of the listed); include_parents (include direct parent refs when relation is None). Returns: relation=None {curie, iri, label, ontology, short_form, synonyms, description, is_obsolete, has_children, parents}; otherwise {root, relation, total_elements, term_count, terms:[...]}.

FieldTypeRequirement and constraints
ontologystringrequired
term_idstringrequired
relationstringoptional; enum: ["parents", "children", "ancestors", "descendants", "hierarchicalParents", "hierarchicalChildren", "hierarchicalAncestors", "hierarchicalDescendants"]
include_parentsbooleanoptional; default: false
const result = await host.mcp("genes", "get_ontology_term", {"ontology": "go", "term_id": "GO:0006281", "relation": "children"})

get_go_annotations

Retrieve GO annotations for a UniProt gene product from QuickGO (complete, count-verified). Args: uniprot_accession (e.g. "P04637", prefix optional); aspect (omit for all aspects, or one of biological_process/molecular_function/cellular_component); evidence (None/all, a preset "experimental_manual"=manually-assigned experimental evidence, "automatic_iea"=electronic/IEA, or an explicit ECO code like "ECO:0000314"; three-letter GO evidence codes like IDA/IEA are NOT accepted — QuickGO silently ignores goEvidence, filter must use ECO codes); taxon_id (optional NCBI taxon, e.g. 9606); include_term_names (hydrate each record with GO term name/aspect/obsolete via one batched ontology lookup); max_records (cap on records; full set still retrieved and summarized; truncated flags the cap). Returns {gene_product, total_annotations, n_records, complete, truncated, distinct_go_ids (across ALL annotations), records:[{go_id, go_aspect, qualifier, go_evidence, eco_id, reference, assigned_by, date, ...}]}.

FieldTypeRequirement and constraints
uniprot_accessionstringrequired
aspectstringoptional; enum: ["biological_process", "molecular_function", "cellular_component"]
evidencestringoptional
taxon_idintegeroptional
include_term_namesbooleanoptional; default: false
max_recordsintegeroptional; default: 200
const result = await host.mcp("genes", "get_go_annotations", {"uniprot_accession": "P04637", "aspect": "molecular_function", "evidence": "experimental_manual"})

get_uniprot_entries

Fetch UniProtKB records for a list of accessions (batched OR-queries, not per-accession). Three modes: fields given → token-lean tabular retrieval of just those UniProt fields (e.g. ["accession","id","protein_name","gene_names","organism_name","length","sequence"]); format is ignored. format="fasta" → per-accession FASTA sequences. format="txt" → per-accession full UniProt flat-file text (complete annotation; can be very large — prefer fields). Args: accessions (e.g. ["P04637","P38398"]); format ("fasta"/"txt", ignored when fields given); fields (optional UniProt REST field names for tabular mode). Returns: fields mode {accessions, fields, n_records, records:[{<column>:value}]}; fasta/txt mode {accessions, format, n_found, missing, records:{accession:text}} — missing lists accessions UniProt returned no record for.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
formatstringoptional; enum: ["fasta", "txt"]
fieldsarray of stringoptional
const result = await host.mcp("genes", "get_uniprot_entries", {"accessions": ["P04637", "P38398"], "fields": ["accession", "id", "protein_name", "gene_names", "organism_name", "length"]})

map_reactome_pathways

Map gene symbols or UniProt accessions to Reactome pathways (AnalysisService token workflow). Args: identifiers (gene symbols if id_type="symbol", UniProt accessions if "uniprot"; no duplicates); id_type ("symbol"/"uniprot"); species (default "Homo sapiens"); resource (AnalysisService molecule-resource view "TOTAL" default; "UNIPROT" restricts to protein-level mappings); include_disease (service default True); compact (True → per-identifier low-level pathways only {stId,name,species} + reactome release version; False → full deterministic result: per-identifier complete pathway sets with entity/reaction statistics (p-values, FDR, found/total) and batch summary incl. identifiers_not_found). Returns: compact {tool, reactome_version, id_type, species, n_input, genes:{identifier:{found, n_lowlevel_pathways, pathways}}}; full adds per-pathway statistics and batch_summary.

FieldTypeRequirement and constraints
identifiersarray of stringrequired
id_typestringrequired; enum: ["symbol", "uniprot"]
speciesstringoptional; default: "Homo sapiens"
resourcestringoptional; default: "TOTAL"
include_diseasebooleanoptional; default: true
compactbooleanoptional; default: true
const result = await host.mcp("genes", "map_reactome_pathways", {"identifiers": ["TP53", "EGFR", "BRCA1"], "id_type": "symbol"})

Genomes

Show operations and parameters

ensembl_lookup

Look up an Ensembl gene/transcript/protein by stable ID or a gene by symbol; returns the core annotation record (location, biotype, canonical transcript, description). Args: query (Ensembl stable ID ENSG.../ENST.../ENSP..., versioned accepted; or a gene symbol/alias like BRAF — true stable IDs [ENS + optional species code + feature letter + >=6-digit block, or LRG_N] route to the ID endpoint; everything else, incl. symbols starting with "ENS" like ENSA, to the symbol endpoint); species (Ensembl species name for symbol lookups, default homo_sapiens; ignored for stable IDs); expand (include the child feature tree — a gene's transcripts/exons/translation; default off). Returns {found, query, species, record}; record is null when nothing matches, else the upstream lookup dict — for a gene {id, display_name, description, biotype, object_type, seq_region_name, start, end, strand, assembly_name, canonical_transcript, version, ...} with 1-based inclusive coordinates.

FieldTypeRequirement and constraints
querystringrequired
speciesstringoptional; default: "homo_sapiens"
expandbooleanoptional; default: false
const result = await host.mcp("genomes", "ensembl_lookup", {"query": "BRAF"})

ensembl_xrefs

External cross-references of an Ensembl stable ID — the bridge from Ensembl gene/transcript IDs to HGNC, NCBI (EntrezGene), UniProt, OMIM, RefSeq, Expression Atlas and others. Args: stable_id (ENSG.../ENST..., versioned accepted); external_db (optional exact upstream database-name filter, e.g. HGNC, EntrezGene, Uniprot_gn, MIM_GENE, RefSeq_mRNA; omit for all). Returns {stable_id, external_db, n_xrefs, xrefs} — the COMPLETE list (never truncated), sorted by (dbname, primary_id); each row {dbname, db_display_name, primary_id, display_id, description, synonyms, info_type}. Unknown IDs return n_xrefs:0.

FieldTypeRequirement and constraints
stable_idstringrequired
external_dbstringoptional
const result = await host.mcp("genomes", "ensembl_xrefs", {"stable_id": "ENSG00000157764", "external_db": "HGNC"})

ensembl_vep_variant

Predict variant consequences with Ensembl VEP — most-severe-first summary of the (often huge) per-transcript consequence list. Pass EITHER variant_id OR region+allele. Args: variant_id (dbSNP rsID rs7412, COSMIC COSV..., or HGMD ID); region (GRCh38 1-based inclusive chrom:start-end, e.g. 7:140753336-140753336; SNV start==end; insertion start=end+1; explicit strand suffix :1/:-1 accepted); allele (variant allele on forward strand for the region route, e.g. T or - for deletion); species (default homo_sapiens); max_consequences (cap on returned per-transcript rows, default 25; full count in n_transcript_consequences, rows kept are most severe HIGH>MODERATE>LOW>MODIFIER; transcript_consequences_truncated flags the cap). Returns {query, n_results, results:[{input, assembly_name, seq_region_name, start, end, strand, allele_string, most_severe_consequence, genes:[{gene_id, gene_symbol, worst_impact, n_transcripts}], n_transcript_consequences, transcript_consequences_truncated, transcript_consequences:[...], n_regulatory_feature_consequences, n_motif_feature_consequences, colocated_variants:[...]}]}. Unknown rsIDs raise with the upstream message.

FieldTypeRequirement and constraints
variant_idstringoptional
regionstringoptional
allelestringoptional
speciesstringoptional; default: "homo_sapiens"
max_consequencesintegeroptional; default: 25
const result = await host.mcp("genomes", "ensembl_vep_variant", {"variant_id": "rs7412", "max_consequences": 25})

ensembl_homology

Orthologues or paralogues of a gene from Ensembl Compara (condensed rows — no alignments/sequences). Args: gene_symbol (resolved to a stable ID in species first; pass exactly one of gene_symbol/gene_id); gene_id (ENSG...); homology_type (orthologues default/paralogues/projections); target_species (restrict to one species); target_taxon (NCBI taxon subtree, e.g. 9443 Primates; combinable with target_species, OR semantics); species (source species, default homo_sapiens); max_homologies (row cap default 200; n_total carries the complete count, homologies_truncated flags the cap). Returns {gene_id, gene_symbol, species, homology_type, target_species, target_taxon, n_total, homologies_truncated, homologies}; rows sorted by (species,id) {type, species, id, protein_id, taxonomy_level, method_link_type}. Quirk: the /homology/symbol route stalls — this tool always resolves symbols itself and queries by stable ID.

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
homology_typestringoptional; default: "orthologues"; enum: ["orthologues", "paralogues", "projections"]
target_speciesstringoptional
target_taxonintegeroptional
speciesstringoptional; default: "homo_sapiens"
max_homologiesintegeroptional; default: 200
const result = await host.mcp("genomes", "ensembl_homology", {"gene_symbol": "BRAF", "target_species": "mus_musculus"})

ensembl_sequence

Fetch sequence from Ensembl — by stable ID (gene/transcript/protein) or by genomic region. Pass EITHER stable_id OR region. Args: stable_id (ENSG.../ENST.../ENSP..., versioned accepted); region (1-based inclusive chrom:start..end or chrom:start-end, GRCh38 for human, max 10Mb); species (for region route, default homo_sapiens; ignored for stable IDs); seq_type (ID route: genomic default/cdna/cds/protein — protein only for ENST/ENSP; ignored for regions which always return genomic); max_bytes (payload guard default 400000 — larger sequences have seq omitted; length/sha256/metadata always returned; re-call with larger max_bytes for full text). Returns {found, query, seq_type, id, description, molecule, length, sha256, seq} — length in the unit implied by molecule (bases for dna, residues for protein); seq replaced by seq_omitted when capped; found:false with null fields for unknown stable IDs; malformed/oversized regions raise with the upstream message.

FieldTypeRequirement and constraints
stable_idstringoptional
regionstringoptional
speciesstringoptional; default: "homo_sapiens"
seq_typestringoptional; default: "genomic"; enum: ["genomic", "cdna", "cds", "protein"]
max_bytesintegeroptional; default: 400000
const result = await host.mcp("genomes", "ensembl_sequence", {"stable_id": "ENSP00000288602", "seq_type": "protein"})

ensembl_overlap_region

List Ensembl features overlapping a genomic region — genes, transcripts, regulatory features (enhancers/promoters), repeats, variants, karyotype bands. Args: region (1-based inclusive chrom:start-end GRCh38, e.g. 7:140719327-140925199; upstream rejects spans >5Mb — split larger); feature (gene default/transcript/exon/cds/regulatory/motif/repeat/variation/structural_variation/band/simple/misc); species (default homo_sapiens); max_features (row cap default 500; n_total carries the complete overlap count, features_truncated flags the cap). Returns {region, species, feature, n_total, features_truncated, features} sorted by (start,id). Row shape varies — genes {id, external_name, biotype, description, start, end, strand, canonical_transcript, ...}; regulatory {id, description, start, end, extended_start/end, ...}. Empty regions return n_total:0.

FieldTypeRequirement and constraints
regionstringrequired
featurestringoptional; default: "gene"; enum: ["gene", "transcript", "exon", "cds", "regulatory", "motif", "repeat", "variation", "structural_variation", "band", "simple", "misc"]
speciesstringoptional; default: "homo_sapiens"
max_featuresintegeroptional; default: 500
const result = await host.mcp("genomes", "ensembl_overlap_region", {"region": "7:140719327-140925199", "feature": "gene"})

ucsc_list_tracks

List data tracks available in a UCSC Genome Browser assembly (leaf tracks only — the queryable ones), optionally filtered. Args: genome (hg38 default/hg19/mm39/danRer11/... ~220 assemblies); filter_text (case-insensitive substring over name/short/long label, e.g. phyloP, TFBS, ClinVar; omit to list everything — hg38 has ~24k leaf tracks, you almost always want a filter); max_tracks (row cap default 200; n_total carries the full match count, tracks_truncated flags the cap). Returns {genome, filter_text, n_total, tracks_truncated, tracks} sorted by track name; each row {track, short_label, long_label, type, group, parent}. Use track with ucsc_track_data. Quirk: first call per genome downloads the full ~17MB listing and caches it for the process.

FieldTypeRequirement and constraints
genomestringoptional; default: "hg38"
filter_textstringoptional
max_tracksintegeroptional; default: 200
const result = await host.mcp("genomes", "ucsc_list_tracks", {"genome": "hg38", "filter_text": "phyloP", "max_tracks": 50})

ucsc_track_data

Fetch raw rows of any UCSC Genome Browser track in a region — the generic escape hatch behind ucsc_conservation / ucsc_tfbs_clusters (gene tracks, ClinVar, GWAS catalog, CpG islands, repeats, ...). Args: track (name from ucsc_list_tracks, e.g. knownGene, cpgIslandExt, clinvarMain); chrom (chr-prefixed, chr7/chrX — UCSC requires the prefix); start (0-based half-open; an Ensembl 1-based start is start-1 here); end (exclusive); genome (default hg38); max_rows (API maxItemsOutput, default 1000; truncated reflects the API's own maxItemsLimit flag). Returns {genome, track, chrom, start, end, track_type, items_returned, truncated, rows} — rows in upstream shape (BED-like {chrom, chromStart, chromEnd, name, score, ...}; wiggle {start, end, value}). Unknown tracks raise. Quirk: for some huge tracks the API caps output itself and points at dataDownloadUrl — echoed when present.

FieldTypeRequirement and constraints
trackstringrequired
chromstringrequired
startintegerrequired
endintegerrequired
genomestringoptional; default: "hg38"
max_rowsintegeroptional; default: 1000
const result = await host.mcp("genomes", "ucsc_track_data", {"track": "cpgIslandExt", "chrom": "chr7", "start": 140700000, "end": 140800000, "genome": "hg38"})

ucsc_conservation

Evolutionary conservation summary for a region from UCSC phyloP / phastCons tracks (base-wise scores over multi-species alignments). Args: chrom (chr-prefixed); start (0-based half-open); end (exclusive; span capped at 100000 bp — split larger); genome (default hg38); track (default phyloP100way; positive=conserved, negative=fast-evolving; alternatives hg38 phastCons100way, phyloP30way, phastCons30way, phyloP447way, phyloP470way; hg19 phyloP100wayAll/phastCons100way); include_values (also return per-base {start,end,value} rows capped at max_values, values_truncated flags the cap; default false = summary only); max_values (per-base cap default 2000). Returns {genome, track, chrom, start, end, span_bp, n_bases_covered, coverage_fraction, mean, min, max} (+values, values_truncated when requested). Stats weighted by each row's base span, clipped to window; uncovered bases lower coverage_fraction, not zero-scored. Non-score tracks raise; an upstream-truncated row list also raises.

FieldTypeRequirement and constraints
chromstringrequired
startintegerrequired
endintegerrequired
genomestringoptional; default: "hg38"
trackstringoptional; default: "phyloP100way"
include_valuesbooleanoptional; default: false
max_valuesintegeroptional; default: 2000
const result = await host.mcp("genomes", "ucsc_conservation", {"chrom": "chr7", "start": 140753330, "end": 140753380, "track": "phyloP100way"})

ucsc_tfbs_clusters

ENCODE transcription-factor binding site clusters overlapping a region (ChIP-seq peak clusters across hundreds of cell types) — which TFs bind where. Args: chrom (chr-prefixed); start (0-based half-open); end (exclusive); genome (hg38 default track encRegTfbsClustered ENCODE 3, or hg19 wgEncodeRegTfbsClusteredV3; other assemblies raise); max_rows (API maxItemsOutput default 1000; truncated reflects maxItemsLimit). Returns {genome, track, chrom, start, end, items_returned, truncated, n_factors, factors, clusters} — clusters sorted by (chromStart,name) {name (TF symbol e.g. CTCF), chrom, chromStart, chromEnd, score (0-1000), sourceCount (supporting experiments)}; factors is the distinct TF list. Score>=~600 and high sourceCount ~ robust binding.

FieldTypeRequirement and constraints
chromstringrequired
startintegerrequired
endintegerrequired
genomestringoptional; default: "hg38"
max_rowsintegeroptional; default: 1000
const result = await host.mcp("genomes", "ucsc_tfbs_clusters", {"chrom": "chr7", "start": 140699000, "end": 140760000, "genome": "hg38"})

ucsc_chrom_sizes

Chromosome/contig names and sizes of a UCSC assembly — for validating coordinates and iterating regions. Args: genome (default hg38); filter_text (case-insensitive substring on the name, e.g. chr1; omit for all — hg38 has 711 sequences, mostly alt/random/unplaced; primary chromosomes sort first); max_chroms (row cap default 100; n_total carries the full post-filter count, chroms_truncated flags the cap). Returns {genome, filter_text, chrom_count (assembly-wide from the API), n_total, chroms_truncated, chromosomes:[{name, size_bp}]} sorted by size descending.

FieldTypeRequirement and constraints
genomestringoptional; default: "hg38"
filter_textstringoptional
max_chromsintegeroptional; default: 100
const result = await host.mcp("genomes", "ucsc_chrom_sizes", {"genome": "hg38", "filter_text": "chr1", "max_chroms": 25})

Variants

Show operations and parameters

get_variant

Look up one gnomAD short variant by ID and return its population frequencies. variant_id is chrom-pos-ref-alt on the dataset's reference build (GRCh38 for r3/r4, GRCh37 for r2.1/ExAC), e.g. 19-44908822-C-T (APOE rs7412); use search_variants to resolve an rsID first.

FieldTypeRequirement and constraints
variant_idstringrequired
datasetstringoptional; default: "gnomad_r4"; enum: ["gnomad_r4", "gnomad_r4_non_ukb", "gnomad_r3", "gnomad_r3_controls_and_biobanks", "gnomad_r3_non_cancer", "gnomad_r3_non_neuro", "gnomad_r3_non_topmed", "gnomad_r3_non_v2", "gnomad_r2_1", "gnomad_r2_1_controls", "gnomad_r2_1_non_cancer", "gnomad_r2_1_non_neuro", "gnomad_r2_1_non_topmed", "exac"]
const result = await host.mcp("variants", "get_variant", {"variant_id": "19-44908822-C-T", "dataset": "gnomad_r4"})

search_variants

Search gnomAD for variant IDs matching a query string (an rsID like rs7412, a variant ID, or a prefix). Use this to resolve rsIDs to chrom-pos-ref-alt IDs for get_variant.

FieldTypeRequirement and constraints
querystringrequired
datasetstringoptional; default: "gnomad_r4"; enum: ["gnomad_r4", "gnomad_r4_non_ukb", "gnomad_r3", "gnomad_r3_controls_and_biobanks", "gnomad_r3_non_cancer", "gnomad_r3_non_neuro", "gnomad_r3_non_topmed", "gnomad_r3_non_v2", "gnomad_r2_1", "gnomad_r2_1_controls", "gnomad_r2_1_non_cancer", "gnomad_r2_1_non_neuro", "gnomad_r2_1_non_topmed", "exac"]
const result = await host.mcp("variants", "search_variants", {"query": "rs7412", "dataset": "gnomad_r4"})

gene_variants

List ALL gnomAD short variants in a gene (complete listing — can be thousands of rows for large genes). Pass exactly one of gene_symbol (HGNC symbol, e.g. APOE) or gene_id (Ensembl gene ID, e.g. ENSG00000130203).

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
datasetstringoptional; default: "gnomad_r4"; enum: ["gnomad_r4", "gnomad_r4_non_ukb", "gnomad_r3", "gnomad_r3_controls_and_biobanks", "gnomad_r3_non_cancer", "gnomad_r3_non_neuro", "gnomad_r3_non_topmed", "gnomad_r3_non_v2", "gnomad_r2_1", "gnomad_r2_1_controls", "gnomad_r2_1_non_cancer", "gnomad_r2_1_non_neuro", "gnomad_r2_1_non_topmed", "exac"]
const result = await host.mcp("variants", "gene_variants", {"gene_symbol": "APOE", "dataset": "gnomad_r4"})

gene_constraint

gnomAD gene constraint metrics: pLI, observed/expected LoF-missense-synonymous counts with oe ratios + 90% CI bounds, and per-class z-scores. Use to judge a gene's intolerance to loss-of-function (pLI >= 0.9 or oe_lof_upper (LOEUF) < 0.6 ~ LoF-intolerant). Pass exactly one of gene_symbol (e.g. TP53) or gene_id.

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
const result = await host.mcp("variants", "gene_constraint", {"gene_symbol": "TP53"})

region_variants

List ALL gnomAD short variants in a genomic region (max 1 Mb — split larger regions into consecutive windows). chrom is a chromosome name without chr prefix (1-22, X, Y); start/stop are 1-based inclusive and stop - start must be <= 1,000,000. The dataset determines the reference build of the coordinates (GRCh38 for r3/r4).

FieldTypeRequirement and constraints
chromstringrequired
startintegerrequired
stopintegerrequired
datasetstringoptional; default: "gnomad_r4"; enum: ["gnomad_r4", "gnomad_r4_non_ukb", "gnomad_r3", "gnomad_r3_controls_and_biobanks", "gnomad_r3_non_cancer", "gnomad_r3_non_neuro", "gnomad_r3_non_topmed", "gnomad_r3_non_v2", "gnomad_r2_1", "gnomad_r2_1_controls", "gnomad_r2_1_non_cancer", "gnomad_r2_1_non_neuro", "gnomad_r2_1_non_topmed", "exac"]
const result = await host.mcp("variants", "region_variants", {"chrom": "1", "start": 55039475, "stop": 55064852, "dataset": "gnomad_r4"})

liftover_variant

Map a variant ID between reference builds (GRCh37 <-> GRCh38) using gnomAD's liftover table. variant_id is chrom-pos-ref-alt on source_build. The route is directional: a GRCh38 ID passed with source_build=GRCh37 returns zero results, not an error.

FieldTypeRequirement and constraints
variant_idstringrequired
source_buildstringoptional; default: "GRCh37"; enum: ["GRCh37", "GRCh38"]
const result = await host.mcp("variants", "liftover_variant", {"variant_id": "1-55516888-G-GA", "source_build": "GRCh37"})

clinvar_variants

List ClinVar variants in a gene as mirrored by gnomAD, with clinical significance, review status and gold stars. The output pins gnomAD's ClinVar snapshot via clinvar_release_date. Pass exactly one of gene_symbol (e.g. BRCA1) or gene_id.

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
const result = await host.mcp("variants", "clinvar_variants", {"gene_symbol": "BRCA1"})

structural_variants

List gnomAD structural variants (deletions, duplications, insertions, inversions, CNVs...) overlapping a gene. Pass exactly one of gene_symbol (e.g. TP53) or gene_id. dataset is an SV pin — gnomad_sv_r4 (default, GRCh38) or gnomad_sv_r2_1 (GRCh37); SV IDs are release-specific.

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
datasetstringoptional; default: "gnomad_sv_r4"; enum: ["gnomad_sv_r4", "gnomad_sv_r2_1"]
const result = await host.mcp("variants", "structural_variants", {"gene_symbol": "TP53", "dataset": "gnomad_sv_r4"})

get_structural_variant

Look up one gnomAD structural variant by its release-specific SV ID (e.g. DEL_CHR17_599B1512 in gnomad_sv_r4). IDs do NOT carry across releases — dataset (gnomad_sv_r4 default, or gnomad_sv_r2_1) must match the release the ID came from.

FieldTypeRequirement and constraints
sv_idstringrequired
datasetstringoptional; default: "gnomad_sv_r4"; enum: ["gnomad_sv_r4", "gnomad_sv_r2_1"]
const result = await host.mcp("variants", "get_structural_variant", {"sv_id": "DEL_CHR17_A5250EA9", "dataset": "gnomad_sv_r4"})

mitochondrial_variants

List gnomAD mitochondrial variants with heteroplasmy-aware counts (ac_het, ac_hom, max_heteroplasmy) for a mitochondrial gene OR a chrM coordinate window. Pass a gene (gene_symbol like MT-TL1, or gene_id) OR a region (region_start + region_stop), not both.

FieldTypeRequirement and constraints
gene_symbolstringoptional
gene_idstringoptional
region_startintegeroptional
region_stopintegeroptional
datasetstringoptional; default: "gnomad_r4"; enum: ["gnomad_r4", "gnomad_r4_non_ukb", "gnomad_r3", "gnomad_r3_controls_and_biobanks", "gnomad_r3_non_cancer", "gnomad_r3_non_neuro", "gnomad_r3_non_topmed", "gnomad_r3_non_v2", "gnomad_r2_1", "gnomad_r2_1_controls", "gnomad_r2_1_non_cancer", "gnomad_r2_1_non_neuro", "gnomad_r2_1_non_topmed", "exac"]
const result = await host.mcp("variants", "mitochondrial_variants", {"gene_symbol": "MT-TL1", "dataset": "gnomad_r4"})

Search ClinVar directly (live NCBI, not gnomAD's snapshot) and return matching variation records with clinical significance, review status and gold stars. Requires a contact email (Settings → Credentials → Literature access → Contact email) per NCBI E-utilities usage policy. Args: query (a ClinVar Entrez query — free text like "TP53 R175H" or an HGVS string works, and fielded terms compose with AND/OR/NOT, e.g. BRCA1[gene], pathogenic[CLIN_SIG], "Lynch syndrome"[dis], single_nucleotide_variant[Type of variation]; an rsID also works but clinvar_variant_by_rsid returns fuller records), max_records (page cap 1-200, default 50). The match TOTAL is always reported; when total > max_records the list is a capped prefix (ClinVar relevance/recency order) and truncated is true. NCBI E-utilities intermittently return HTTP 500 under load — retry once a few seconds later if that surfaces.

FieldTypeRequirement and constraints
querystringrequired
max_recordsintegeroptional; default: 50
const result = await host.mcp("variants", "clinvar_search", {"query": "BRCA1 pathogenic[CLIN_SIG]", "max_records": 50})

clinvar_get_records

Fetch full ClinVar records for a batch of VCV/RCV accessions or bare variation IDs. Requires a contact email (Settings → Credentials → Literature access → Contact email) per NCBI E-utilities usage policy. Args: accessions (up to 50 identifiers, mixed forms accepted — VCV000045122 (versioned VCV000045122.3 ok; resolved locally, free), RCV000019428 (each RCV costs one extra esearch), or a bare ClinVar variation ID (45122). rsIDs are rejected — use clinvar_variant_by_rsid. An RCV (one variant-condition pair) resolves to its parent VCV variation record). Never silently drops an input.

FieldTypeRequirement and constraints
accessions['string', 'array']required
const result = await host.mcp("variants", "clinvar_get_records", {"accessions": ["VCV000045122", "RCV000019428", "45123"]})

clinvar_variant_by_rsid

All ClinVar variation records that reference a dbSNP rsID, with full classifications (an rsID can map to several VCVs — one per alternate allele, e.g. rs121913529 covers KRAS G12D/G12V/G12A). Requires a contact email (Settings → Credentials → Literature access → Contact email) per NCBI E-utilities usage policy. Args: rsid (dbSNP reference SNP ID, e.g. rs7412; case-insensitive, must match rs<digits>), max_records (cap 1-200, default 50). total always carries the true match count and truncated flags a capped listing; total == 0 means ClinVar has no record for the rsID.

FieldTypeRequirement and constraints
rsidstringrequired
max_recordsintegeroptional; default: 50
const result = await host.mcp("variants", "clinvar_variant_by_rsid", {"rsid": "rs121913529", "max_records": 50})

dbsnp_get_rsids

Canonical dbSNP RefSNP records for a batch of rsIDs: GRCh38+GRCh37 placements, alleles, gene context, per-study allele frequencies, and ClinVar cross-references. Requires a contact email (Settings → Credentials → Literature access → Contact email) per NCBI E-utilities usage policy; without one the tool returns {error: 'contact_email_required', message}. Args: rsids (up to 20 rs<digits>, case-insensitive) — each costs one paced NCBI Variation Services request, so large batches take ~1 s per rsID. Returns {n_requested, records, not_found (rs numbers dbSNP doesn't know), not_processed (rsIDs skipped when the wall-clock budget ran out — re-request just those)}. Each record: {rsid, status, create_date, last_update_date, last_update_build_id, n_citations, citations_pmids (capped at 20; citations_truncated flags the cap), variant_type, mane_select_ids, placements, alleles}. status is 'live', 'merged' (record instead carries merged_into — re-query those rsIDs) or 'no_data' (withdrawn/unsupported). placements give 1-based chromosome coordinates with ref/alts per assembly (GRCh38 first, is_primary true). Each alt-allele entry: {allele, ref, spdi (0-based interbase), hgvs, frequencies: [{study, study_version, allele_count, total_count, af}] (ALFA, 1000Genomes, TOPMED, gnomAD...), clinvar: [{rcv_accession, clinical_significances, review_status, last_evaluated_date, disease_names}], genes: [{symbol, gene_id, name, orientation, consequences (SO terms), mane_select: [{transcript_hgvs, protein_spdi}]}]}.

FieldTypeRequirement and constraints
rsidsarray of stringrequired
const result = await host.mcp("variants", "dbsnp_get_rsids", {"rsids": ["rs7412", "rs429358"]})

dbsnp_search_by_region

List dbSNP rsIDs in a genomic window (esearch db=snp positional index — NCBI Variation Services has no region endpoint). Requires a contact email (Settings → Credentials → Literature access → Contact email) per NCBI E-utilities usage policy; without one the tool returns {error: 'contact_email_required', message}. Args: chrom (1-22, X, Y or MT; 'chr' prefix tolerated), start (1-based inclusive), stop (inclusive; span capped at 1 Mb — split larger regions into consecutive windows; dense regions hold many thousands of rsIDs per kb, so keep windows small or raise max_rsids), assembly (which positional index — 'GRCh38' default -> [CPOS], or 'GRCh37' -> [CPOS_GRCH37]; coordinates must be on the chosen assembly), max_rsids (listing cap 1-1000, default 200). Returns {chrom, start, stop, assembly, term (the exact Entrez query used), total (the API's own count), n_returned, truncated, rsids}. truncated is true when total > n_returned — the list is then a prefix in Entrez default order (descending rs number), never a silent truncation. Feed rsIDs (<= 20 at a time) to dbsnp_get_rsids for full records.

FieldTypeRequirement and constraints
chromstringrequired
startintegerrequired
stopintegerrequired
assemblystringoptional; default: "GRCh38"; enum: ["GRCh38", "GRCh37"]
max_rsidsintegeroptional; default: 200
const result = await host.mcp("variants", "dbsnp_search_by_region", {"chrom": "19", "start": 44905000, "stop": 44910000, "assembly": "GRCh38"})

Clinical Trials

Show operations and parameters

search_trials

PRIMARY search over ClinicalTrials.gov. Filter by condition, intervention, sponsor, location, status (e.g. ["RECRUITING"]), phase (["PHASE1".."PHASE4"]) and study_type. condition/intervention/sponsor/location accept Essie query syntax (boolean AND/OR/NOT, "quoted phrases", grouping, automatic synonyms). Page with page_token; set count_total for the total match count. advanced_query merges a raw Essie expression into filter.advanced.

FieldTypeRequirement and constraints
conditionstringoptional
interventionstringoptional
sponsorstringoptional
locationstringoptional
statusarray of stringoptional
phasearray of stringoptional
study_typestringoptional; enum: ["INTERVENTIONAL", "OBSERVATIONAL", "EXPANDED_ACCESS"]
advanced_querystringoptional
page_sizeintegeroptional; default: 10; minimum: 1; maximum: 1000
page_tokenstringoptional
count_totalbooleanoptional; default: false
const result = await host.mcp("clinical-trials", "search_trials", {"condition": "lung cancer", "status": ["RECRUITING"], "phase": ["PHASE3"], "count_total": true, "page_size": 10})

get_trial_details

Get comprehensive details for one trial by NCT id (format "NCT" + 8 digits; a bare number is prefixed, case-insensitive). Returns full eligibility criteria, study design, primary/secondary/other endpoints, all locations, sponsor and collaborators, dates, enrollment, and a results link.

FieldTypeRequirement and constraints
nct_idstringrequired
const result = await host.mcp("clinical-trials", "get_trial_details", {"nct_id": "NCT03661411"})

search_by_sponsor

Find trials sponsored by a company or organization (partial name match, e.g. "Pfizer" matches "Pfizer Inc"). Optionally narrow by condition, phase and status. Set count_total for the total number of trials by the sponsor. Page with page_token.

FieldTypeRequirement and constraints
sponsor_namestringrequired
conditionstringoptional
phasearray of stringoptional
statusarray of stringoptional
page_sizeintegeroptional; default: 10; minimum: 1; maximum: 1000
page_tokenstringoptional
count_totalbooleanoptional; default: false
const result = await host.mcp("clinical-trials", "search_by_sponsor", {"sponsor_name": "Pfizer", "phase": ["PHASE3"], "count_total": true})

search_investigators

Find principal investigators and research sites by condition, institution, location or investigator_name. institution filters on the site facility and takes precedence over location; investigator_name searches OverallOfficialName and ResponsiblePartyInvestigatorFullName. Returns site contacts (names, roles, affiliations, facilities, cities) with their trial NCT ids. page_size caps how many trials are scanned.

FieldTypeRequirement and constraints
conditionstringoptional
institutionstringoptional
locationstringoptional
investigator_namestringoptional
statusarray of stringoptional
page_sizeintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("clinical-trials", "search_investigators", {"condition": "Alzheimer", "institution": "Mayo Clinic", "page_size": 20})

analyze_endpoints

Analyze primary/secondary/other outcome measures (endpoints). Provide ONLY nct_id (single-trial mode) OR condition (aggregate mode across trials); if both are given, nct_id takes precedence. Aggregate mode may be narrowed by phase and start_date_after (YYYY-MM-DD) and scans up to page_size trials. Returns the endpoint lists plus the most common measure names across the analyzed trials.

FieldTypeRequirement and constraints
nct_idstringoptional
conditionstringoptional
phasearray of stringoptional
start_date_afterstringoptional
page_sizeintegeroptional; default: 50; minimum: 1; maximum: 1000
const result = await host.mcp("clinical-trials", "analyze_endpoints", {"nct_id": "NCT03661411"})

search_by_eligibility

Patient-trial matching. DEFAULTS to RECRUITING trials unless status is set. min_age/max_age are the PATIENT's age ("65 Years", "6 Months") and match trials whose age window admits the patient; sex matches trials accepting that sex or all comers; eligibility_keywords searches the inclusion/exclusion criteria text (e.g. "HbA1c > 8", "BRCA mutation", "ECOG 0-1"). At least one of condition, eligibility_keywords, min_age, max_age or sex is required. Page with page_token.

FieldTypeRequirement and constraints
conditionstringoptional
eligibility_keywordsstringoptional
min_agestringoptional
max_agestringoptional
sexstringoptional; enum: ["ALL", "MALE", "FEMALE"]
statusarray of stringoptional
page_sizeintegeroptional; default: 10; minimum: 1; maximum: 1000
page_tokenstringoptional
const result = await host.mcp("clinical-trials", "search_by_eligibility", {"condition": "diabetes", "min_age": "65 Years", "sex": "FEMALE"})

Clinical Genomics

Show operations and parameters

clingen_gene_validity

ClinGen gene-disease validity curations (how strong the evidence is that variation in a gene causes a disease: Definitive/Strong/Moderate/Limited/Disputed/Refuted/No Known Disease Relationship). Omit gene to list all 3,600+ curations.

FieldTypeRequirement and constraints
genestringoptional
const result = await host.mcp("clinical-genomics", "clingen_gene_validity", {"gene": "BRCA2"})

clingen_dosage_sensitivity

ClinGen dosage sensitivity curations: haploinsufficiency and triplosensitivity assertions for genes (and optionally ISCA genomic/CNV regions). A gene symbol or an ISCA region id filters exactly; omit for the full table.

FieldTypeRequirement and constraints
genestringoptional
include_regionsbooleanoptional; default: false
const result = await host.mcp("clinical-genomics", "clingen_dosage_sensitivity", {"gene": "TP53"})

clingen_actionability

ClinGen clinical actionability curations: for disorders associated with a gene, whether early intervention in pre-symptomatic carriers is actionable (intervention/outcome pairs with severity, likelihood, effectiveness, nature-of-intervention component scores and the total score). Gene filter matches any member of multi-gene topics.

FieldTypeRequirement and constraints
genestringoptional
contextstringoptional; default: "both"; enum: ["adult", "pediatric", "both"]
const result = await host.mcp("clinical-genomics", "clingen_actionability", {"gene": "BRCA1", "context": "adult"})

clingen_variant_classifications

ClinGen Evidence Repository (ERepo) expert-panel variant pathogenicity classifications (VCEP interpretations under ACMG criteria). Provide EXACTLY ONE of gene (HGNC symbol), caid (ClinGen canonical allele id, e.g. CA114360), or hgvs (e.g. NM_000277.2:c.1222C>T). Complete retrieval (matchLimit=none).

FieldTypeRequirement and constraints
genestringoptional
caidstringoptional
hgvsstringoptional
const result = await host.mcp("clinical-genomics", "clingen_variant_classifications", {"gene": "BRCA1"})

civic_search_genes

Find CIViC gene records by exact Entrez symbol (e.g. "BRAF"). Fully paginated, count-verified. Use the returned CIViC gene id with civic_gene_variants.

FieldTypeRequirement and constraints
entrez_symbolstringrequired
const result = await host.mcp("clinical-genomics", "civic_search_genes", {"entrez_symbol": "BRAF"})

civic_gene_variants

All variants of one CIViC gene (by CIViC gene id), fully paginated — complete even for genes with hundreds of variants. Sorted by variant id.

FieldTypeRequirement and constraints
gene_idintegerrequired
const result = await host.mcp("clinical-genomics", "civic_gene_variants", {"gene_id": 5})

civic_get_variant

One CIViC variant by its CIViC variant id (aliases, variant types, feature/gene linkage, coordinates for gene variants). Returns found=false if absent.

FieldTypeRequirement and constraints
variant_idintegerrequired
const result = await host.mcp("clinical-genomics", "civic_get_variant", {"variant_id": 12})

civic_search_variants

Search CIViC variants by name substring (e.g. "V600"), optionally scoped to a CIViC gene id. Fully paginated; sorted by variant id.

FieldTypeRequirement and constraints
namestringrequired
gene_idintegeroptional
const result = await host.mcp("clinical-genomics", "civic_search_variants", {"name": "V600", "gene_id": 5})

civic_get_evidence_item

One CIViC evidence item by id: clinical significance of a molecular profile in a disease/therapy context (evidence level A-E, type, direction, significance, rating, disease, therapies, source). Returns found=false if absent.

FieldTypeRequirement and constraints
evidence_idintegerrequired
const result = await host.mcp("clinical-genomics", "civic_get_evidence_item", {"evidence_id": 1409})

civic_search_evidence

Search CIViC evidence items by any combination of filters; fully paginated, count-verified, sorted by ascending evidence id. Enum filters take CIViC GraphQL enum values verbatim (evidence_level "A".."E"; evidence_type PREDICTIVE|PROGNOSTIC|DIAGNOSTIC|PREDISPOSING|ONCOGENIC|FUNCTIONAL; evidence_direction SUPPORTS|DOES_NOT_SUPPORT; status ACCEPTED|SUBMITTED|REJECTED|ALL). Provide at least one filter — no filters walks the entire 10k+ corpus.

FieldTypeRequirement and constraints
disease_namestringoptional
therapy_namestringoptional
evidence_levelstringoptional
evidence_typestringoptional
evidence_directionstringoptional
significancestringoptional
variant_originstringoptional
evidence_ratingintegeroptional
statusstringoptional
molecular_profile_namestringoptional
molecular_profile_idintegeroptional
variant_idintegeroptional
disease_idintegeroptional
therapy_idintegeroptional
phenotype_idintegeroptional
source_idintegeroptional
assertion_idintegeroptional
const result = await host.mcp("clinical-genomics", "civic_search_evidence", {"disease_name": "melanoma", "evidence_level": "A"})

civic_get_assertion

One CIViC assertion by id: an expert-curated summary claim (AMP/ASCO/CAP tier, ACMG/ClinGen codes, FDA companion-test flags) aggregating evidence for a molecular profile in a disease/therapy context. Returns found=false if absent.

FieldTypeRequirement and constraints
assertion_idintegerrequired
const result = await host.mcp("clinical-genomics", "civic_get_assertion", {"assertion_id": 7})

civic_search_assertions

Search CIViC assertions by any combination of filters; fully paginated, count-verified, sorted by ascending assertion id. assertion_type PREDICTIVE|PROGNOSTIC|DIAGNOSTIC|PREDISPOSING|ONCOGENIC; assertion_direction SUPPORTS|DOES_NOT_SUPPORT; amp_level e.g. TIER_I_LEVEL_A; status ACCEPTED|SUBMITTED|REJECTED|ALL. No filters walks the full corpus.

FieldTypeRequirement and constraints
disease_namestringoptional
therapy_namestringoptional
assertion_typestringoptional
assertion_directionstringoptional
significancestringoptional
amp_levelstringoptional
statusstringoptional
molecular_profile_namestringoptional
molecular_profile_idintegeroptional
variant_idintegeroptional
variant_namestringoptional
disease_idintegeroptional
therapy_idintegeroptional
phenotype_idintegeroptional
evidence_idintegeroptional
summarystringoptional
const result = await host.mcp("clinical-genomics", "civic_search_assertions", {"disease_name": "melanoma"})

civic_get_molecular_profile

One CIViC molecular profile by id (variant combination that evidence/assertions attach to), incl. parsed name, score, and component variants. Returns found=false if absent.

FieldTypeRequirement and constraints
mp_idintegerrequired
const result = await host.mcp("clinical-genomics", "civic_get_molecular_profile", {"mp_id": 12})

civic_search_molecular_profiles

Search CIViC molecular profiles by name substring (e.g. "BRAF V600E"). Fully paginated; sorted by id.

FieldTypeRequirement and constraints
namestringrequired
const result = await host.mcp("clinical-genomics", "civic_search_molecular_profiles", {"name": "BRAF V600E"})

civic_search_diseases

Search CIViC disease records by name substring (e.g. "melanoma"). Returns DOIDs + display names; fully paginated; sorted by id.

FieldTypeRequirement and constraints
namestringrequired
const result = await host.mcp("clinical-genomics", "civic_search_diseases", {"name": "melanoma"})

civic_search_therapies

Search CIViC therapy records by name substring (e.g. "vemurafenib"). Returns NCIt ids + names; fully paginated; sorted by id.

FieldTypeRequirement and constraints
namestringrequired
const result = await host.mcp("clinical-genomics", "civic_search_therapies", {"name": "vemurafenib"})

open_targets_graphql

Run an arbitrary GraphQL query against the Open Targets Platform API (targets, diseases, drugs, target-disease association scores, evidence, tractability, safety, known drugs). Introspection queries work for schema discovery. Note knownDrugs was renamed to drugAndClinicalCandidates upstream.

FieldTypeRequirement and constraints
querystringrequired
variablesobjectoptional
const result = await host.mcp("clinical-genomics", "open_targets_graphql", {"query": "query($id: String!){ target(ensemblId: $id){ approvedSymbol associatedDiseases{ count } } }", "variables": {"id": "ENSG00000157764"}})

open_targets_disease_drugs

Known/investigational drugs for a disease (Open Targets Platform) — wraps Disease.drugAndClinicalCandidates. efo_id is a disease ontology id (EFO/MONDO/etc., e.g. "MONDO_0004992").

FieldTypeRequirement and constraints
efo_idstringrequired
sizeintegeroptional; default: 25
const result = await host.mcp("clinical-genomics", "open_targets_disease_drugs", {"efo_id": "MONDO_0004992", "size": 25})

open_targets_disease_targets

Top associated targets for a disease, ranked by Open Targets overall association score — wraps Disease.associatedTargets. efo_id is a disease ontology id (EFO/MONDO/etc.).

FieldTypeRequirement and constraints
efo_idstringrequired
sizeintegeroptional; default: 25
const result = await host.mcp("clinical-genomics", "open_targets_disease_targets", {"efo_id": "MONDO_0004992", "size": 25})

open_targets_drug

Drug details by ChEMBL id (Open Targets Platform) — name, type, maximum clinical stage, and mechanisms of action (target + action type). chembl_id e.g. "CHEMBL1201583".

FieldTypeRequirement and constraints
chembl_idstringrequired
const result = await host.mcp("clinical-genomics", "open_targets_drug", {"chembl_id": "CHEMBL1201583"})

Structures & Interactions

Show operations and parameters

emdb_get_entries

Fetch structured metadata records for EMDB cryo-EM 3D map entries. Accepts accessions as 'EMD-1234', 'emd-1234' or '1234'. Each record carries title, structure determination method (singleParticle / helical / tomography / subtomogramAveraging / electronCrystallography), resolution in Angstrom (null for entries with no reported resolution, e.g. raw tomograms) and the resolution method, deposition/release dates, sample and macromolecule/supramolecule names, fitted PDB model IDs (empty list when no model is fitted), primary citation (journal, year, first author, DOI, PMID), map dimensions and voxel size, and status. Obsolete entries report is_obsolete=true plus superseded_by accessions. Unknown accessions come back as {"emdb_id", "error": "not_found"} — never silently dropped. Metadata only; map volumes are never downloaded.

FieldTypeRequirement and constraints
emdb_idsarray of stringrequired
const result = await host.mcp("structures", "emdb_get_entries", {"emdb_ids": ["EMD-11638", "emd-3061", "1234"]})

emdb_search_entries

Search EMDB with a Solr-style query; complete paged retrieval of compact rows. Query examples: 'title:"apoferritin" AND resolution:[0 TO 1.5]', 'structure_determination_method:"singleParticle"', 'current_status:"REL" AND release_date:[2024-01-01T00:00:00Z TO *]'. Args: query (Solr query string); max_rows (row cap, default 1000). Returns num_found_released (the API's own released-entry count from the facet route — ground truth), rows_retrieved, rows_by_status (REL vs OBS — the search route returns obsolete entries too but they are NOT counted as released), released_complete (true iff every released match was retrieved; false means max_rows truncated the sweep or the counts disagree), and records: compact per-entry rows (emdb_id, title, resolution, structure_determination_method, current_status, release_date, fitted_pdbs) sorted by EMD accession.

FieldTypeRequirement and constraints
querystringrequired
max_rowsintegeroptional; default: 1000
const result = await host.mcp("structures", "emdb_search_entries", {"query": "title:\"apoferritin\" AND resolution:[0 TO 1.5]", "max_rows": 500})

emdb_get_entry_section

Fetch one detailed metadata section for EMDB entries. Sections: 'publications' — primary citation with complete ordered author list, auxiliary citations, external references (PMID/DOI/ISSN/CSD); 'map' — file, format, data type, dimensions, voxel spacing, origin, axis order, cell, voxel statistics, contour levels, symmetry; 'sample' — per-macromolecule records (type, molecular weight, copies, EC number, source organism + NCBI taxid, sequence cross-refs) and per-supramolecule records; 'imaging' — microscope, voltage, electron source, detector, dose, imaging modes, defocus range, magnification, Cs, cryogen, grid/buffer/vitrification conditions (one record per microscopy session — entries can carry several). Args: emdb_ids (accession list, any of EMD-1234/emd-1234/1234); section (one of publications/map/sample/imaging). Unknown accessions are reported with "error": "not_found". Use emdb_get_entries first when you only need the headline record.

FieldTypeRequirement and constraints
emdb_idsarray of stringrequired
sectionstringrequired; enum: ["publications", "map", "sample", "imaging"]
const result = await host.mcp("structures", "emdb_get_entry_section", {"emdb_ids": ["EMD-11638"], "section": "imaging"})

emdb_get_validation

Fetch numeric validation-analysis metrics for EMDB entries. Per entry (from the EMDB /analysis route): Q-score, atom inclusion, recommended/predicted/rawmap contour levels, model/mask volumes, model-map ratio, surface metrics — where the validation pipeline has computed them. available_blocks lists every block the validation service returned; sparse payloads (tomograms, model-free or historical entries) yield explicit nulls. Entries with no validation analysis report has_validation_analysis=false — never silently dropped.

FieldTypeRequirement and constraints
emdb_idsarray of stringrequired
const result = await host.mcp("structures", "emdb_get_validation", {"emdb_ids": ["EMD-11638", "EMD-3061"]})

complexportal_get_complexes

Fetch curated Complex Portal records by CPX accession. Each record: complex AC, recommended/systematic names + synonyms, species and taxid, participant list with stoichiometry (min/max copies), biological role and interactor type, evidence ECO code, GO annotations, and cross-references — the manually curated description of a stable macromolecular complex. Records come back in input order; unknown accessions are listed in not_found rather than silently dropped. For binary interaction evidence (who binds whom in which experiment) use the intact_* tools instead.

FieldTypeRequirement and constraints
complex_acsarray of stringrequired
const result = await host.mcp("structures", "complexportal_get_complexes", {"complex_acs": ["CPX-2158", "CPX-2419"]})

complexportal_search_by_participant

Search Complex Portal for complexes containing a molecule. accession is a participant accession — UniProt (e.g. 'P04637'), ChEBI, or RNAcentral. With participants_only=true (default) the search is field-qualified (pxref:<accession>) so only complexes that actually contain the molecule as a curated participant are returned; with false the bare accession is matched as free text too (descriptions, names), which over-reports but can catch mentions. All result pages are retrieved and the row count is verified against the service-reported total (total_reported == total_retrieved, or the call fails loudly). Hits are compact records (complex_ac, name, species, interactors) sorted by complex accession; fetch full detail with complexportal_get_complexes.

FieldTypeRequirement and constraints
accessionstringrequired
participants_onlybooleanoptional; default: true
const result = await host.mcp("structures", "complexportal_search_by_participant", {"accession": "P69905", "participants_only": true})

intact_fetch_interactions

Retrieve ALL IntAct binary interactions matching a query, MI-score filtered. query is a UniProt accession (e.g. 'P04637'), gene symbol, free text, or any IntAct Solr query. Retrieval is a complete paginated sweep verified against the server-reported total (n_records == total_elements, or the call FAILS LOUDLY — silent truncation is impossible). min_mi_score/max_mi_score filter server-side on the IntAct MI confidence score (0.45 is a common medium-confidence floor); interactor_species filters by species name or taxid (e.g. ["Homo sapiens"] or ["9606"]). Records are slim and structured: interactor pair (IntAct ACs, database identifiers, molecule names, species/taxids), interaction type, detection method (+MI id), experimental roles, host organism, MI score, PubMed id, first author, source database — sorted by DESCENDING MI score. Output lists at most max_records_returned records (records_truncated=true when the full verified sweep was larger; n_records always reports the true total). Large queries (e.g. CFTR ~10k interactions) take a while — narrow with min_mi_score or species when possible.

FieldTypeRequirement and constraints
querystringrequired
min_mi_scorenumberoptional; default: 0
max_mi_scorenumberoptional; default: 1
interactor_speciesarray of stringoptional
max_records_returnedintegeroptional; default: 500
const result = await host.mcp("structures", "intact_fetch_interactions", {"query": "P04637", "min_mi_score": 0.45, "interactor_species": ["Homo sapiens"], "max_records_returned": 200})

intact_get_interactor

Resolve a molecule to its IntAct interactor record(s). query is a UniProt accession, gene symbol, or IntAct interactor AC (e.g. 'EBI-7090529'). Returns ALL matching interactor records with an explicit n_matches — a UniProt accession can resolve to the canonical protein plus chain/isoform interactors, and this tool never silently picks one. Each record: interactor_ac, preferred_identifier, name, species, taxid, interactor_type, and the interaction_count seen by IntAct (useful for sizing an intact_fetch_interactions sweep).

FieldTypeRequirement and constraints
querystringrequired
const result = await host.mcp("structures", "intact_get_interactor", {"query": "P04637"})

intact_get_interaction_details

Full curated detail for ONE IntAct interaction AC (e.g. 'EBI-15635490'). Returns interaction type, host organism, detection method, publication, cross-references, annotations, kinetic/affinity parameters and confidences, plus per-participant records (identifier, species, biological and experimental role, participant detection methods) unless include_participants=false. Get interaction ACs from intact_fetch_interactions records (the interaction_ac field). Unknown ACs return { interaction_ac, error: 'not_found' }.

FieldTypeRequirement and constraints
interaction_acstringrequired
include_participantsbooleanoptional; default: true
const result = await host.mcp("structures", "intact_get_interaction_details", {"interaction_ac": "EBI-15635490", "include_participants": true})

intact_build_network

Build a depth-1 IntAct interaction network around seed proteins. seed_accessions are UniProt accessions. Step 1: a complete, count-verified MI-score-filtered interaction sweep per seed. Step 2: the partners of every seed edge plus the seeds form the node set. Step 3: partner-partner edges are only discoverable by querying the partners themselves, so up to max_interactors_expanded partners are queried (most-connected first, ties by identifier) and edges with BOTH endpoints inside the node set are kept. The expansion block reports exactly which partners were / were not expanded (expansion.complete=false means more partner-partner edges may exist). Output: nodes, edges (with MI score, detection method, PubMed id), per-seed sweep stats. Keep seeds few and min_mi_score >= 0.45 — every expansion is a full paginated sweep.

FieldTypeRequirement and constraints
seed_accessionsarray of stringrequired
min_mi_scorenumberoptional; default: 0.45
max_interactors_expandedintegeroptional; default: 25
interactor_speciesarray of stringoptional
const result = await host.mcp("structures", "intact_build_network", {"seed_accessions": ["P04637", "Q00987"], "min_mi_score": 0.45, "max_interactors_expanded": 25})

pdb_search_structures

Search RCSB PDB entries by attribute filters; paged, capped + flagged. All filters AND together; at least one is required. text is a full-text relevance query ('p53 DNA binding domain'); organism is an exact source-organism lineage name ('Homo sapiens' — matches at any lineage level, so 'Eukaryota' works too); taxonomy_id an NCBI taxid (9606); uniprot_accession finds entries whose polymer entities map to that UniProt ('P04637' -> every p53 structure); experimental_method is the PDB vocabulary ('X-RAY DIFFRACTION', 'ELECTRON MICROSCOPY', 'SOLUTION NMR', ... — case-insensitive, unknown values error with the full list); max_resolution_angstrom keeps entries at or below that resolution; ligand_comp_id requires a bound nonpolymer component by chem-comp id ('ZN', 'ATP', 'HEM'). include_computed_models=true adds computed structure models (e.g. AlphaFold) to the default experimental-only results. Returns total_count (the API's own match total — ground truth), n_retrieved, truncated (true iff total_count > n_retrieved; max_rows, 1..1000, caps retrieval), and records [{pdb_id, score}] in relevance order. Identifiers only — chain to pdb_get_structures for metadata.

FieldTypeRequirement and constraints
textstringoptional
organismstringoptional
taxonomy_idintegeroptional
uniprot_accessionstringoptional
experimental_methodstringoptional
max_resolution_angstromnumberoptional
ligand_comp_idstringoptional
include_computed_modelsbooleanoptional; default: false
max_rowsintegeroptional; default: 100
const result = await host.mcp("structures", "pdb_search_structures", {"uniprot_accession": "P04637", "experimental_method": "X-RAY DIFFRACTION", "max_rows": 50})

pdb_get_structures

Fetch entry-level summaries for PDB entries (batch, max 25 ids). Accepts 4-character PDB ids in any case ('1tup' == '1TUP'; duplicates are de-duplicated). Each record: title, experimental methods, resolution in Angstrom (null for methods without one, e.g. NMR), determination methodology (experimental vs computational), deposit/release/revision dates and status, molecular weight (kDa), assembly and entity counts (protein/DNA/RNA polymer + nonpolymer), bound ligand chem-comp ids, polymer/nonpolymer entity id lists (inputs for pdb_get_entities / pdb_get_ligands), and the primary citation (title, journal, year, authors, PubMed id, DOI). Unknown ids come back as {"pdb_id", "error": "not_found"} — never silently dropped. Metadata only; coordinate files are never downloaded.

FieldTypeRequirement and constraints
pdb_idsarray of stringrequired
const result = await host.mcp("structures", "pdb_get_structures", {"pdb_ids": ["1TUP", "1tup", "6XYZ"]})

pdb_get_entities

Polymer entity details for one PDB entry, incl. UniProt mappings. With entity_ids=null every polymer entity of the entry is fetched, capped at 25 with truncated=true and n_polymer_entities reporting the entry's true count (large assemblies like ribosomes carry 50+ — get the full id list from pdb_get_structures' polymer_entity_ids and page with explicit subsets like ["26", "27"]); with an explicit entity_ids subset the entry total is not fetched, so n_polymer_entities is null; an explicit entity_ids list larger than 25 errors. Each record: description, polymer type (Protein / DNA / RNA), sequence length, mutation count, deposited copies, chain ids (asym + author), source organisms with taxids, UniProt accessions with per-entity sequence coverage (SIFTS), and UniProt-aligned regions (entity-seq vs reference-seq coordinates). Unknown entity ids are listed in not_found; an unknown entry id errors. include_sequences=true adds the canonical one-letter sequence per entity; if the combined sequences exceed max_bytes (default 400000) they are omitted and sequences_omitted explains why — metadata always survives.

FieldTypeRequirement and constraints
pdb_idstringrequired
entity_idsarray of stringoptional
include_sequencesbooleanoptional; default: false
max_bytesintegeroptional; default: 400000
const result = await host.mcp("structures", "pdb_get_entities", {"pdb_id": "1TUP", "include_sequences": true})

pdb_get_ligands

Bound ligands (nonpolymer components) of one PDB entry, with chemistry. Walks the entry's nonpolymer entities and resolves each chemical component: per ligand — entity id, chem-comp id ('ZN', 'ATP'), description, deposited copy count, author chain ids, and a chem_comp block (name, formula, formula weight, formal charge, component type, InChIKey, stereo SMILES). Waters are not nonpolymer entities in the PDB data model and never appear. Entries with no ligands return ligands: []. n_nonpolymer_entities is the entry's true count; truncated=true when it exceeds max_ligands (clamped to 1..25, which bounds the request budget) — never silently dropped. Entities/components the data API no longer serves are reported inline with "error": "not_found" (partial results, not an aborted call). An unknown entry id errors.

FieldTypeRequirement and constraints
pdb_idstringrequired
max_ligandsintegeroptional; default: 25
const result = await host.mcp("structures", "pdb_get_ligands", {"pdb_id": "1TUP"})

alphafold_get_prediction

AlphaFold DB predicted-structure metadata for one UniProt accession. Returns has_model, n_models and per-model records. A single accession can carry several models (canonical + isoforms like 'P04637-9', and community providers beyond the Google DeepMind monomer pipeline — provider_id / tool_used identify them). Each model: entry id, UniProt annotation (id, description, gene, organism, taxid, reviewed flags), sequence coordinates and length, global pLDDT (global_plddt, 0-100) plus the fraction of residues per pLDDT confidence bin (very_low <50, low 50-70, confident 70-90, very_high >90), model version info and creation date, and download URLs (cif/bcif/pdb coordinates, PAE JSON + image, per-residue pLDDT JSON, MSA, AlphaMissense CSV where available) — URLs only, payloads are never downloaded; fetch them yourself if needed. Accessions without a prediction return has_model=false (not an error); malformed identifiers return an explicit error field. include_sequence=true adds the model sequence (protein one-letter).

FieldTypeRequirement and constraints
uniprot_accessionstringrequired
include_sequencebooleanoptional; default: false
const result = await host.mcp("structures", "alphafold_get_prediction", {"uniprot_accession": "P04637"})

alphafold_check_coverage

Batch AlphaFold DB coverage check (max 40 unique UniProt accessions). Blank entries and duplicates are stripped before the batch cap applies, and disclosed: n_requested == n_unique + n_blank_skipped + n_duplicate_skipped always reconciles. One compact record per unique accession, in input order: has_model, n_models, and the primary (first-listed) model's model_entity_id, latest_version, global_plddt and sequence_length. Accessions with no prediction report has_model=false; malformed ones carry an explicit error field — never silently dropped. Use to triage which proteins of a set have usable predicted structures before pulling full records with alphafold_get_prediction.

FieldTypeRequirement and constraints
uniprot_accessionsarray of stringrequired
const result = await host.mcp("structures", "alphafold_check_coverage", {"uniprot_accessions": ["P04637", "P38398", "Q9Y6K9"]})

ChEMBL

Show operations and parameters

Search ChEMBL chemical compounds by name (default), ChEMBL id, or molecular structure. By name: case-insensitive synonym substring match (falls back to a preferred-name match). By chembl_id: direct record lookup. By smiles: Tanimoto similarity search when similarity_threshold is set, else a substructure search (structure walks are capped and disclose walk_truncated/upstream_total). Optional max_phase filters by clinical stage. Pass at least one of name, chembl_id, or smiles. Use drug_search instead when searching by therapeutic indication.

FieldTypeRequirement and constraints
namestringoptional
chembl_idstringoptional
smilesstringoptional
similarity_thresholdintegeroptional; minimum: 70; maximum: 100
max_phaseintegeroptional; enum: [0, 1, 2, 3, 4]
limitintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("chembl", "compound_search", {"name": "aspirin", "limit": 5})

Search approved drugs and clinical candidates by therapeutic indication (EFO term, partial match). Joins drug_indication rows to distinct parent molecules, then to molecule records and withdrawal/black-box warnings. only_approved restricts to phase 4. Optional post-filters molecule_chembl_id, drug_name (preferred-name substring), and max_phase (>=) narrow the joined set. Use compound_search for name/id/structure lookups.

FieldTypeRequirement and constraints
indicationstringrequired
drug_namestringoptional
molecule_chembl_idstringoptional
max_phaseintegeroptional; enum: [0, 1, 2, 3, 4]
only_approvedbooleanoptional; default: false
limitintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("chembl", "drug_search", {"indication": "hypertension", "only_approved": true, "limit": 10})

get_admet

Retrieve ChEMBL calculated molecular properties for drug-likeness / ADMET assessment of one molecule (ALogP, molecular weight, PSA, HBA/HBD, rotatable bonds, aromatic rings, heavy atoms, Rule-of-5 violations, Rule-of-3 pass, QED, molecular formula). These are computed from structure, not experimental measurements.

FieldTypeRequirement and constraints
molecule_chembl_idstringrequired
const result = await host.mcp("chembl", "get_admet", {"molecule_chembl_id": "CHEMBL25"})

get_bioactivity

Retrieve ChEMBL bioactivity measurements (IC50, Ki, Kd, EC50, ...) for compound-target interactions. Filter by molecule_chembl_id and/or target_chembl_id, activity_type (standard_type), a pChEMBL floor (min_pchembl), a standard_value range (min_value/max_value), and unit (standard_units). Returns one page ordered by activity_id with a most-potent summary.

FieldTypeRequirement and constraints
molecule_chembl_idstringoptional
target_chembl_idstringoptional
activity_typestringoptional; enum: ["IC50", "EC50", "Ki", "Kd", "AC50", "GI50", "ED50", "Potency"]
min_pchemblnumberoptional; minimum: 0; maximum: 14
min_valuenumberoptional
max_valuenumberoptional
unitstringoptional; enum: ["nM", "uM", "mM", "pM", "M"]
limitintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("chembl", "get_bioactivity", {"molecule_chembl_id": "CHEMBL25", "activity_type": "IC50", "limit": 10})

get_mechanism

Retrieve ChEMBL mechanism-of-action records for approved drugs and clinical candidates. Filter by molecule_chembl_id, target_chembl_id, and/or action_type. When a molecule id yields nothing, retries against the parent molecule so salt-form ids resolve. Returns one page ordered by mec_id with an action-type summary.

FieldTypeRequirement and constraints
molecule_chembl_idstringoptional
target_chembl_idstringoptional
action_typestringoptional; enum: ["INHIBITOR", "AGONIST", "ANTAGONIST", "BLOCKER", "MODULATOR", "OPENER", "ACTIVATOR", "POSITIVE ALLOSTERIC MODULATOR", "NEGATIVE ALLOSTERIC MODULATOR", "PARTIAL AGONIST", "INVERSE AGONIST"]
limitintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("chembl", "get_mechanism", {"molecule_chembl_id": "CHEMBL25"})

Search ChEMBL biological targets (proteins, complexes, families, organisms). Filter by target_chembl_id, gene_symbol (exact component-synonym match), target_name (preferred-name substring), organism (substring), and/or target_type. Each result carries its components with UniProt accessions, a gene_symbol, and bounded cross-reference lists.

FieldTypeRequirement and constraints
target_namestringoptional
gene_symbolstringoptional
target_chembl_idstringoptional
organismstringoptional
target_typestringoptional; enum: ["SINGLE PROTEIN", "PROTEIN COMPLEX", "PROTEIN FAMILY", "ORGANISM", "TISSUE", "CELL-LINE", "NUCLEIC-ACID", "SUBCELLULAR"]
limitintegeroptional; default: 20; minimum: 1; maximum: 1000
const result = await host.mcp("chembl", "target_search", {"gene_symbol": "EGFR", "organism": "Homo sapiens", "limit": 5})

bioRxiv

Show operations and parameters

get_categories

List all 27 bioRxiv subject categories and their API-compatible slugs (e.g. "cancer biology" -> "cancer_biology"). Use before search_preprints to discover valid category values.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("biorxiv", "get_categories", {})

search_preprints

Search bioRxiv/medRxiv preprints by date and (optionally) category. Use exactly ONE search method: date_from+date_to, recent_days (last N days), or recent_count (N most recent within a 90-day window); with none, the last 60 days. There is NO keyword/text search. cursor paginates. Returns DOI, title, authors, date, category, version, and a 200-char abstract preview.

FieldTypeRequirement and constraints
serverstringoptional; default: "biorxiv"; enum: ["biorxiv", "medrxiv"]
categorystringoptional; enum: ["animal behavior and cognition", "biochemistry", "bioengineering", "bioinformatics", "biophysics", "cancer biology", "cell biology", "clinical trials", "developmental biology", "ecology", "epidemiology", "evolutionary biology", "genetics", "genomics", "immunology", "microbiology", "molecular biology", "neuroscience", "paleontology", "pathology", "pharmacology and toxicology", "physiology", "plant biology", "scientific communication and education", "synthetic biology", "systems biology", "zoology"]
date_fromstringoptional
date_tostringoptional
recent_daysintegeroptional; minimum: 1
recent_countintegeroptional; minimum: 1
limitintegeroptional; default: 10; minimum: 1; maximum: 100
cursorintegeroptional; default: 0; minimum: 0
const result = await host.mcp("biorxiv", "search_preprints", {"recent_days": 30, "category": "neuroscience", "limit": 20})

get_preprint

Get complete metadata for one preprint by DOI (bare "10.1101/..." or a full https://doi.org/ URL). Uses the latest version. Returns title, authors, corresponding author + institution, full abstract, category, license, version, JATS XML, funding, published journal DOI (if linked), PDF and web URLs, and version count. Preprints are NOT peer-reviewed.

FieldTypeRequirement and constraints
doistringoptional
serverstringoptional; default: "biorxiv"; enum: ["biorxiv", "medrxiv"]
const result = await host.mcp("biorxiv", "get_preprint", {"doi": "10.1101/339747"})

search_published_preprints

Find preprints that were later published in peer-reviewed journals (preprint -> journal-article links). Same ONE-OF search methods as search_preprints (date_from+date_to / recent_days / recent_count). include_details=false returns a compact summary. publisher filters by journal DOI prefix (e.g. "10.1038" for Nature) via the bioRxiv-only /publisher route.

FieldTypeRequirement and constraints
serverstringoptional; default: "biorxiv"; enum: ["biorxiv", "medrxiv"]
publisherstringoptional
include_detailsbooleanoptional; default: true
date_fromstringoptional
date_tostringoptional
recent_daysintegeroptional; minimum: 1
recent_countintegeroptional; minimum: 1
limitintegeroptional; default: 10; minimum: 1; maximum: 100
cursorintegeroptional; default: 0; minimum: 0
const result = await host.mcp("biorxiv", "search_published_preprints", {"publisher": "10.1038", "date_from": "2024-01-01", "date_to": "2024-01-05", "limit": 10})

search_by_funder

Find preprints acknowledging a funder, identified by ROR id (9-char, e.g. "021nxhr62" for NIH; a full https://ror.org/ URL is also accepted). Requires an explicit date_from + date_to; funder metadata begins 2025-04-10. Optional category filter. cursor paginates. Same compact result shape as search_preprints.

FieldTypeRequirement and constraints
funder_ror_idstringoptional
date_fromstringoptional
date_tostringoptional
serverstringoptional; default: "biorxiv"; enum: ["biorxiv", "medrxiv"]
categorystringoptional; enum: ["animal behavior and cognition", "biochemistry", "bioengineering", "bioinformatics", "biophysics", "cancer biology", "cell biology", "clinical trials", "developmental biology", "ecology", "epidemiology", "evolutionary biology", "genetics", "genomics", "immunology", "microbiology", "molecular biology", "neuroscience", "paleontology", "pathology", "pharmacology and toxicology", "physiology", "plant biology", "scientific communication and education", "synthetic biology", "systems biology", "zoology"]
limitintegeroptional; default: 10; minimum: 1; maximum: 100
cursorintegeroptional; default: 0; minimum: 0
const result = await host.mcp("biorxiv", "search_by_funder", {"funder_ror_id": "021nxhr62", "date_from": "2025-04-10", "date_to": "2025-05-10", "limit": 10})

get_content_statistics

bioRxiv submission statistics over all history — new vs revised paper counts per period, with running cumulative totals. interval is "monthly" (default) or "yearly".

FieldTypeRequirement and constraints
intervalstringoptional; default: "monthly"; enum: ["monthly", "yearly"]
const result = await host.mcp("biorxiv", "get_content_statistics", {"interval": "yearly"})

get_usage_statistics

bioRxiv usage/engagement statistics over all history — abstract views, full-text views, and PDF downloads per period, with running cumulative totals. interval is "monthly" (default) or "yearly".

FieldTypeRequirement and constraints
intervalstringoptional; default: "monthly"; enum: ["monthly", "yearly"]
const result = await host.mcp("biorxiv", "get_usage_statistics", {"interval": "yearly"})

Drug Regulatory

Show operations and parameters

search_drug_applications

Search Drugs@FDA applications (NDA/ANDA/BLA) by any combination of exact-phrase filters (brand, generic, active_ingredient, sponsor, marketing_status, dosage_form, route, pharm_class). generic and pharm_class query the harmonized openfda block (absent on older applications, so silently skipped there). A broad search returns the first max_records with the true total and truncated=true; to page beyond ~26,000 records, narrow with submission_date_from/to.

FieldTypeRequirement and constraints
brandstringoptional
genericstringoptional
active_ingredientstringoptional
sponsorstringoptional
marketing_statusstringoptional; enum: ["Prescription", "Over-the-counter", "Discontinued", "None (Tentative Approval)"]
dosage_formstringoptional
routestringoptional
pharm_classstringoptional
pharm_class_typestringoptional; enum: ["epc", "moa", "cs", "pe"]
search_typestringoptional; default: "and"; enum: ["and", "or"]
submission_date_fromstringoptional
submission_date_tostringoptional
raw_searchstringoptional
max_recordsintegeroptional; default: 50
const result = await host.mcp("drug-regulatory", "search_drug_applications", {"generic": "ATORVASTATIN CALCIUM", "marketing_status": "Prescription", "max_records": 25})

get_drug_application

Fetch one Drugs@FDA application by its number (e.g. "NDA020702", "ANDA076543", "BLA125514"). Returns the full record — sponsor, products (brand, active ingredients + strengths, dosage form, route, marketing status, TE code), complete submissions history, and harmonized openfda fields when present.

FieldTypeRequirement and constraints
application_numberstringrequired
const result = await host.mcp("drug-regulatory", "get_drug_application", {"application_number": "NDA020702"})

count_drug_applications

Aggregate Drugs@FDA bucket counts over one field, optionally narrowed by the same filters as search_drug_applications. count_field accepts friendly names (sponsor_name, application_number, dosage_form, route, marketing_status, te_code, pharm_class_epc/moa/cs/pe) or a raw openFDA field path (append .exact yourself for analyzed fields).

FieldTypeRequirement and constraints
count_fieldstringrequired
brandstringoptional
genericstringoptional
active_ingredientstringoptional
sponsorstringoptional
marketing_statusstringoptional
dosage_formstringoptional
routestringoptional
pharm_classstringoptional
pharm_class_typestringoptional; enum: ["epc", "moa", "cs", "pe"]
search_typestringoptional; default: "and"; enum: ["and", "or"]
submission_date_fromstringoptional
submission_date_tostringoptional
raw_searchstringoptional
max_bucketsintegeroptional; default: 100
const result = await host.mcp("drug-regulatory", "count_drug_applications", {"count_field": "marketing_status"})

get_drug_statistics

Corpus-level Drugs@FDA statistics in one call — total applications, marketing-status split, top dosage forms and routes (with distinct counts), and top sponsors by application count.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("drug-regulatory", "get_drug_statistics", {})

list_pharmacologic_classes

Enumerate pharmacologic classes with their application counts, counted over the harmonized openfda.pharm_class_<type> block. Counts reflect only applications carrying that block.

FieldTypeRequirement and constraints
class_typestringoptional; default: "epc"; enum: ["epc", "moa", "cs", "pe"]
max_bucketsintegeroptional; default: 100
const result = await host.mcp("drug-regulatory", "list_pharmacologic_classes", {"class_type": "epc", "max_buckets": 50})

get_generic_equivalents

Find generic equivalents of a brand drug: resolve the brand to its reference application(s), extract the exact active-ingredient name set(s), then return every Drugs@FDA application with a product whose active-ingredient set matches (including TE codes and marketing status).

FieldTypeRequirement and constraints
brandstringrequired
const result = await host.mcp("drug-regulatory", "get_generic_equivalents", {"brand": "Lipitor"})

search_drug_labels

Retrieve FDA drug product labels (SPL) by ingredient/name/route with targeted section extraction. Filters (active_ingredient, generic_name, brand_name, route, product_type) hit the openfda label block; set exact to query the non-analyzed .exact variants. Pass sections to extract raw openFDA label sections instead of the default structured record. raw_search is mutually exclusive with the mapped filters.

FieldTypeRequirement and constraints
active_ingredientstringoptional
generic_namestringoptional
brand_namestringoptional
routestringoptional
product_typestringoptional; enum: ["HUMAN PRESCRIPTION DRUG", "HUMAN OTC DRUG"]
exactbooleanoptional; default: false
raw_searchstringoptional
sectionsarray of stringoptional
max_recordsintegeroptional; default: 25
const result = await host.mcp("drug-regulatory", "search_drug_labels", {"brand_name": "Tylenol", "max_records": 5})

Human Genetics

Show operations and parameters

gwas_associations_for_variant

GWAS Catalog associations reported for one variant (rsID), most significant first. Args: rs_id (dbSNP rsID e.g. rs7412 APOE or rs699 AGT; must be the catalog's current rsID — merged/retired IDs may return zero rows rather than an error); max_records (output cap default 500; trait-hub variants can carry 1000+ associations; rows are server-sorted by p-value ascending, so a capped result is the top-signal prefix). Returns {rs_id, api_total, returned, truncated, associations}. api_total is the catalog's own total; truncated flags a capped fetch. Each association row: {association_id, p_value, pvalue_mantissa, pvalue_exponent, pvalue_description, or_value, beta, ci_lower, ci_upper, range, risk_frequency, snp_effect_alleles, rs_ids, locations, mapped_genes, efo_traits:[{efo_id, efo_trait}], bg_efo_traits, reported_trait, multi_snp_haplotype, snp_interaction, study_accession_id, pubmed_id, first_author}. or_value and beta are mutually exclusive per row (binary vs quantitative); p_value of 0.0 means p < ~1e-308 (use mantissa/exponent).

FieldTypeRequirement and constraints
rs_idstringrequired
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "gwas_associations_for_variant", {"rs_id": "rs7412", "max_records": 100})

gwas_associations_for_gene

GWAS Catalog associations whose variants are MAPPED to a gene (catalog's Ensembl pipeline mapping, not author-reported), most significant first. Args: gene_symbol (HGNC symbol, exact match, e.g. PCSK9, APOE; case-sensitive upstream — pass canonical uppercase; intergenic variants map to flanking genes, so rows may sit outside the gene body); max_records (cap default 500; rows server-sorted by p-value ascending). Returns {gene_symbol, api_total, returned, truncated, associations} with the same row shape as gwas_associations_for_variant. A nonexistent symbol returns api_total=0, not an error.

FieldTypeRequirement and constraints
gene_symbolstringrequired
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "gwas_associations_for_gene", {"gene_symbol": "PCSK9", "max_records": 100})

gwas_associations_for_trait

GWAS Catalog associations annotated to one EFO trait, most significant first. Args: efo_id (ontology term short form as used by the catalog, e.g. MONDO_0005010, EFO_0004340, HP_0003124; the catalog migrated many historical EFO ids to MONDO/HP — resolve current ids with gwas_search_traits first; pass exactly one of efo_id/efo_trait); efo_trait (exact trait LABEL alternative); max_records (cap default 500; rows p-value ascending). Returns {efo_id|efo_trait, api_total, returned, truncated, associations} with the same row shape as gwas_associations_for_variant. An unknown id/label returns api_total=0, not an error.

FieldTypeRequirement and constraints
efo_idstringoptional
efo_traitstringoptional
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "gwas_associations_for_trait", {"efo_id": "MONDO_0005010", "max_records": 100})

gwas_search_traits

Search GWAS Catalog EFO trait annotations by label substring — the entry point for resolving a disease/phenotype name to the ontology ids that gwas_associations_for_trait / gwas_search_studies take. Args: query (case-insensitive substring of the trait label, e.g. "coronary" matches coronary artery disorder MONDO_0005010 etc.; the catalog mixes EFO, MONDO, HP and OBA ids — don't assume an EFO_ prefix); max_records (cap default 500). Returns {query, api_total, returned, truncated, efo_traits}; each row {efo_id, efo_trait, uri} sorted by label. Count-verified against the catalog's own total when not capped.

FieldTypeRequirement and constraints
querystringrequired
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "gwas_search_traits", {"query": "coronary", "max_records": 50})

gwas_search_studies

Search GWAS Catalog studies by trait annotation or publication. Args: efo_id (ontology short form, e.g. MONDO_0005010, resolve via gwas_search_traits; filters combine AND — usually pass one); efo_trait (exact trait label alternative); pubmed_id (PubMed ID of the study's publication, e.g. 38714703); max_records (cap default 500). Returns {filters, api_total, returned, truncated, studies}; each study row {accession_id, disease_trait, efo_traits, bg_efo_traits, pubmed_id, initial_sample_size, replication_sample_size, discovery_ancestry, replication_ancestry, genotyping_technologies, platforms, cohort, full_summary_stats_available, imputed, gxe, gxg}. Count-verified against the catalog total when not capped. At least one filter is required (the unfiltered catalog is ~90k studies).

FieldTypeRequirement and constraints
efo_idstringoptional
efo_traitstringoptional
pubmed_idstringoptional
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "gwas_search_studies", {"efo_id": "MONDO_0005010", "max_records": 50})

gwas_get_study

Fetch one GWAS Catalog study by its GCST accession. Args: accession_id (study accession, e.g. GCST90841394; listed in every association row as study_accession_id and in study search results). Returns {found, accession_id, study} where study is the same row shape as gwas_search_studies (null when the accession is unknown).

FieldTypeRequirement and constraints
accession_idstringrequired
const result = await host.mcp("human-genetics", "gwas_get_study", {"accession_id": "GCST90841394"})

gwas_get_variant

Fetch one GWAS Catalog variant record (position, mapped genes, consequence) by rsID — lighter than pulling its associations. Args: rs_id (dbSNP rsID e.g. rs7412). Returns {found, rs_id, variant}; variant is {rs_id, merged, functional_class, most_severe_consequence, alleles (e.g. "C/T (forward)"), mapped_genes, locations:[{chromosome, position, region}], last_update_date} — positions GRCh38 — or null when the rsID is not in the catalog. merged=1 means the rsID was merged into another record upstream.

FieldTypeRequirement and constraints
rs_idstringrequired
const result = await host.mcp("human-genetics", "gwas_get_variant", {"rs_id": "rs7412"})

eqtl_list_datasets

List eQTL Catalogue datasets (one dataset = one study x tissue/cell type x quantification method). Args: study_label (exact study name, e.g. GTEx, Alasoo_2018, BLUEPRINT); tissue_label (exact tissue/cell-type label, e.g. liver, macrophage, LCL — lowercase in the catalogue); quant_method (ge=gene expression, exon, tx, txrev, microarray, leafcutter, aptamer=plasma protein; for conventional gene-level eQTLs use ge); max_records (cap default 1000; the full unfiltered catalogue is ~760 datasets). Returns {filters, returned, truncated, datasets} sorted by dataset_id; each {dataset_id (QTD...), study_id (QTS...), study_label, sample_group, tissue_id, tissue_label, condition_label, quant_method, sample_size}. The API publishes no total count; truncated=false proves the listing is complete.

FieldTypeRequirement and constraints
study_labelstringoptional
tissue_labelstringoptional
quant_methodstringoptional
max_recordsintegeroptional; default: 1000
const result = await host.mcp("human-genetics", "eqtl_list_datasets", {"study_label": "Alasoo_2018", "quant_method": "ge"})

eqtl_associations

Molecular-QTL association rows from one eQTL Catalogue dataset, filtered by gene, variant or region. Args: dataset_id (QTD accession from eqtl_list_datasets, e.g. QTD000266); gene_id (unversioned Ensembl gene ID e.g. ENSG00000130203 APOE; at least one of gene_id/rsid/variant/pos is required); rsid (dbSNP rsID); variant (eQTL Catalogue variant string chr19_44908822_C_T, chr-prefixed underscore GRCh38); pos (genomic window chromosome:start-end GRCh38 no chr prefix, e.g. 19:44900000-44920000); nlog10p_min (significance floor: only rows with -log10(p) >= this, applied upstream); max_records (cap default 1000 = one page). Returns {dataset_id, filters, returned, truncated, associations}; each row {molecular_trait_id, gene_id, variant, rsid, chromosome, position, ref, alt, type, beta, se, pvalue, nlog10p, maf, ac, an, r2, median_tpm}. Rows cover ONLY the cis window the dataset tested (±1 Mb of each gene); empty means "not tested / not present". No total count is published: truncated=false proves exhaustion, truncated=true means the cap was hit.

FieldTypeRequirement and constraints
dataset_idstringrequired
gene_idstringoptional
rsidstringoptional
variantstringoptional
posstringoptional
nlog10p_minnumberoptional
max_recordsintegeroptional; default: 1000
const result = await host.mcp("human-genetics", "eqtl_associations", {"dataset_id": "QTD000266", "gene_id": "ENSG00000130203", "nlog10p_min": 2})

phewas_instances

List the public PheWeb PheWAS portals this server can query, with genome build and capability registry. Returns {instances:{key:{label, base_url, genome_build, capabilities, notes}}}. capabilities name the endpoints each instance exposes: variant (phewas_variant), gene (phewas_finngen_gene), phenotypes (phewas_list_phenotypes), autocomplete (phewas_search_phenotypes). NOTE the build split: FinnGen R12 variant IDs are GRCh38; BioBank Japan (pheweb.jp) is GRCh37/hg19 — liftover coordinates before cross-querying.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("human-genetics", "phewas_instances", {})

phewas_variant

PheWAS for one variant: its association statistics against every phenotype in a biobank PheWeb portal, most significant first. Args: instance (finngen FinnGen R12 GRCh38, or bbj BioBank Japan GRCh37; variant coords MUST be on the instance's build); variant (chrom-pos-ref-alt, :/_ separators and chr prefix tolerated, e.g. 19-44908822-C-T APOE rs7412 GRCh38/finngen or 1-55505647-G-T PCSK9 rs11591147 GRCh37/bbj); max_phenos (cap default 200; FinnGen returns ~2470 rows; sorted by p-value ascending before capping). Returns {instance, genome_build, variant, variant_meta, total, returned, truncated, phenotypes}; variant_meta {chrom, pos, ref, alt, rsids, nearest_genes, gnomad (FinnGen only)}. Each phenotype row {phenocode, phenostring, category, pval, mlogp, beta, sebeta, af|maf, maf_case, maf_control, n_cases, n_controls, n_samples} (unpublished fields null; BBJ rows have af, FinnGen rows have maf triplets + mlogp). Unknown variants raise a not-found error.

FieldTypeRequirement and constraints
instancestringrequired; enum: ["finngen", "bbj"]
variantstringrequired
max_phenosintegeroptional; default: 200
const result = await host.mcp("human-genetics", "phewas_variant", {"instance": "finngen", "variant": "19-44908822-C-T", "max_phenos": 50})

phewas_finngen_gene

Gene-level PheWAS from FinnGen R12: for every disease endpoint, the best-associated variant in the gene region, most significant first. Args: gene_symbol (HGNC symbol e.g. PCSK9, APOE; unknown symbols raise a not-found error); max_phenos (cap default 200; FinnGen has 2470 endpoints, one row each; sorted by p-value ascending before capping). Returns {instance:"finngen", genome_build:"GRCh38", gene_symbol, total, returned, truncated, phenotypes}; each row is the phewas_variant row shape plus variant:{chrom, pos, ref, alt, varid, rsids} — the top variant for that endpoint in this gene's region (region != gene body; PheWeb pads gene boundaries). Most rows are null results (pval1) — the per-endpoint BEST variant is still reported; filter by pval yourself for significant hits.

FieldTypeRequirement and constraints
gene_symbolstringrequired
max_phenosintegeroptional; default: 200
const result = await host.mcp("human-genetics", "phewas_finngen_gene", {"gene_symbol": "PCSK9", "max_phenos": 50})

phewas_list_phenotypes

Complete phenotype (disease endpoint) catalogue of a PheWeb instance, with case/control counts. Args: instance (currently only finngen exposes this endpoint; BBJ does not — use phewas_search_phenotypes there); max_records (cap default 3000 > FinnGen's ~2470 endpoints, so the default returns the complete catalogue). Returns {instance, total, returned, truncated, phenotypes} sorted by phenocode; each row {phenocode (e.g. "T2D"), phenostring, category, num_cases, num_controls, num_gw_significant (count of genome-wide-significant loci for that endpoint)}.

FieldTypeRequirement and constraints
instancestringoptional; default: "finngen"; enum: ["finngen"]
max_recordsintegeroptional; default: 3000
const result = await host.mcp("human-genetics", "phewas_list_phenotypes", {"instance": "finngen", "max_records": 3000})

phewas_search_phenotypes

Search a PheWeb instance's phenotypes (and entities) by name — the entry point for resolving a disease name to a phenocode. Args: query (free-text phenotype query e.g. "diabetes", "asthma"; matches phenotype names/codes; some instances also match gene names and rsIDs); instance (finngen default or bbj — both expose autocomplete); max_records (cap default 500; autocomplete responses are short lists, rarely capped). Returns {instance, query, total, returned, truncated, matches}; each match {display, phenocode, url}. Use the phenocode with phewas_list_phenotypes rows or the instance website; BBJ display strings embed the code in parentheses.

FieldTypeRequirement and constraints
querystringrequired
instancestringoptional; default: "finngen"; enum: ["finngen", "bbj"]
max_recordsintegeroptional; default: 500
const result = await host.mcp("human-genetics", "phewas_search_phenotypes", {"query": "diabetes", "instance": "finngen"})

Expression

Show operations and parameters

gtex_tissue_sites

List all tissue sites with metadata for a pinned GTEx release (54 in gtex_v8): sample counts, eGene/sGene counts, colour codes, and UBERON ontology ids.

FieldTypeRequirement and constraints
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_tissue_sites", {"dataset_id": "gtex_v8"})

gtex_dataset_info

List all GTEx dataset releases with metadata: datasetId, GENCODE version, genome build, dbSNP build, and sample/subject/tissue counts.

FieldTypeRequirement and constraints
dataset_idstringoptional
organization_namestringoptional
const result = await host.mcp("expression", "gtex_dataset_info", {})

gtex_sample_info

Sample and donor metadata for a pinned GTEx release, optionally filtered by tissue_site_detail_id, data_type (e.g. RNASEQ, WGS), or subject_id. Paged and count-verified; an unfiltered call matches tens of thousands of samples, so filter or set max_samples.

FieldTypeRequirement and constraints
tissue_site_detail_idstringoptional
data_typestringoptional
subject_idstringoptional
max_samplesintegeroptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_sample_info", {"tissue_site_detail_id": "Liver", "data_type": "RNASEQ", "max_samples": 100})

gtex_resolve_genes

Resolve gene symbols or unversioned Ensembl ids to versioned GENCODE ids for a pinned release, e.g. GAPDH -> ENSG00000111640.14. Feed the ids to the expression / eQTL tools.

FieldTypeRequirement and constraints
genesarray of stringrequired
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_resolve_genes", {"genes": ["GAPDH", "BRCA2"]})

gtex_median_expression

Median gene expression (TPM) for one or more VERSIONED GENCODE ids across tissues (omit tissues for all). Paged and count-verified over (gene, tissue) rows.

FieldTypeRequirement and constraints
gencode_idsarray of stringrequired
tissue_site_detail_idsarray of stringoptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_median_expression", {"gencode_ids": ["ENSG00000111640.14"]})

gtex_expression_summary

Summarize a gene’s expression across ALL tissues ranked by descending median TPM. Accepts a symbol or Ensembl id and auto-resolves it to a versioned GENCODE id first.

FieldTypeRequirement and constraints
genestringrequired
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_expression_summary", {"gene": "GAPDH"})

gtex_gene_expression

Sample-level (not aggregated) expression TPM arrays for one VERSIONED GENCODE id, per tissue (omit tissues for all). Returns the full per-sample TPM array and n_samples for each tissue.

FieldTypeRequirement and constraints
gencode_idstringrequired
tissue_site_detail_idsarray of stringoptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_gene_expression", {"gencode_id": "ENSG00000111640.14", "tissue_site_detail_ids": ["Whole_Blood"]})

gtex_top_expressed_genes

Top-n genes by median TPM in one tissue, using the API-side ranking. filter_mt_gene (default true) drops mitochondrial genes from the ranking.

FieldTypeRequirement and constraints
tissue_site_detail_idstringrequired
nintegeroptional; default: 100
filter_mt_genebooleanoptional; default: true
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_top_expressed_genes", {"tissue_site_detail_id": "Whole_Blood", "n": 20})

gtex_eqtl_genes

All eGenes (genes with ≥1 significant cis-eQTL) for a tissue. Walked page-by-page and count-verified (e.g. Pancreas gtex_v8 = 9,660). max_genes caps how many rows are returned.

FieldTypeRequirement and constraints
tissue_site_detail_idstringrequired
max_genesintegeroptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_eqtl_genes", {"tissue_site_detail_id": "Pancreas", "max_genes": 100})

gtex_single_tissue_eqtls

Significant single-tissue cis-eQTL associations for a gene and/or a variant (precomputed). Provide gencode_id and/or variant_id; tissue_site_detail_id optionally narrows. Paged and count-verified.

FieldTypeRequirement and constraints
gencode_idstringoptional
variant_idstringoptional
tissue_site_detail_idstringoptional
max_resultsintegeroptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_single_tissue_eqtls", {"gencode_id": "ENSG00000111640.14"})

gtex_multi_tissue_eqtls

Multi-tissue cis-eQTL meta-analysis (METASOFT) for a VERSIONED GENCODE id. variant_id optionally narrows to one variant. Returns per-variant rows with per-tissue m-values, NES, p-values, and SEs.

FieldTypeRequirement and constraints
gencode_idstringrequired
variant_idstringoptional
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_multi_tissue_eqtls", {"gencode_id": "ENSG00000111640.14"})

gtex_calculate_eqtl

Calculate an eQTL on the fly for any gene-variant pair in one tissue, including non-significant pairs. Returns p-value, NES, t-statistic, MAF, and the per-sample genotype/expression arrays.

FieldTypeRequirement and constraints
gencode_idstringrequired
variant_idstringrequired
tissue_site_detail_idstringrequired
dataset_idstringoptional; default: "gtex_v8"
const result = await host.mcp("expression", "gtex_calculate_eqtl", {"gencode_id": "ENSG00000111640.14", "variant_id": "chr12_6452899_G_A_b38", "tissue_site_detail_id": "Whole_Blood"})

Protein Annotation

Show operations and parameters

get_domain_architecture

Complete InterPro domain architecture for one or more UniProt proteins (all matching entries, member-DB signatures, fragment coordinates), with pagination verified against the API count.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
const result = await host.mcp("protein-annotation", "get_domain_architecture", {"accessions": ["P04637"]})

search_interpro_entries

Keyword search over InterPro or member-database entries (Pfam, SMART, PROSITE, PANTHER, CDD), complete cursor walk verified against the API count.

FieldTypeRequirement and constraints
querystringoptional
entry_typestringoptional
source_dbstringoptional; default: "interpro"
go_termstringoptional
const result = await host.mcp("protein-annotation", "search_interpro_entries", {"query": "kinase", "source_db": "pfam"})

get_interpro_entry

Detail record for an InterPro entry (IPRxxxxxx) or Pfam family (PFxxxxx) — route chosen by accession prefix.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("protein-annotation", "get_interpro_entry", {"accession": "IPR000719"})

search_pfam_clans

Keyword search over Pfam clans (InterPro sets, accessions CLxxxx).

FieldTypeRequirement and constraints
querystringoptional
const result = await host.mcp("protein-annotation", "search_pfam_clans", {"query": "kinase"})

get_pfam_clan

Pfam clan detail including the complete sorted member-family list.

FieldTypeRequirement and constraints
clan_accessionstringrequired
const result = await host.mcp("protein-annotation", "get_pfam_clan", {"clan_accession": "CL0016"})

get_pfam_family_proteins

Member proteins of a Pfam family (complete count-verified walk or count only). Use count_only for very large families.

FieldTypeRequirement and constraints
pfam_accessionstringrequired
reviewed_onlybooleanoptional; default: false
tax_idintegeroptional
count_onlybooleanoptional; default: false
const result = await host.mcp("protein-annotation", "get_pfam_family_proteins", {"pfam_accession": "PF00069", "count_only": true})

get_pfam_family_proteomes

Proteomes containing members of a Pfam family. count_only defaults true — the upstream proteome cursor pagination is defective for deep walks.

FieldTypeRequirement and constraints
pfam_accessionstringrequired
count_onlybooleanoptional; default: true
const result = await host.mcp("protein-annotation", "get_pfam_family_proteomes", {"pfam_accession": "PF00069"})

get_protein_atlas_gene

Human Protein Atlas per-gene record (release 25.x): tissue/subcellular/pathology/blood/brain expression and antibody info. Accepts an Ensembl gene ID or a gene symbol.

FieldTypeRequirement and constraints
genestringrequired
fullbooleanoptional; default: false
const result = await host.mcp("protein-annotation", "get_protein_atlas_gene", {"gene": "TP53"})

search_protein_atlas

Column-selected bulk search over the Human Protein Atlas (search_download).

FieldTypeRequirement and constraints
querystringrequired
columnsstringoptional; default: "g,gs,eg,gd,up,chr,chrp,scl"
const result = await host.mcp("protein-annotation", "search_protein_atlas", {"query": "kinase"})

map_string_ids

Map gene symbols/aliases to STRING protein identifiers (v12.0). Every input symbol is either mapped or listed in unmapped — the two partition the input.

FieldTypeRequirement and constraints
symbolsarray of stringrequired
speciesintegeroptional; default: 9606
const result = await host.mcp("protein-annotation", "map_string_ids", {"symbols": ["TP53", "BRCA1", "EGFR"]})

get_string_network

STRING protein-protein interaction network for a gene list (v12.0) at a confidence threshold. Maps symbols first (unmapped reported), then retrieves nodes, edges, summary and provenance.

FieldTypeRequirement and constraints
symbolsarray of stringrequired
speciesintegeroptional; default: 9606
required_scoreintegeroptional; default: 700
const result = await host.mcp("protein-annotation", "get_string_network", {"symbols": ["TP53", "BRCA1", "EGFR"], "required_score": 700})

get_string_similarity_scores

Smith-Waterman protein similarity bitscores among a gene set (STRING /homology). Sparse: pairs absent from STRING's data are not listed (absence means no recorded similarity, not zero).

FieldTypeRequirement and constraints
symbolsarray of stringrequired
speciesintegeroptional; default: 9606
const result = await host.mcp("protein-annotation", "get_string_similarity_scores", {"symbols": ["TP53", "MDM2", "MDM4"]})

get_string_best_similarity_hits

Best homology hit per input protein in a target species (STRING /homology_best). target_species=null asks for the best hit across all species.

FieldTypeRequirement and constraints
symbolsarray of stringrequired
speciesintegeroptional; default: 9606
target_speciesintegeroptional
const result = await host.mcp("protein-annotation", "get_string_best_similarity_hits", {"symbols": ["TP53"], "target_species": 10090})

Cancer Models

Show operations and parameters

cbioportal_list_studies

List cBioPortal cancer studies, optionally filtered by a free-text keyword (name/description/cancer type) and/or an exact cancer-type id; returns study id, name, cancer type, reference genome, citation, and per-data-type sample counts.

FieldTypeRequirement and constraints
keywordstringoptional
cancer_type_idstringoptional
max_recordsintegeroptional; default: 500
const result = await host.mcp("cancer-models", "cbioportal_list_studies", {"keyword": "glioma"})

cbioportal_get_study

Get a cBioPortal cancer study by id: metadata, per-data-type sample counts, true sample/patient counts (from the study collections, not the display field), and its molecular profiles.

FieldTypeRequirement and constraints
study_idstringrequired
const result = await host.mcp("cancer-models", "cbioportal_get_study", {"study_id": "msk_impact_2017"})

cbioportal_mutations_in_gene

All mutations of one gene (HUGO symbol) in a cBioPortal study, with recurrence aggregates: total mutations, mutated-sample count, mutation-type and protein-change distributions, and the most recurrent protein changes.

FieldTypeRequirement and constraints
gene_symbolstringrequired
study_idstringrequired
max_recordsintegeroptional; default: 100
const result = await host.mcp("cancer-models", "cbioportal_mutations_in_gene", {"gene_symbol": "IDH1", "study_id": "difg_msk_2023"})

cbioportal_mutation_frequency

Mutation frequency of one gene across several cBioPortal studies (1–12): mutated-sample fraction of the sequenced cohort per study, ranked most-frequent first.

FieldTypeRequirement and constraints
gene_symbolstringrequired
study_idsarray of stringrequired; minItems: 1; maxItems: 12
const result = await host.mcp("cancer-models", "cbioportal_mutation_frequency", {"gene_symbol": "KRAS", "study_ids": ["msk_impact_2017", "difg_msk_2023"]})

cbioportal_cna_in_gene

Discrete copy-number alterations of one gene in a cBioPortal study, filtered by event type (deep deletion / amplification by default), with the full per-sample alteration distribution.

FieldTypeRequirement and constraints
gene_symbolstringrequired
study_idstringrequired
event_typestringoptional; default: "HOMDEL_AND_AMP"; enum: ["HOMDEL_AND_AMP", "HOMDEL", "AMP", "GAIN", "HETLOSS", "DIPLOID", "ALL"]
max_recordsintegeroptional; default: 100
const result = await host.mcp("cancer-models", "cbioportal_cna_in_gene", {"gene_symbol": "CDKN2A", "study_id": "msk_impact_2017"})

cbioportal_clinical_attributes

Clinical attributes defined in a cBioPortal study (patient- and sample-level fields), highlighting survival endpoints and whether overall-survival data is present.

FieldTypeRequirement and constraints
study_idstringrequired
max_recordsintegeroptional; default: 200
const result = await host.mcp("cancer-models", "cbioportal_clinical_attributes", {"study_id": "brca_tcga_pan_can_atlas_2018"})

RNA

Show operations and parameters

get_family

Rfam family metadata for an accession (RF00005) or family id (tRNA) — both resolve. Flattened record plus the full upstream JSON in "raw".

FieldTypeRequirement and constraints
familystringrequired
const result = await host.mcp("rna", "get_family", {"family": "RF00005"})

get_seed_alignment

Seed alignment of an Rfam family in Stockholm (default, with consensus secondary-structure line) or aligned gapped FASTA.

FieldTypeRequirement and constraints
familystringrequired
fmtstringoptional; default: "stockholm"; enum: ["stockholm", "fasta"]
max_bytesintegeroptional; default: 400000
const result = await host.mcp("rna", "get_seed_alignment", {"family": "RF00162", "fmt": "stockholm"})

get_covariance_model

Infernal covariance model (CM file) of an Rfam family, usable directly with cmsearch/cmscan, plus parsed header fields.

FieldTypeRequirement and constraints
familystringrequired
max_bytesintegeroptional; default: 400000
const result = await host.mcp("rna", "get_covariance_model", {"family": "RF00162"})

get_tree

Seed phylogenetic tree of an Rfam family (NHX/Newick text).

FieldTypeRequirement and constraints
familystringrequired
const result = await host.mcp("rna", "get_tree", {"family": "RF00162"})

get_sequence_regions

All full-region hits of an Rfam family across sequence databases (parsed TSV). Check num_full via get_family first — rfam.org 403s this route for very large families (e.g. RF00005).

FieldTypeRequirement and constraints
familystringrequired
const result = await host.mcp("rna", "get_sequence_regions", {"family": "RF00162"})

get_structure_mapping

PDB residue-level structure mappings of an Rfam family, deterministically sorted.

FieldTypeRequirement and constraints
familystringrequired
const result = await host.mcp("rna", "get_structure_mapping", {"family": "RF00162"})

accession_to_id

Convert an Rfam accession to its family id (e.g. RF00005 -> "tRNA").

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("rna", "accession_to_id", {"accession": "RF00005"})

id_to_accession

Convert an Rfam family id to its accession (e.g. "tRNA" -> RF00005).

FieldTypeRequirement and constraints
family_idstringrequired
const result = await host.mcp("rna", "id_to_accession", {"family_id": "tRNA"})

search_sequence

Search an RNA sequence through the official Rfam batch endpoint. Keep the returned job identity while waiting; an unfinished response is not a zero-hit result. Inspect the completed matches and source information. After a failed response, diagnose or resume the existing job rather than submitting it repeatedly.

FieldTypeRequirement and constraints
sequencestringrequired
max_wait_snumberoptional; default: 300
poll_interval_snumberoptional; default: 5
const result = await host.mcp("rna", "search_sequence", {"sequence": "GGUUCCGGGAAGGCAGCAGGUGGAAACCUGCCA"})

Omics Archives

Show operations and parameters

arrayexpress_search_experiments

Search ArrayExpress functional-genomics experiments (BioStudies) with complete, totalHits-verified retrieval; filters (query, organism, study_type, technology, release-date range, extra facets) combine with AND.

FieldTypeRequirement and constraints
querystringoptional
organismstringoptional
study_typestringoptional
technologystringoptional
released_afterstringoptional
released_beforestringoptional
extra_facetsobjectoptional
max_recordsintegeroptional; default: 50
const result = await host.mcp("omics-archives", "arrayexpress_search_experiments", {"organism": "Homo sapiens", "study_type": "ChIP-seq", "max_records": 50})

arrayexpress_get_experiment

Fetch one ArrayExpress experiment (BioStudies) as a flattened analyst record — study type, organisms, assay/sample counts, designs/factors, authors, publications, protocols, array designs, and file summary.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("omics-archives", "arrayexpress_get_experiment", {"accession": "E-MTAB-5061"})

arrayexpress_get_experiment_files

List every file of an ArrayExpress experiment (name, size, type, format, description) with download URLs, plus the /info endpoint file count carried alongside for comparison.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("omics-archives", "arrayexpress_get_experiment_files", {"accession": "E-MTAB-5061"})

arrayexpress_get_experiment_samples

Fetch per-sample SDRF annotation rows for an ArrayExpress experiment (MAGE-TAB headers verbatim, repeats suffixed #2/#3). Experiments with no SDRF return {"error":"no_sdrf"}.

FieldTypeRequirement and constraints
accessionstringrequired
max_rows_returnedintegeroptional; default: 200
const result = await host.mcp("omics-archives", "arrayexpress_get_experiment_samples", {"accession": "E-MTAB-5061", "max_rows_returned": 200})

geo_search_series

Search NCBI GEO DataSets (db=gds) and return series-level records (trimmed esummary docs). term is full E-utilities syntax; add gse[ETYP] to restrict to series.

FieldTypeRequirement and constraints
termstringrequired
retmaxintegeroptional; default: 500
const result = await host.mcp("omics-archives", "geo_search_series", {"term": "asthma AND gse[ETYP]", "retmax": 20})

geo_get_series

Fetch structured metadata for GEO series (GSE accessions) with samples included — series title/summary/design, platforms, samples with characteristics and library info, and supplementary-file URLs. Data tables are never downloaded.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
const result = await host.mcp("omics-archives", "geo_get_series", {"accessions": ["GSE131907"]})

metabolights_list_studies

List every public MetaboLights study accession (numerically sorted) with the API's own reported count. There is no server-side study search — filter fetched candidates by title/descriptor instead.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("omics-archives", "metabolights_list_studies", {})

metabolights_get_studies

Fetch structured metadata for MetaboLights studies (MTBLSxxx) from the parsed ISA payload — title, status, years, organisms, assays, factors, descriptors, sample count, protocols; optional per-sample table. Unknown/private accessions go in not_found.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
include_samplesbooleanoptional; default: false
max_sample_rows_returnedintegeroptional; default: 200
const result = await host.mcp("omics-archives", "metabolights_get_studies", {"accessions": ["MTBLS1"], "include_samples": false})

metabolights_get_study_files

Complete file inventory for a public MetaboLights study — the top-level study folder (ISA-Tab, MAF, folder entries) and, by default, the recursive FILES data folder.

FieldTypeRequirement and constraints
accessionstringrequired
include_data_filesbooleanoptional; default: true
const result = await host.mcp("omics-archives", "metabolights_get_study_files", {"accession": "MTBLS1"})

metabolights_search_data_files

Glob search over a MetaboLights study's raw-data folder (FILES tree). pattern is a filename glob (e.g. '.mzML', '.raw'); omit it to list every data file.

FieldTypeRequirement and constraints
accessionstringrequired
patternstringoptional
const result = await host.mcp("omics-archives", "metabolights_search_data_files", {"accession": "MTBLS1", "pattern": "*.zip"})

mgnify_search_studies

Find MGnify metagenomics studies by free text OR biome lineage (provide exactly one). Full listing is paginated to completion and count-verified against the API.

FieldTypeRequirement and constraints
querystringoptional
biome_lineagestringoptional
const result = await host.mcp("omics-archives", "mgnify_search_studies", {"query": "coral"})

mgnify_get_studies

Fetch structured records for MGnify studies (MGYS accessions). With include_analyses, each study also carries its complete analyses listing plus by-pipeline/by-experiment breakdowns. Unknown accessions go in missing.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
include_analysesbooleanoptional; default: false
const result = await host.mcp("omics-archives", "mgnify_get_studies", {"accessions": ["MGYS00000410"], "include_analyses": false})

mgnify_get_study_analyses

List ALL analyses of one MGnify study (complete, count-verified pagination) — one record per MGYA analysis with pipeline version, experiment type, status, and run/assembly/sample accessions.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("omics-archives", "mgnify_get_study_analyses", {"accession": "MGYS00000410"})

pride_search_projects

Search PRIDE Archive proteomics projects (complete, api_total-verified retrieval); filters (keyword, organism, instrument, disease, extra_filters) combine with AND. Sorted by accession ASC — a bounded walk is a stable prefix.

FieldTypeRequirement and constraints
keywordstringoptional
organismstringoptional
instrumentstringoptional
diseasestringoptional
extra_filtersobjectoptional
max_records_returnedintegeroptional; default: 50
const result = await host.mcp("omics-archives", "pride_search_projects", {"keyword": "phosphoproteome", "organism": "Homo sapiens (human)", "max_records_returned": 50})

pride_get_projects

Fetch full metadata for PRIDE projects by accession (e.g. PXD010154) — the same normalized record shape as pride_search_projects, so the two are directly comparable. Unknown accessions go in not_found.

FieldTypeRequirement and constraints
accessionsarray of stringrequired
const result = await host.mcp("omics-archives", "pride_get_projects", {"accessions": ["PXD010154"]})

pride_search_project_proteins

List protein evidence rows for one PRIDE affinity-proteomics project (paged to exhaustion). NOTE: only affinity-proteomics projects are served here; for classic MS (PXD) projects use pride_find_projects_for_protein instead.

FieldTypeRequirement and constraints
project_accessionstringrequired
keywordstringoptional
const result = await host.mcp("omics-archives", "pride_search_project_proteins", {"project_accession": "PXD010154"})

pride_find_projects_for_protein

Find PRIDE projects containing a protein (MS-archive direction). protein_accession is a UniProt accession (e.g. P04637). Feed the returned project accessions to pride_get_projects for full metadata.

FieldTypeRequirement and constraints
protein_accessionstringrequired
const result = await host.mcp("omics-archives", "pride_find_projects_for_protein", {"protein_accession": "P04637"})

CellGuide

Show operations and parameters

get_cell_type_info

CellGuide (CELLxGENE) cell-type info by Cell Ontology id or name: name, synonyms, ontology description, and curated/GPT description.

FieldTypeRequirement and constraints
cell_typestringrequired
const result = await host.mcp("cellguide", "get_cell_type_info", {"cell_type": "acinar cell"})

search_cell_types

Search CellGuide cell types by free text over name and synonyms (the CDN has no search endpoint, so celltype_metadata.json is filtered client-side).

FieldTypeRequirement and constraints
querystringrequired
limitintegeroptional; default: 25
const result = await host.mcp("cellguide", "search_cell_types", {"query": "T cell", "limit": 25})

get_marker_genes

CellGuide marker genes for a cell type (id or name): computational (data-derived, scored) or canonical (literature-curated).

FieldTypeRequirement and constraints
cell_typestringrequired
marker_typestringoptional; default: "computational"; enum: ["computational", "canonical"]
limitintegeroptional; default: 25
const result = await host.mcp("cellguide", "get_marker_genes", {"cell_type": "CL:0000084", "marker_type": "computational", "limit": 25})

get_source_data

CellGuide source datasets and publications contributing to a cell type (id or name): collection name/url, publication, and the tissues/diseases/organisms each covers.

FieldTypeRequirement and constraints
cell_typestringrequired
const result = await host.mcp("cellguide", "get_source_data", {"cell_type": "CL:0000622"})

get_cell_tissues

Anatomical tissues where a cell type (id or name) is observed, aggregated (deduplicated) across CellGuide source collections.

FieldTypeRequirement and constraints
cell_typestringrequired
const result = await host.mcp("cellguide", "get_cell_tissues", {"cell_type": "T cell"})

Regulation

Show operations and parameters

encode_search_experiments

Search ENCODE functional-genomics experiments (ChIP-seq, ATAC-seq, ...). Filters: assay_title (e.g. "TF ChIP-seq"), target (protein label, e.g. "CTCF"), organism (scientific name), status (default "released"), date_released_before (ISO date — a closed window), plus arbitrary portal field filters via extra_filters. The full result set is paged and count-verified; accessions lists every match, at most max_rows row summaries are returned.

FieldTypeRequirement and constraints
assay_titlestringoptional
targetstringoptional
organismstringoptional
statusstringoptional; default: "released"
date_released_beforestringoptional
extra_filtersobjectoptional
max_rowsintegeroptional; default: 100
const result = await host.mcp("regulation", "encode_search_experiments", {"target": "CTCF", "assay_title": "TF ChIP-seq", "max_rows": 50})

encode_search_biosamples

Search ENCODE biosamples (cell lines, tissues, primary cells). Filters: term_name (ontology term, e.g. "K562"), classification ("cell line", "tissue", ...), organism (scientific name), status (default "released"), date_created_before (ISO date), plus arbitrary portal field filters via extra_filters. Complete, count-verified: accessions is the full match list, at most max_rows row summaries are returned.

FieldTypeRequirement and constraints
term_namestringoptional
classificationstringoptional
organismstringoptional
statusstringoptional; default: "released"
date_created_beforestringoptional
extra_filtersobjectoptional
max_rowsintegeroptional; default: 100
const result = await host.mcp("regulation", "encode_search_biosamples", {"term_name": "K562", "classification": "cell line", "max_rows": 25})

encode_list_files

List ENCODE data files by format / assay / biosample. Filters: file_format ("fastq", "bam", "bigWig", "bed", ...), assay_term_name (the ontology term e.g. "ChIP-seq" — NOT the display assay_title like "TF ChIP-seq", which matches nothing; pass titles via extra_filters={"assay_title": ...}), biosample_term_name (e.g. "K562"), status (default "released"), date_created_before, plus arbitrary portal field filters via extra_filters. File queries match millions of rows unfiltered — always combine several filters. Complete + count-verified; at most max_rows row summaries returned.

FieldTypeRequirement and constraints
file_formatstringoptional
assay_term_namestringoptional
biosample_term_namestringoptional
statusstringoptional; default: "released"
date_created_beforestringoptional
extra_filtersobjectoptional
max_rowsintegeroptional; default: 100
const result = await host.mcp("regulation", "encode_list_files", {"file_format": "bed", "assay_term_name": "ChIP-seq", "biosample_term_name": "K562", "extra_filters": {"output_type": "peaks", "assembly": "GRCh38"}, "max_rows": 50})

encode_get_experiment

Get one ENCODE experiment by accession (e.g. "ENCSR000AKP"). Returns a stable-field record: assay, target, biosample ontology + summary, description, lab, award project, release/submission dates, assemblies, replicate counts, replication type, dbxrefs, DOI and uuid. Volatile portal fields (audits, analyses, internal status) are excluded.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("regulation", "encode_get_experiment", {"accession": "ENCSR000AKP"})

encode_get_file

Get one ENCODE file by accession (e.g. "ENCFF002JUR"). Returns a stable-field record: format, output type/category, assay, assembly, parent dataset, biological replicates, file size, md5sums, run type, read length, lab, creation date, download href and uuid.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("regulation", "encode_get_file", {"accession": "ENCFF002JUR"})

encode_get_biosample

Get one ENCODE biosample by accession (e.g. "ENCBS013JZP"). Returns a stable-field record: ontology term + classification, organism, summary/description, source, donor, treatments, genetic modifications, life stage, age, sex, lab, creation date, status and uuid.

FieldTypeRequirement and constraints
accessionstringrequired
const result = await host.mcp("regulation", "encode_get_biosample", {"accession": "ENCBS013JZP"})

jaspar_get_matrix

Get one JASPAR TF binding profile by VERSIONED matrix id (e.g. "MA0002.2"). Returns the full record: position frequency matrix (pfm), TF name/class/family, species, data type, literature references (pubmed/medline), sequence logo URL. Requires a versioned id ("MA0002.2", not "MA0002") — use jaspar_matrix_versions to enumerate versions. Versioned matrices are immutable, so results are reproducible.

FieldTypeRequirement and constraints
matrix_idstringrequired
const result = await host.mcp("regulation", "jaspar_get_matrix", {"matrix_id": "MA0002.2"})

jaspar_matrix_versions

List all versions of a JASPAR base matrix id (e.g. "MA0002"). Returns every released version with its matrix_id, name, collection and URL — count-verified. Use to pin an exact version before jaspar_get_matrix, or to track how a profile changed across releases. A versioned id ("MA0002.2") is accepted and reduced to its base.

FieldTypeRequirement and constraints
base_idstringrequired
const result = await host.mcp("regulation", "jaspar_matrix_versions", {"base_id": "MA0002"})

jaspar_list_matrices

Search/list JASPAR TF binding profiles (the full profile catalog). Filters (all optional): collection ("CORE", "UNVALIDATED"), tax_group ("vertebrates", "plants", ...), tax_id (NCBI taxonomy id, e.g. 9606 for human — this is how you filter by species; enumerate ids with jaspar_list_species), name (exact TF name, e.g. "FOXA1"), search (free text), version="latest" (restrict to latest versions only). The full filtered catalog is paginated and count-verified; at most max_rows summary rows are returned.

FieldTypeRequirement and constraints
collectionstringoptional
tax_groupstringoptional
tax_idintegeroptional
namestringoptional
searchstringoptional
versionstringoptional
max_rowsintegeroptional; default: 1000
const result = await host.mcp("regulation", "jaspar_list_matrices", {"tax_id": 9606, "collection": "CORE", "version": "latest", "max_rows": 200})

jaspar_list_species

List all species with JASPAR profiles (NCBI tax_id + name); count-verified full listing. Use the tax_id values to filter jaspar_list_matrices (e.g. 9606 = Homo sapiens, 10090 = Mus musculus).

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("regulation", "jaspar_list_species", {})

jaspar_list_taxa

List all JASPAR taxonomic groups (vertebrates, plants, fungi, insects, ...); count-verified full listing. Use the group names as the tax_group filter of jaspar_list_matrices.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("regulation", "jaspar_list_taxa", {})

jaspar_list_collections

List all JASPAR collections (CORE, UNVALIDATED, ...); count-verified full listing. Use the collection names as the collection filter of jaspar_list_matrices (CORE = curated, non-redundant profiles).

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("regulation", "jaspar_list_collections", {})

jaspar_list_releases

List all JASPAR database releases (year, release number, active flag); count-verified full listing. Record the active release when selecting motifs for reproducibility, or check release history before comparing results across JASPAR versions.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("regulation", "jaspar_list_releases", {})

unibind_search_tfbs

Search UniBind ChIP-seq datasets with high-confidence TFBS predictions (unibind.uio.no, 2021 release; direct TF-DNA interactions from ~10k datasets across 9 species). Each dataset is one (experiment, cell type, TF) triple. Filters (all optional, AND-combined, exact-match unless noted): tf_name (gene symbol, e.g. "CTCF"), cell_line (verbose UniBind title — prefer search for fuzzy matching), species (scientific name), collection ("Robust" = best-model / high confidence, or "Permissive"), jaspar_id (versioned, e.g. "MA0139.1"), search (free text). total is the API's exact count; at most max_rows rows are returned (a stable prefix).

FieldTypeRequirement and constraints
tf_namestringoptional
cell_linestringoptional
speciesstringoptional
collectionstringoptional; enum: ["Robust", "Permissive"]
jaspar_idstringoptional
searchstringoptional
max_rowsintegeroptional; default: 200
const result = await host.mcp("regulation", "unibind_search_tfbs", {"tf_name": "CTCF", "collection": "Robust", "max_rows": 50})

unibind_get_dataset

Get one UniBind dataset's detail: per-model TFBS counts + file URLs. tf_id is the dataset key "<identifier>.<cell_line>.<TF>" as returned by unibind_search_tfbs (e.g. "ENCSR000AUE.A549_lung_carcinoma.CTCF"). Returns the TF name, source identifiers (ENCODE/GEO/GTRD), cell lines, biological conditions, JASPAR matrix ids, ChIP-seq peak count, and one row per TFBS prediction model (DAMO/PWM/...) with total_tfbs, score/distance thresholds, adjusted CentriMo p-value, and direct BED/FASTA download URLs — use those URLs (not an MCP call) to retrieve the complete site list.

FieldTypeRequirement and constraints
tf_idstringrequired
const result = await host.mcp("regulation", "unibind_get_dataset", {"tf_id": "ENCSR000AUE.A549_lung_carcinoma.CTCF"})

unibind_tfbs_in_region

TF binding sites overlapping a genomic region (UniBind 2021 maps), served via the UCSC hubApi against UniBind's registered public track hubs (UniBind's own REST API has no region endpoint). Coordinates are 0-based half-open. genome: UCSC assembly — Robust hub: hg38, mm10, ce11, dm6, danRer11, sacCer3, rn6, araTha1; Permissive adds spo2 (no hg19 — lift first). chrom: with "chr" prefix. start/end: interval, end-start <= 1,000,000 bp. HONEST-CAP: at most 20,000 items are scanned per call; region_scan_complete=false means the region has more sites than were scanned (narrow the window) and, with tf_name set, matches may be missing. n_matching counts scanned sites passing the filter; returned/truncated describe the max_sites cap.

FieldTypeRequirement and constraints
genomestringrequired
chromstringrequired
startintegerrequired
endintegerrequired
tf_namestringoptional
collectionstringoptional; default: "Robust"; enum: ["Robust", "Permissive"]
max_sitesintegeroptional; default: 2000
const result = await host.mcp("regulation", "unibind_tfbs_in_region", {"genome": "hg38", "chrom": "chr1", "start": 1000000, "end": 1010000, "collection": "Robust"})

Research Resources

Show operations and parameters

search_grants

Search Grants.gov funding opportunities via the search2 API (complete, count-verified retrieval). At least one criterion is required (keyword, opportunity_number, aln/CFDA, agencies, eligibilities, funding_categories, or funding_instruments). opportunity_statuses defaults to ["forecasted","posted"] (current opportunities); add "closed"/"archived" for historical ones. agencies takes codes like ["HHS-NIH11"] (NIH), ["HHS-FDA"], ["NSF"]. Set count_only for just the hit count + facets; max_records caps returned records (the walk still retrieves the complete set and flags truncated).

FieldTypeRequirement and constraints
keywordstringoptional
opportunity_numberstringoptional
alnstringoptional
agenciesarray of stringoptional
opportunity_statusesarray of stringoptional
eligibilitiesarray of stringoptional
funding_categoriesarray of stringoptional
funding_instrumentsarray of stringoptional
count_onlybooleanoptional; default: false
max_recordsintegeroptional; default: 100
include_facetsbooleanoptional; default: true
const result = await host.mcp("research-resources", "search_grants", {"keyword": "cancer", "agencies": ["HHS-NIH11"], "max_records": 25})

search_antibodies

Full-text search the Antibody Registry (antibodyregistry.org, ~3.2M records). Token-based matching against antibody name/target/catalog text ("TP53" and "p53" are different queries). With page omitted, all pages are walked up to max_records or the anonymous depth cap (rows beyond offset 500 need authentication upstream, flagged as anonymous_limit_hit — never silently dropped). Pass a 1-based page for single-page retrieval (page*page_size must stay <= 500).

FieldTypeRequirement and constraints
querystringrequired
pageintegeroptional
page_sizeintegeroptional; default: 100
max_recordsintegeroptional; default: 500
const result = await host.mcp("research-resources", "search_antibodies", {"query": "CD4", "max_records": 100})

get_antibody

Fetch Antibody Registry detail record(s) for one antibody accession / RRID. Accepts a plain number ("3643095"), "AB_3643095", or "RRID:AB_3643095". The upstream route is list-valued (an accession can map to several curated records, e.g. multi-vendor duplicates). A nonexistent id yields record_count 0, not an error.

FieldTypeRequirement and constraints
antibody_idstringrequired
const result = await host.mcp("research-resources", "get_antibody", {"antibody_id": "RRID:AB_3643095"})

find_antibodies_by_catalog

Find antibodies by vendor catalog number (exact, case-insensitive). Implemented as a full-text search plus client-side exact matching on the catalog number (or its listed alternatives), because the upstream column-filter route returns HTTP 500 for every key. Pass an optional vendor name (exact, case-insensitive) to further narrow the matches.

FieldTypeRequirement and constraints
catalog_numberstringrequired
vendorstringoptional
page_sizeintegeroptional; default: 100
const result = await host.mcp("research-resources", "find_antibodies_by_catalog", {"catalog_number": "ab32572"})

get_antibody_registry_stats

Antibody Registry statistics: total antibody count and last-update date. Returns the upstream /api/datainfo payload.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("research-resources", "get_antibody_registry_stats", {})

BioMart

Show operations and parameters

list_marts

List available Ensembl BioMart marts (databases). BioMart organizes data as MART -> DATASET -> ATTRIBUTES/FILTERS; a mart name feeds list_datasets.

FieldTypeRequirement and constraints
objectNo fields; pass an empty object.
const result = await host.mcp("biomart", "list_marts", {})

list_datasets

List the datasets available in a given mart (e.g. hsapiens_gene_ensembl for human genes). A dataset name feeds the attribute/filter/query tools.

FieldTypeRequirement and constraints
martstringrequired
const result = await host.mcp("biomart", "list_datasets", {"mart": "ENSEMBL_MART_ENSEMBL"})

list_common_attributes

List the commonly used attributes for a dataset (a curated high-signal subset). Use this before list_all_attributes to pick attributes for get_data. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
const result = await host.mcp("biomart", "list_common_attributes", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl"})

list_all_attributes

List all attributes available for a dataset, minus homologs and microarray probes (which are bulky and rarely needed). Can be large; prefer list_common_attributes first. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
const result = await host.mcp("biomart", "list_all_attributes", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl"})

list_filters

List the filters available for a dataset. Filters narrow a get_data query (e.g. chromosome_name, biotype) and are passed to get_data as a filters dict. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
const result = await host.mcp("biomart", "list_filters", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl"})

get_data

Run a BioMart query: retrieve the requested attributes for a dataset, optionally narrowed by filters. This is the main data-retrieval tool. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
attributesarray of stringrequired
filtersobjectoptional
const result = await host.mcp("biomart", "get_data", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl", "attributes": ["ensembl_gene_id", "external_gene_name", "chromosome_name"], "filters": {"chromosome_name": "Y", "biotype": "protein_coding"}})

get_translation

Translate a single identifier from one attribute type to another (e.g. an HGNC symbol to an Ensembl gene ID) within a dataset. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
from_attrstringrequired
to_attrstringrequired
targetstringrequired
const result = await host.mcp("biomart", "get_translation", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl", "from_attr": "hgnc_symbol", "to_attr": "ensembl_gene_id", "target": "TP53"})

batch_translate

Translate many identifiers from one attribute type to another in a single query — more efficient than repeated get_translation calls. mart is accepted for signature parity but ignored; the query keys off dataset.

FieldTypeRequirement and constraints
martstringrequired
datasetstringrequired
from_attrstringrequired
to_attrstringrequired
targetsarray of stringrequired
const result = await host.mcp("biomart", "batch_translate", {"mart": "ENSEMBL_MART_ENSEMBL", "dataset": "hsapiens_gene_ensembl", "from_attr": "hgnc_symbol", "to_attr": "ensembl_gene_id", "targets": ["TP53", "BRCA1", "BRCA2"]})

ZINC

Show operations and parameters

zinc_search_by_id

Look up purchasable compounds in ZINC22/ZINC20 by ZINC identifier — answers "what is this compound and who sells it". Batched: pass up to 100 ids in one call rather than many single-id calls. Async upstream (submit + poll); can take up to timeout_s seconds.

FieldTypeRequirement and constraints
zinc_ids['string', 'array']required
max_resultsintegeroptional; default: 50
timeout_snumberoptional; default: 25
const result = await host.mcp("zinc", "zinc_search_by_id", {"zinc_ids": ["ZINC000000000012"]})

zinc_search_by_smiles

Search ZINC22's purchasable chemical space by structure — answers "what purchasable compounds look like this SMILES". This is BOTH the exact-match and the analog-discovery (similarity) tool: CartBlanche22 exposes one structure-search endpoint whose dist parameter spans exact through diverse, so there is deliberately no separate similarity-search tool. The slowest ZINC query — raise dist gradually rather than starting loose.

FieldTypeRequirement and constraints
smilesstringrequired
distintegeroptional; default: 0
adistintegeroptional
max_resultsintegeroptional; default: 50
timeout_snumberoptional; default: 25
const result = await host.mcp("zinc", "zinc_search_by_smiles", {"smiles": "CC(=O)Oc1ccccc1C(=O)O", "dist": 2})

zinc_search_by_supplier

Resolve vendor catalog numbers to ZINC compounds — answers "which ZINC substance is this supplier code, and what's its structure". Batched: up to 100 supplier codes per call. Async upstream (submit + poll).

FieldTypeRequirement and constraints
supplier_codes['string', 'array']required
max_resultsintegeroptional; default: 50
timeout_snumberoptional; default: 25
const result = await host.mcp("zinc", "zinc_search_by_supplier", {"supplier_codes": ["MCULE-2311834287"]})

zinc_random_sample

Draw a random sample of purchasable compounds from ZINC22 — for building screening decks, property baselines, or decoy sets. count doubles as this tool's max_results; re-calling draws a fresh sample. Async upstream (submit + poll).

FieldTypeRequirement and constraints
countintegeroptional; default: 50
subsetstringoptional
timeout_snumberoptional; default: 25
const result = await host.mcp("zinc", "zinc_random_sample", {"count": 25, "subset": "lead-like"})

zinc_get_3d

Locate docking-ready 3D structures for ZINC compounds. ZINC22 ships pre-generated 3D conformers (DOCK .db2.gz, .mol2.gz, .sdf.gz) in its file repository, organized by tranche — this tool resolves each id to its tranche and returns the repository locations to download from for docking prep (DOCK6, AutoDock Vina, etc.). Max 50 ids per call (3D retrieval is per-compound work). Async upstream (submit + poll).

FieldTypeRequirement and constraints
zinc_ids['string', 'array']required
timeout_snumberoptional; default: 25
const result = await host.mcp("zinc", "zinc_get_3d", {"zinc_ids": ["ZINC000000000012"]})

Example response records

The example response records include exact inputs, capped response excerpts and per-operation outcomes. Distinguish a returned record, an empty match and a failed request. Results may be metadata, schemas or identifiers; check the source fields and completeness flags before using them in your research.