# Documentation > Complete documentation for Large Language Models --- ## Document: SeqHub API URL: /introduction # SeqHub API The SeqHub API lets you search 130,000+ microbial genomes by any combination of protein sequences, Pfam domains, Rfam RNA families and gLM2-derived SAE features, and annotate proteins. ## Quickstart Every request needs a personal access token (see [Authentication](#authentication)). **Search with a protein and features.** Find neighborhoods with a protein similar to E. coli's vitamin B12-binding protein BtuF, an ABC transporter domain (`PF00005`) and a cobalamin riboswitch (`RF00174`): ```bash curl https://api.seqhub.org/api/v1/protein-contexts/search/multi-query \ -H "Authorization: Bearer $SEQHUB_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "sequences": ["MAKSLFRALVALSFLAPLWLNAAPRVITLSPANTELAFAAGITPVGVSSYSDYPPQAQKIEQVSTWQGMNLERIVALKPDLVIAWRGGNAERQVDQLASLGIKVMWVDATSIEQIANALRQLAPWSPQPDKAEQAAQSLLDQYAQLKAQYADKPKKRVFLQFGINPPFTSGKESIQNQVLEVCGGENIFKDSRVPWPQVSREQVLARSPQAIVITGGPDQIPKIKQYWGEQLKIPVIPLTSDWFERASPRIILAAQQLCNALSQVD"], "featureFilter": {"and": [ {"field": "pfam", "value": "PF00005"}, {"field": "rfam", "value": "RF00174"} ]} }' ``` **Search with features only.** Find ABC transporter domains within 5 genes of an iron-box operator (SAE feature `8581`): ```bash curl https://api.seqhub.org/api/v1/feature-contexts/search \ -H "Authorization: Bearer $SEQHUB_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "featureFilter": {"and": [ {"field": "saeFeature", "value": 8581}, {"field": "pfam", "value": "PF00005"} ]}, "window": 5 }' ``` Each search result is a neighborhood: its proteins with their annotations and Pfam domains, its taxonomy, and where each requested feature was found. **Annotate proteins.** Get a functional annotation for up to 128 sequences: ```bash curl https://api.seqhub.org/api/v1/protein-annotations \ -H "Authorization: Bearer $SEQHUB_TOKEN" \ -H "Content-Type: application/json" \ -d '{"sequences": ["MKVLAAGIVGLLLAAPAQAE", "MSTNPKPQRKTKRNTNRRPQDVKFPGG"]}' ``` The [API reference](/api) has full example responses for every endpoint. ## Endpoints ### Protein Context Search Search our database of 130,000+ microbial genomes for contigs containing proteins most similar (by embedding distance) to your query protein(s). A **contig** is a contiguous stretch of sequenced DNA encoding multiple proteins. Each result is the **neighborhood** around a match: a window of up to 10 genes on either side of the matched protein, on the same contig, with functional annotations and taxonomic information. - **Single query**: [`POST /api/v1/protein-contexts/search`](/api/protein-context-search) — find the proteins most similar to one sequence. - **Multi-query**: [`POST /api/v1/protein-contexts/search/multi-query`](/api/protein-context-search#protein-context-search-multi-query) — find neighborhoods where a combination of 1–5 proteins and optional Pfam domains, Rfam families or intergenic SAE (sparse autoencoder) features occurs together. For example: two proteins in the same neighborhood, or one protein near a cobalamin riboswitch. Both endpoints accept a `taxonomyFilter` to limit results to part of the taxonomic tree. ### Feature Context Search Search the same genomes for neighborhoods where a combination of features occurs, without a query sequence. Features are **Pfam** domains, **Rfam** RNA families and intergenic **SAE features**, combined with `and`, `or` and `not`. SAE features are interpretable features that a sparse autoencoder (SAE) learned from the gLM2 genomic language model, such as operator sites and ribosome binding sites. A neighborhood matches when the required features are found within `window` genes of each other; `window` also sets how many genes on either side are returned. - **Search**: [`POST /api/v1/feature-contexts/search`](/api/feature-context-search) — find neighborhoods where a combination of features occurs. Features are given by identifier, such as `PF00069`. Use the [vocabulary endpoints](#vocabularies) to look up valid identifiers. ### Protein Annotation Annotate a batch of protein sequences with biological function by finding their closest match in SwissProt (UniProt's curated database). Returns a functional description, accession ID, similarity score, percent identity, and query coverage for each input sequence. - **Annotate**: [`POST /api/v1/protein-annotations`](/api/protein-annotation) — annotate one or more protein sequences. ### Contig DNA Fetch the DNA sequence for any region of a contig returned by a search, such as a gene, an Rfam hit, an SAE feature's run, or a whole neighborhood. - **Fetch DNA**: [`POST /api/v1/contig-dna`](/api/contig-dna) — fetch up to 100 ranges in one call. Each range is a `sampleId` and `contigId` from a search match, plus 1-based inclusive `start` and `end` coordinates taken from the same response. Sequences are always returned on the plus strand, so reverse-complement minus-strand features yourself. Each range can be up to 100 kb. ### Vocabularies These endpoints list the identifiers you can use in feature expressions and taxonomy filters. The lists rarely change, so fetch them once and cache them. - **Pfam families**: [`GET /api/v1/pfam-families`](/api/vocabularies#list-pfam-families) — every Pfam accession and name. - **Rfam families**: [`GET /api/v1/rfam-families`](/api/vocabularies#list-rfam-families) — every Rfam accession and name. - **SAE features**: [`GET /api/v1/sae-features`](/api/vocabularies#list-sae-features) — every searchable SAE feature, with its name. - **Taxa**: [`GET /api/v1/taxa`](/api/vocabularies#list-taxa) — every taxon across the seven ranks, as `{rank, value}` pairs for use in `taxonomyFilter`. ## Authentication All requests require a personal access token (PAT). Generate one from the **API Tokens** section of your [SeqHub profile](https://seqhub.org). Pass the token in the `Authorization` header: ``` Authorization: Bearer ``` Your token is shown once when you create it. It starts with `seqhub_` — copy it exactly as shown. ## Limits **Rate limits:** 1000 requests per 7 days for each of four groups: search (protein and feature context search combined), annotation, contig DNA, and vocabularies. **Batch size:** The annotation endpoint accepts up to 128 sequences per request. If you need higher limits, let us know! Please reach out to us at [team@tatta.bio](mailto:team@tatta.bio). ## OpenAPI spec The full OpenAPI 3.1 spec is available at seqhub-public.json, for use with client generators and AI agents. The guide pages (not the API reference) are also available as plain text at llms.txt and llms-full.txt. ## Versioning The SeqHub API is currently in beta. The API is subject to breaking changes while in beta.