SeqHub API
The SeqHub API lets you search 130,000+ microbial genomes by any combination of protein sequences, Pfam domains, Rfam RNA families and gLM2-derived SAE features, and annotate proteins.
Quickstart
Every request needs a personal access token (see Authentication).
Search with a protein and features. Find neighborhoods with a protein similar to E. coli's vitamin B12-binding protein BtuF, an ABC transporter domain (PF00005) and a cobalamin riboswitch (RF00174):
Code
Search with features only. Find ABC transporter domains within 5 genes of an iron-box operator (SAE feature 8581):
Code
Each search result is a neighborhood: its proteins with their annotations and Pfam domains, its taxonomy, and where each requested feature was found.
Annotate proteins. Get a functional annotation for up to 128 sequences:
Code
The API reference has full example responses for every endpoint.
Endpoints
Protein Context Search
Search our database of 130,000+ microbial genomes for contigs containing proteins most similar (by embedding distance) to your query protein(s). A contig is a contiguous stretch of sequenced DNA encoding multiple proteins. Each result is the neighborhood around a match: a window of up to 10 genes on either side of the matched protein, on the same contig, with functional annotations and taxonomic information.
- Single query:
POST /api/v1/protein-contexts/search— find the proteins most similar to one sequence. - Multi-query:
POST /api/v1/protein-contexts/search/multi-query— find neighborhoods where a combination of 1–5 proteins and optional Pfam domains, Rfam families or intergenic SAE (sparse autoencoder) features occurs together. For example: two proteins in the same neighborhood, or one protein near a cobalamin riboswitch.
Both endpoints accept a taxonomyFilter to limit results to part of the taxonomic tree.
Feature Context Search
Search the same genomes for neighborhoods where a combination of features occurs, without a query sequence. Features are Pfam domains, Rfam RNA families and intergenic SAE features, combined with and, or and not. SAE features are interpretable features that a sparse autoencoder (SAE) learned from the gLM2 genomic language model, such as operator sites and ribosome binding sites. A neighborhood matches when the required features are found within window genes of each other; window also sets how many genes on either side are returned.
- Search:
POST /api/v1/feature-contexts/search— find neighborhoods where a combination of features occurs.
Features are given by identifier, such as PF00069. Use the vocabulary endpoints to look up valid identifiers.
Protein Annotation
Annotate a batch of protein sequences with biological function by finding their closest match in SwissProt (UniProt's curated database). Returns a functional description, accession ID, similarity score, percent identity, and query coverage for each input sequence.
- Annotate:
POST /api/v1/protein-annotations— annotate one or more protein sequences.
Contig DNA
Fetch the DNA sequence for any region of a contig returned by a search, such as a gene, an Rfam hit, an SAE feature's run, or a whole neighborhood.
- Fetch DNA:
POST /api/v1/contig-dna— fetch up to 100 ranges in one call.
Each range is a sampleId and contigId from a search match, plus 1-based inclusive start and end coordinates taken from the same response. Sequences are always returned on the plus strand, so reverse-complement minus-strand features yourself. Each range can be up to 100 kb.
Vocabularies
These endpoints list the identifiers you can use in feature expressions and taxonomy filters. The lists rarely change, so fetch them once and cache them.
- Pfam families:
GET /api/v1/pfam-families— every Pfam accession and name. - Rfam families:
GET /api/v1/rfam-families— every Rfam accession and name. - SAE features:
GET /api/v1/sae-features— every searchable SAE feature, with its name. - Taxa:
GET /api/v1/taxa— every taxon across the seven ranks, as{rank, value}pairs for use intaxonomyFilter.
Authentication
All requests require a personal access token (PAT). Generate one from the API Tokens section of your SeqHub profile.
Pass the token in the Authorization header:
Code
Your token is shown once when you create it. It starts with seqhub_ — copy it exactly as shown.
Limits
Rate limits: 1000 requests per 7 days for each of four groups: search (protein and feature context search combined), annotation, contig DNA, and vocabularies.
Batch size: The annotation endpoint accepts up to 128 sequences per request.
If you need higher limits, let us know! Please reach out to us at team@tatta.bio.
OpenAPI spec
The full OpenAPI 3.1 spec is available at seqhub-public.json, for use with client generators and AI agents. The guide pages (not the API reference) are also available as plain text at llms.txt and llms-full.txt.
Versioning
The SeqHub API is currently in beta. The API is subject to breaking changes while in beta.