Protein Context Search
Protein Context Search
Search our database of 130,000+ microbial genomes for protein contexts containing a protein most similar (by embedding distance) to your query sequence. A protein context is the set of proteins encoded on the same contig (genomic neighborhood). Returns taxonomic information and functionally annotated proteins within each matching context.
Protein Context Search › Request Body
sequenceThe protein sequence to search for (max length 15000). Standard 20 and ambiguous B, X, Z, U residues are allowed.
Only return matches from genomes matching this taxonomy expression: a single {rank, value}, or several combined with and, or and not. At most 10 taxa. See /api/v1/taxa.
diversityLevelHow diverse the results are. Higher levels search only one representative protein per cluster of similar sequences, so near-identical proteins don't crowd out the results. low (default): clusters at 90% identity, returning the most similar results. medium: 70% identity. high: 50% identity. max: 30% identity, the most diverse results. For distant homologs, use high or max: at low, results can fill up with near-identical close matches.
maxResultsMaximum number of results to return, up to 100.
Protein Context Search › Responses
Successful Response
Genomic neighborhoods around the closest matches to the query sequence, most similar first.
Protein Context Search (Multi-Query)
Search our database of 130,000+ microbial genomes for genomic neighborhoods where a combination of proteins and features occurs together. A query is 1-5 protein sequences, optionally combined with Pfam domains, Rfam families and intergenic SAE features (featureFilter) and a taxonomy filter (taxonomyFilter). Each protein is matched by embedding similarity, and only neighborhoods that contain a match for every protein and satisfy the filters are returned.
With a single sequence, this works like single-query search, plus filters and the option to return every SAE feature and Rfam hit in each neighborhood (includeOtherSaes, includeOtherRfams).
Each entry in sequences is either a plain sequence string or an object with an optional id and pfamFilter. All entries must use the same form, so to add a pfamFilter to one query, send every query as an object.
Protein Context Search (Multi-Query) › Request Body
1-5 protein sequences to search for (each max 15000 characters). Standard 20 and ambiguous B, X, Z, U residues are allowed.
Either all plain sequence strings, or all ProteinQuery objects. A ProteinQuery adds an optional id and an optional pfamFilter on the protein that query matches.
Only return matches whose window satisfies this expression of Pfam, Rfam and SAE features, combined with and, or and not. At most 25 features. Valid identifiers are listed by /api/v1/pfam-families, /api/v1/rfam-families and /api/v1/sae-features.
Only return matches from genomes matching this taxonomy expression: a single {rank, value}, or several combined with and, or and not. At most 10 taxa. See /api/v1/taxa.
diversityLevelHow diverse the results are. Higher levels search only one representative protein per cluster of similar sequences, so near-identical proteins don't crowd out the results. low (default): clusters at 90% identity, returning the most similar results. medium: 70% identity. high: 50% identity. max: 30% identity, the most diverse results. For distant homologs, use high or max: at low, results can fill up with near-identical close matches.
maxResultsMaximum number of results to return, up to 100.
includeOtherSaesAlso return every SAE feature in each match's window, not just those named in featureFilter. Makes the response much larger.
includeOtherRfamsAlso return every Rfam hit in each match's window, not just those named in featureFilter.
Protein Context Search (Multi-Query) › Responses
Successful Response
Genomic neighborhoods where every query sequence found a match, best first.