
BLAST Search
Find protein sequence matches in Swiss-Prot and PDB.
Input
Not available for new jobs
BLAST Search webserver overview
BLAST Search runs NCBI BLAST+ 2.17.0+ protein searches against fixed Swiss-Prot or PDB protein database snapshots on ProteinIQ infrastructure. It supports blastp, blastp-fast, and blastp-short, returning hit tables, alignments, native reports, and database provenance. Queries are not sent to NCBI's public search service.
BLAST Search is currently unavailable while release verification is completed.
Pricing
Searches start at 10 credits. The exact quote is calculated before submission from the total query residues across all records in a job. FASTA headers, whitespace, and gap characters (- and .) do not count toward the quote. Optional filter and strategy files do not add query residues.
| Total query residues in one job | Credits |
|---|---|
| 60 | 10 |
| 300 | 10 |
| 1,000 | 13 |
| 3,000 | 22 |
| 10,000 | 45 |
These examples use protein FASTA input and the current calculator. Database choice, task, search controls, and report format do not change the quote. Batch runs are quoted as separate jobs for their expanded query inputs.
Inputs
| Input | Accepted formats | Purpose |
|---|---|---|
Query proteins (query) | .fasta, .fa, .fas, .faa, .txt | Required protein FASTA with one or multiple records, or plain protein sequence text. |
Include/exclude GI lists (gilist, negative_gilist) | .txt | Restrict or exclude database sequences by GI identifier. |
Include/exclude sequence ID lists (seqidlist, negative_seqidlist) | .txt | Restrict or exclude database sequences by native sequence identifier. |
Include/exclude taxonomy ID lists (taxidlist, negative_taxidlist) | .txt | Restrict or exclude taxa and, by default, their descendants. |
Include/exclude IPG lists (ipglist, negative_ipglist) | .txt | Restrict or exclude Identical Protein Group identifiers. |
Native search strategy (import_search_strategy) | .asn1, .txt | Import a native ASN.1 search strategy. |
Each input accepts pasted text or one file, with a 50 MiB file limit. Staged inputs also have a combined 50 MiB submission limit. Query and optional file bytes are preserved; BLAST handles sequence parsing, ambiguous amino acids, and matrix/gap feasibility. Identifier lists use the native BLAST file syntax, normally one identifier per line.
Only one GI, sequence ID, or taxonomy include/exclude filter may be supplied, counting both files and taxonomy settings. IPG lists are separate controls. Taxonomy descendant expansion cannot be disabled when a GI or sequence ID list is supplied.
When a strategy is imported, BLAST reads its search settings from that file. The separately submitted query and selected hosted database take precedence. Import and export are mutually exclusive, so strategy.asn1 retains the imported strategy unchanged instead of exporting a replacement. See NCBI's strategy documentation.
Batch mode expands up to 10 query inputs into separate jobs. Multiple records within one query input remain together in that job; optional lists and settings are shared across the batch. Workflows can consume the native files and HSP rows.
Databases and attribution
These are NCBI-distributed snapshots, rather than live database searches. Prepared sequence and taxonomy files are used without scientific modification. Snapshot dates, archive checksums, individual file checksums, and database identity accompany successful results.
BLAST is public-domain software. UniProt and NCBI cannot grant unrestricted permission for third-party patents or other rights covering individual data. The exact BLAST license and database attribution accompany the downloads.
Settings
Empty optional fields leave the corresponding option unset, preserving BLAST's task-dependent defaults. Numeric controls shown as text fields accept numeric text. Native BLAST checks additional task, matrix, and gap-cost constraints.
Search selection
| Parameter | Type | Default | Description |
|---|---|---|---|
database | enum | swissprot-20260929 | swissprot-20260929 or pdbaa-20260929; both are immutable hosted snapshots. |
task | enum | blastp | blastp for standard protein searches, blastp-fast for the faster search task, or blastp-short for short protein queries. |
evalue | string | optional; native task default | Nonnegative E-value cutoff for retaining hits: 10 for blastp/blastp-fast, 20000 for blastp-short. |
With scoring fields empty, blastp-short uses word size 2, matrix PAM30, gap opening cost 9, and gap extension cost 1. These values are task defaults, not values automatically entered into the form.
Advanced search settings
| Parameter | Type | Default | Description |
|---|---|---|---|
word_size | string | optional | Seed-word length; integer at least 2. |
gapopen | string | optional | Integer cost of opening an alignment gap. |
gapextend | string | optional | Integer cost of extending an alignment gap. |
threshold | string | optional | Minimum seed-word score for the lookup table; number at least 0. |
qcov_hsp_perc | string | optional | Minimum query coverage per HSP, from 0 to 100 percent. |
max_hsps | string | optional | Maximum HSPs retained per query/subject pair; integer at least 1. |
culling_limit | string | optional | Remove a hit whose query region is contained in this many higher-scoring hits; integer at least 0, with 0 disabling culling. |
best_hit_overhang | string | optional | Best-Hit filtering allowance for an HSP extending beyond a competing query region; number greater than 0 and less than 0.5. |
best_hit_score_edge | string | optional | Best-Hit filtering margin for comparing score per alignment length; number greater than 0 and less than 0.5. |
max_target_seqs | string | optional | Maximum aligned target sequences retained; integer at least 1, native default 500. |
dbsize | string | optional | Integer effective database length used for statistics. |
searchsp | string | optional | Effective search-space size used for statistics; integer at least 0. |
xdrop_ungap | string | optional | Score drop allowed during ungapped extension, in bits. |
xdrop_gap | string | optional | Score drop allowed during preliminary gapped extension, in bits. |
xdrop_gap_final | string | optional | Score drop allowed during final gapped alignment, in bits. |
window_size | string | optional | Integer window for multiple seed hits; at least 0, with 0 selecting the one-hit algorithm. |
matrix | string | optional | Native scoring-matrix name, such as BLOSUM62 or PAM30; compatible gap costs depend on the matrix. |
comp_based_stats | string | optional | 0/F disables composition statistics; 1 uses composition statistics; 2/D/T uses conditional score adjustment; 3 uses unconditional adjustment. Letters also accept lowercase. |
seg | string | optional | Query low-complexity filtering: no, yes, or window locut hicut, such as 12 2.2 2.5. |
soft_masking | string | optional | true masks seed finding while allowing extension through masked regions; false selects hard masking. |
query_loc | string | optional | Query region as inclusive 1-based start-stop, such as 1-100. |
taxids | string | optional | Comma-separated taxonomy IDs to include, with descendants unless expansion is disabled. |
negative_taxids | string | optional | Comma-separated taxonomy IDs to exclude, with descendants unless expansion is disabled. |
db_soft_mask | string | optional | Database masking algorithm ID used for seed finding. |
db_hard_mask | string | optional | Database masking algorithm ID used for both seed finding and extension. |
lcase_masking | boolean | false | Treat lowercase sequence regions as masks. |
ungapped | boolean | false | Perform ungapped alignments only. |
use_sw_tback | boolean | false | Use Smith-Waterman traceback for locally optimal alignments. |
parse_deflines | boolean | false | Parse native sequence identifiers in FASTA headers. |
subject_besthit | boolean | false | Retain the best HSP for each nonoverlapping query region. |
no_taxid_expansion | boolean | false | Use only the supplied taxonomy IDs without adding descendants. |
Culling and Best-Hit controls are mutually exclusive, as are soft and hard database masks. max_target_seqs cannot be combined with num_descriptions or num_alignments. Database masking requires an algorithm present in the selected snapshot; an unavailable algorithm produces a native BLAST error. NCBI describes the Best-Hit filtering criteria in detail.
Native report format
| Parameter | Type | Default | Description |
|---|---|---|---|
output_format | enum | 0 | Selected native report: 0 pairwise; 1/2 query-anchored with/without identities; 3/4 flat query-anchored with/without identities; 5 XML; 6 tabular; 7 tabular with comments; 8 text ASN.1; 9 binary ASN.1; 10 CSV; 11 archive; 12 Seqalign JSON; 13 multiple-file JSON; 14 multiple-file XML2; 15 single-file JSON; 16 single-file XML2; 18 organism report; 20 CSV with headers. |
output_fields | string | optional | Space-separated native field names for formats 6, 7, 10, or 20; optional delim= must precede field names, such as delim=@ qacc sacc score. |
num_descriptions | string | optional | Number of one-line subject descriptions for formats 0 to 4; integer at least 0, native default 500. |
num_alignments | string | optional | Number of subject sequences with alignments; integer at least 0, native default 250; also passed to the fixed native report searches. |
line_length | string | optional | Alignment report line width for formats 0 to 4; integer at least 1, native default 60. |
sorthits | string | optional | Hit sort for formats 0 to 4: 0 E-value, 1 bit score, 2 total score, 3 identity percentage, 4 query coverage. |
sorthsps | string | optional | HSP sort for format 0: 0 E-value, 1 score, 2 query start, 3 identity percentage, 4 subject start. |
show_gis | boolean | false | Include NCBI GI identifiers in report headers where available. |
html | boolean | false | Request native HTML formatting. |
With output_fields empty, tabular formats use BLAST's standard columns. This setting changes the selected report, while hits.tsv retains its fixed 35-column schema. BLAST determines which report options apply to each format.
Outputs
The Hits tab contains one row per native high-scoring segment pair (HSP), in BLAST's original order. Alignments presents the pairwise report, and Files provides native downloads.
| Download | Contents |
|---|---|
archive.asn1 | Native BLAST search archive. |
pairwise.txt | Native pairwise alignment report and search statistics. |
hits.tsv | Fixed 35-column native table with identifiers, accessions, scores, coverage, coordinates, aligned sequences, taxonomy, and BTOP. |
results.xml, results.json | Native XML and single-file JSON reports. |
strategy.asn1 | Exported search strategy, or the unchanged imported strategy when provided. |
selected-report.* and associated files | Report selected by output_format, including child files produced by multiple-file formats. |
query.fasta and optional input files | Original submitted query, lists, and imported strategy bytes. |
stdout.txt, stderr.txt | Native output and warning/error logs; either may be empty on success. |
provenance.json | Actual commands, settings, BLAST version, executable checksum, query/list checksums, and database snapshot identity. |
BLAST-license.txt, attribution.txt | BLAST public-domain notice and database attribution. |
A valid search with no matches retains its empty hit table and native reports. Failed searches retain available partial files and logs.
Native reports are generated with separate blastp invocations using the same query, database, and search settings. This preserves statistics that archive replay can omit. Report formatting controls apply to the selected report, except that num_alignments is also supplied to the fixed report searches. ProteinIQ does not calculate replacement scientific scores.
Understanding results
| Field | Meaning |
|---|---|
qseqid, sseqid, qaccver, saccver | Native query/subject identifiers and versioned accessions; numeric-looking identifiers remain text. |
pident, nident | Identical aligned residues as a percentage and count. |
ppos, positive | Positive-scoring aligned residue pairs as a percentage and count. |
evalue | Expected number of matches with this score or better occurring by chance in the searched space; smaller values indicate stronger statistical evidence. |
bitscore, score | Normalized alignment score in bits and unnormalized native score. |
length, mismatch, gapopen, gaps | Alignment length, mismatch count, gap-opening count, and total gaps. |
qstart, qend, sstart, send | Inclusive 1-based alignment coordinates in the query and subject. |
qcovs, qcovhsp | Query coverage percentages per subject and per individual HSP. |
qseq, sseq, btop | Aligned query/subject sequences and native BLAST traceback operations. |
staxids, sscinames, stitle, salltitles | Subject taxonomy IDs, scientific names, and sequence descriptions. |
A sequence match alone does not establish biological function. Coverage, alignment extent, database context, and annotation should be considered alongside scores.
The NCBI BLAST+ manual defines native options and reports. The method is described by Altschul et al. (1997).
Related tools
HMMER
Sensitive sequence homology search using profile hidden Markov models
MAFFT
Align protein or nucleotide sequences with selectable accuracy and speed trade-offs.
MMseqs2
Search and cluster protein or nucleotide sequences for homology discovery at large scale.
MUSCLE5
Align multiple protein or nucleotide sequences with high-accuracy PPP refinement.
ANARCI
Number antibody and T cell receptor sequences with multiple numbering schemes
IgBLAST
Analyze antibody and T cell receptor variable domain sequences
StringZilla v5
Hardware-accelerated edit distances and global or local sequence scores
FoldSeek
Search AlphaFold DB, compare structures, or cluster by 3D similarity
USAlign
Universal structure alignment for proteins, RNA, and DNA molecules
MUMmer4
Align and compare whole genomes to detect SNPs, indels, and structural variants.