BLAST Search icon

BLAST Search

(2.17.0+)Code (opens in a new tab)Docs

Find protein sequence matches in Swiss-Prot and PDB.

Input

Not available for new jobs

BLAST Search is being prepared for release while scientific verification is completed.

0 credits

Output

Configure inputs to begin

Set options on the left, then click “Search proteins”.

BLAST Search webserver overview

BLAST Search runs NCBI BLAST+ 2.17.0+ protein searches against fixed Swiss-Prot or PDB protein database snapshots on ProteinIQ infrastructure. It supports blastp, blastp-fast, and blastp-short, returning hit tables, alignments, native reports, and database provenance. Queries are not sent to NCBI's public search service.

BLAST Search is currently unavailable while release verification is completed.

Pricing

Searches start at 10 credits. The exact quote is calculated before submission from the total query residues across all records in a job. FASTA headers, whitespace, and gap characters (- and .) do not count toward the quote. Optional filter and strategy files do not add query residues.

Total query residues in one jobCredits
6010
30010
1,00013
3,00022
10,00045

These examples use protein FASTA input and the current calculator. Database choice, task, search controls, and report format do not change the quote. Batch runs are quoted as separate jobs for their expanded query inputs.

Inputs

InputAccepted formatsPurpose
Query proteins (query).fasta, .fa, .fas, .faa, .txtRequired protein FASTA with one or multiple records, or plain protein sequence text.
Include/exclude GI lists (gilist, negative_gilist).txtRestrict or exclude database sequences by GI identifier.
Include/exclude sequence ID lists (seqidlist, negative_seqidlist).txtRestrict or exclude database sequences by native sequence identifier.
Include/exclude taxonomy ID lists (taxidlist, negative_taxidlist).txtRestrict or exclude taxa and, by default, their descendants.
Include/exclude IPG lists (ipglist, negative_ipglist).txtRestrict or exclude Identical Protein Group identifiers.
Native search strategy (import_search_strategy).asn1, .txtImport a native ASN.1 search strategy.

Each input accepts pasted text or one file, with a 50 MiB file limit. Staged inputs also have a combined 50 MiB submission limit. Query and optional file bytes are preserved; BLAST handles sequence parsing, ambiguous amino acids, and matrix/gap feasibility. Identifier lists use the native BLAST file syntax, normally one identifier per line.

Only one GI, sequence ID, or taxonomy include/exclude filter may be supplied, counting both files and taxonomy settings. IPG lists are separate controls. Taxonomy descendant expansion cannot be disabled when a GI or sequence ID list is supplied.

When a strategy is imported, BLAST reads its search settings from that file. The separately submitted query and selected hosted database take precedence. Import and export are mutually exclusive, so strategy.asn1 retains the imported strategy unchanged instead of exporting a replacement. See NCBI's strategy documentation.

Batch mode expands up to 10 query inputs into separate jobs. Multiple records within one query input remain together in that job; optional lists and settings are shared across the batch. Workflows can consume the native files and HSP rows.

Databases and attribution

DatabaseSnapshot dateSequencesLettersTerms
NCBI Swiss-Prot (swissprot-20260929)2026-09-29487,728186,063,409UniProtKB/Swiss-Prot, UniProt Consortium, CC BY 4.0.
PDB proteins, NCBI pdbaa (pdbaa-20260929)2026-09-29190,52955,496,864PDB archive contributors, CC0 1.0.

These are NCBI-distributed snapshots, rather than live database searches. Prepared sequence and taxonomy files are used without scientific modification. Snapshot dates, archive checksums, individual file checksums, and database identity accompany successful results.

BLAST is public-domain software. UniProt and NCBI cannot grant unrestricted permission for third-party patents or other rights covering individual data. The exact BLAST license and database attribution accompany the downloads.

Settings

Empty optional fields leave the corresponding option unset, preserving BLAST's task-dependent defaults. Numeric controls shown as text fields accept numeric text. Native BLAST checks additional task, matrix, and gap-cost constraints.

Search selection

ParameterTypeDefaultDescription
databaseenumswissprot-20260929swissprot-20260929 or pdbaa-20260929; both are immutable hosted snapshots.
taskenumblastpblastp for standard protein searches, blastp-fast for the faster search task, or blastp-short for short protein queries.
evaluestringoptional; native task defaultNonnegative E-value cutoff for retaining hits: 10 for blastp/blastp-fast, 20000 for blastp-short.

With scoring fields empty, blastp-short uses word size 2, matrix PAM30, gap opening cost 9, and gap extension cost 1. These values are task defaults, not values automatically entered into the form.

Advanced search settings

ParameterTypeDefaultDescription
word_sizestringoptionalSeed-word length; integer at least 2.
gapopenstringoptionalInteger cost of opening an alignment gap.
gapextendstringoptionalInteger cost of extending an alignment gap.
thresholdstringoptionalMinimum seed-word score for the lookup table; number at least 0.
qcov_hsp_percstringoptionalMinimum query coverage per HSP, from 0 to 100 percent.
max_hspsstringoptionalMaximum HSPs retained per query/subject pair; integer at least 1.
culling_limitstringoptionalRemove a hit whose query region is contained in this many higher-scoring hits; integer at least 0, with 0 disabling culling.
best_hit_overhangstringoptionalBest-Hit filtering allowance for an HSP extending beyond a competing query region; number greater than 0 and less than 0.5.
best_hit_score_edgestringoptionalBest-Hit filtering margin for comparing score per alignment length; number greater than 0 and less than 0.5.
max_target_seqsstringoptionalMaximum aligned target sequences retained; integer at least 1, native default 500.
dbsizestringoptionalInteger effective database length used for statistics.
searchspstringoptionalEffective search-space size used for statistics; integer at least 0.
xdrop_ungapstringoptionalScore drop allowed during ungapped extension, in bits.
xdrop_gapstringoptionalScore drop allowed during preliminary gapped extension, in bits.
xdrop_gap_finalstringoptionalScore drop allowed during final gapped alignment, in bits.
window_sizestringoptionalInteger window for multiple seed hits; at least 0, with 0 selecting the one-hit algorithm.
matrixstringoptionalNative scoring-matrix name, such as BLOSUM62 or PAM30; compatible gap costs depend on the matrix.
comp_based_statsstringoptional0/F disables composition statistics; 1 uses composition statistics; 2/D/T uses conditional score adjustment; 3 uses unconditional adjustment. Letters also accept lowercase.
segstringoptionalQuery low-complexity filtering: no, yes, or window locut hicut, such as 12 2.2 2.5.
soft_maskingstringoptionaltrue masks seed finding while allowing extension through masked regions; false selects hard masking.
query_locstringoptionalQuery region as inclusive 1-based start-stop, such as 1-100.
taxidsstringoptionalComma-separated taxonomy IDs to include, with descendants unless expansion is disabled.
negative_taxidsstringoptionalComma-separated taxonomy IDs to exclude, with descendants unless expansion is disabled.
db_soft_maskstringoptionalDatabase masking algorithm ID used for seed finding.
db_hard_maskstringoptionalDatabase masking algorithm ID used for both seed finding and extension.
lcase_maskingbooleanfalseTreat lowercase sequence regions as masks.
ungappedbooleanfalsePerform ungapped alignments only.
use_sw_tbackbooleanfalseUse Smith-Waterman traceback for locally optimal alignments.
parse_deflinesbooleanfalseParse native sequence identifiers in FASTA headers.
subject_besthitbooleanfalseRetain the best HSP for each nonoverlapping query region.
no_taxid_expansionbooleanfalseUse only the supplied taxonomy IDs without adding descendants.

Culling and Best-Hit controls are mutually exclusive, as are soft and hard database masks. max_target_seqs cannot be combined with num_descriptions or num_alignments. Database masking requires an algorithm present in the selected snapshot; an unavailable algorithm produces a native BLAST error. NCBI describes the Best-Hit filtering criteria in detail.

Native report format

ParameterTypeDefaultDescription
output_formatenum0Selected native report: 0 pairwise; 1/2 query-anchored with/without identities; 3/4 flat query-anchored with/without identities; 5 XML; 6 tabular; 7 tabular with comments; 8 text ASN.1; 9 binary ASN.1; 10 CSV; 11 archive; 12 Seqalign JSON; 13 multiple-file JSON; 14 multiple-file XML2; 15 single-file JSON; 16 single-file XML2; 18 organism report; 20 CSV with headers.
output_fieldsstringoptionalSpace-separated native field names for formats 6, 7, 10, or 20; optional delim= must precede field names, such as delim=@ qacc sacc score.
num_descriptionsstringoptionalNumber of one-line subject descriptions for formats 0 to 4; integer at least 0, native default 500.
num_alignmentsstringoptionalNumber of subject sequences with alignments; integer at least 0, native default 250; also passed to the fixed native report searches.
line_lengthstringoptionalAlignment report line width for formats 0 to 4; integer at least 1, native default 60.
sorthitsstringoptionalHit sort for formats 0 to 4: 0 E-value, 1 bit score, 2 total score, 3 identity percentage, 4 query coverage.
sorthspsstringoptionalHSP sort for format 0: 0 E-value, 1 score, 2 query start, 3 identity percentage, 4 subject start.
show_gisbooleanfalseInclude NCBI GI identifiers in report headers where available.
htmlbooleanfalseRequest native HTML formatting.

With output_fields empty, tabular formats use BLAST's standard columns. This setting changes the selected report, while hits.tsv retains its fixed 35-column schema. BLAST determines which report options apply to each format.

Outputs

The Hits tab contains one row per native high-scoring segment pair (HSP), in BLAST's original order. Alignments presents the pairwise report, and Files provides native downloads.

DownloadContents
archive.asn1Native BLAST search archive.
pairwise.txtNative pairwise alignment report and search statistics.
hits.tsvFixed 35-column native table with identifiers, accessions, scores, coverage, coordinates, aligned sequences, taxonomy, and BTOP.
results.xml, results.jsonNative XML and single-file JSON reports.
strategy.asn1Exported search strategy, or the unchanged imported strategy when provided.
selected-report.* and associated filesReport selected by output_format, including child files produced by multiple-file formats.
query.fasta and optional input filesOriginal submitted query, lists, and imported strategy bytes.
stdout.txt, stderr.txtNative output and warning/error logs; either may be empty on success.
provenance.jsonActual commands, settings, BLAST version, executable checksum, query/list checksums, and database snapshot identity.
BLAST-license.txt, attribution.txtBLAST public-domain notice and database attribution.

A valid search with no matches retains its empty hit table and native reports. Failed searches retain available partial files and logs.

Native reports are generated with separate blastp invocations using the same query, database, and search settings. This preserves statistics that archive replay can omit. Report formatting controls apply to the selected report, except that num_alignments is also supplied to the fixed report searches. ProteinIQ does not calculate replacement scientific scores.

Understanding results

FieldMeaning
qseqid, sseqid, qaccver, saccverNative query/subject identifiers and versioned accessions; numeric-looking identifiers remain text.
pident, nidentIdentical aligned residues as a percentage and count.
ppos, positivePositive-scoring aligned residue pairs as a percentage and count.
evalueExpected number of matches with this score or better occurring by chance in the searched space; smaller values indicate stronger statistical evidence.
bitscore, scoreNormalized alignment score in bits and unnormalized native score.
length, mismatch, gapopen, gapsAlignment length, mismatch count, gap-opening count, and total gaps.
qstart, qend, sstart, sendInclusive 1-based alignment coordinates in the query and subject.
qcovs, qcovhspQuery coverage percentages per subject and per individual HSP.
qseq, sseq, btopAligned query/subject sequences and native BLAST traceback operations.
staxids, sscinames, stitle, salltitlesSubject taxonomy IDs, scientific names, and sequence descriptions.

A sequence match alone does not establish biological function. Coverage, alignment extent, database context, and annotation should be considered alongside scores.

The NCBI BLAST+ manual defines native options and reports. The method is described by Altschul et al. (1997).

Table of contents

Related tools

HMMER

HMMER

Sensitive sequence homology search using profile hidden Markov models

MAFFT

MAFFT

Align protein or nucleotide sequences with selectable accuracy and speed trade-offs.

MMseqs2

MMseqs2

Search and cluster protein or nucleotide sequences for homology discovery at large scale.

MUSCLE5

MUSCLE5

Align multiple protein or nucleotide sequences with high-accuracy PPP refinement.

ANARCI

ANARCI

Number antibody and T cell receptor sequences with multiple numbering schemes

IgBLAST

IgBLAST

Analyze antibody and T cell receptor variable domain sequences

StringZilla v5

StringZilla v5

Hardware-accelerated edit distances and global or local sequence scores

FoldSeek

FoldSeek

Search AlphaFold DB, compare structures, or cluster by 3D similarity

USAlign

USAlign

Universal structure alignment for proteins, RNA, and DNA molecules

MUMmer4

MUMmer4

Align and compare whole genomes to detect SNPs, indels, and structural variants.