
FASTA converter
Paste or upload a sequence file in any common format and download FASTA. The format is detected automatically, and nothing leaves your browser.
Input
How to convert a file to FASTA format?
Paste a sequence file or upload it, and the FASTA converter detects the format and writes a .fasta file with one record per sequence. GenBank, EMBL, FASTQ, Clustal, Stockholm, PHYLIP, NEXUS, MSF, PDB, mmCIF, CSV, and plain text all convert without choosing a format first, and nothing is uploaded to a server.
A GenBank record keeps its versioned accession and definition line as the FASTA header, while the feature table and position numbers are dropped:
LOCUS DEMO0001 60 bp DNA linear SYN 01-JAN-2024
DEFINITION Synthetic construct with a short open reading frame.
ACCESSION DEMO0001
VERSION DEMO0001.1
FEATURES Location/Qualifiers
CDS 4..36
/gene="demoA"
ORIGIN
1 ccgatggctt cgaaaggcga agagctgttc taataggtca ttgaccgcat gcatgcaaat
//>DEMO0001.1 Synthetic construct with a short open reading frame.
CCGATGGCTTCGAAAGGCGAAGAGCTGTTCTAATAGGTCATTGACCGCATGCATGCAAATAlignment formats become aligned FASTA when Preserve alignment gaps is on. This Clustal input, converted with gaps kept and no line wrapping:
CLUSTAL W (1.83) multiple sequence alignment
KRAS_HUMAN MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
HRAS_HUMAN MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
RHEB_HUMAN MPQSKSRKIAILGYRSVGKSSLTIQFVEGQFVDSYDPTIENTF>KRAS_HUMAN
MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
>HRAS_HUMAN
MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
>RHEB_HUMAN
MPQSKSRKIAILGYRSVGKSSLTIQFVEGQFVDSYDPTIENTFWith the default settings the six gap characters are removed instead, giving ungapped sequences, and a notice says how many were dropped.
Supported formats
The format is read from the content, not the file extension, so a GenBank record saved as .txt still converts as GenBank.
| Format | Recognized by | FASTA header comes from |
|---|---|---|
GenBank (.gb, .gbk, .gbff) | LOCUS and ORIGIN lines | VERSION (or ACCESSION, then LOCUS) plus DEFINITION |
EMBL (.embl) | ID and SQ lines | First AC accession plus DE description |
FASTQ (.fastq, .fq) | @ header, sequence, +, and a quality line of equal length | The read header; quality lines are dropped |
Clustal (.aln) | A first line starting with CLUSTAL, MUSCLE, or MAFFT | Sequence names; interleaved blocks are joined |
Stockholm (.sto) | # STOCKHOLM on the first line | Sequence names; #= markup lines are skipped |
PHYLIP (.phy) | A first line with sequence and site counts | Sequence names; sequential and interleaved layouts |
NEXUS (.nex) | #NEXUS and a matrix block | Taxon names, including quoted names with spaces |
MSF (.msf) | An MSF: line and a // divider | Sequence names |
PDB (.pdb) | ATOM or SEQRES records | Structure ID and chain, one record per chain |
mmCIF (.cif) | A data_ block with _atom_site coordinates | Entry ID and chain from the first model |
| CSV or TSV | A delimiter on nearly every row and a column of sequences | The first column that isn't the sequence, or a column named id or name |
| FASTA or plain text | Anything else | Existing > headers, or name lines found in the text |
Plain text is read with the same rules as TXT to FASTA, which also explains how name lines, wrapped lines, and lists are handled. Word (.docx) and PDF uploads are converted to text in the browser first.
Settings
| Setting | Description |
|---|---|
Sequence type | Auto-detect (default), DNA, RNA, Nucleotide (DNA or RNA), or Protein. A specific type removes letters outside that alphabet and reports them. A U in DNA mode or a T in RNA mode skips that sequence with an explanation. |
Sequence names | Use names from input (default) keeps the headers listed in the table above. Numbered (seq_1, seq_2, …) replaces them. |
Name prefix | Shown for Numbered names. The text before the number. Default: seq. |
Line wrapping | 80 characters per line (standard) (default), 60 characters per line, No wrapping (single line), or Custom characters per line. |
Custom line length | Shown for custom wrapping. 0 turns wrapping off; an invalid value uses 80. Default: 80. |
Under Advanced settings:
| Setting | Description |
|---|---|
Preserve alignment gaps | Keeps - and . so converted alignments stay column-aligned. Default: off. |
Preserve stop codons | Keeps a terminal *. Default: off. |
Case format | Uppercase (default), lowercase, or Preserve original. Stockholm insert columns are lowercase, so keep the original case when that distinction matters. |
Validation | Lenient (clean and report) (default), Strict (reject anything that needs cleanup), or Off. |
Keep descriptions | Shown for Numbered names. Writes the original header after the new name, for example >seq_1 DEMO0001.1 Synthetic construct with a short open reading frame. Default: off. |
Results
The preview shows the FASTA text under a summary line that names the detected format, such as 1 sequence · 60 residues · DNA · read as GenBank. Notices above it list anything that changed the sequences, such as removed gaps or stop markers. Files holds the .fasta file, named after the uploaded file, and a run.log with the detected format, statistics, and cleanup counts.
What is FASTA format?
FASTA is a plain-text format for nucleotide and protein sequences: a header line starting with >, then the sequence as one-letter codes. It comes from the FASTA similarity-search program by Pearson and Lipman (1988), and its simplicity made it the default input for BLAST, aligners, structure predictors, and most sequence tools.
Richer formats hold more than FASTA can. GenBank and EMBL carry annotation and feature tables, FASTQ carries per-base quality scores, and Clustal, Stockholm, PHYLIP, and NEXUS encode alignments for specific programs. Converting to FASTA keeps the identifiers and sequences and drops the rest, which is usually what the next tool needs. When the annotation matters, extract it first: GenBank to FASTA writes coding sequences or protein translations, and the GenBank Feature Extractor pulls genes and other features.
Which converter should I use?
| Goal | Tool |
|---|---|
| Convert any sequence file to FASTA without choosing a format | FASTA converter |
| Turn notes, pasted text, or a Word file into FASTA | TXT to FASTA |
| Extract CDS or translated proteins from GenBank | GenBank to FASTA |
| Filter FASTQ reads by quality while converting | FASTQ to FASTA |
| Choose CSV columns and header prefixes | CSV to FASTA |
| Choose chains or use SEQRES from a PDB file | PDB to FASTA |
| Turn FASTA into a spreadsheet or plain text | FASTA to CSV |
Converted sequences can go straight into multiple sequence alignment with Clustal Omega or be divided with the FASTA splitter.
FAQ
How do I convert GenBank to FASTA?
Paste the GenBank record or upload the .gb file. Each record becomes a FASTA entry named by its versioned accession, followed by the definition line.
How do I convert a Clustal alignment to FASTA?
Paste the .aln file and turn on Preserve alignment gaps to keep the alignment. Interleaved blocks are joined per sequence, and the conservation line is ignored.
Can I convert FASTQ to FASTA online?
Yes. Headers are kept and the + and quality lines are dropped. For quality filtering during conversion, use FASTQ to FASTA.
Can I convert a PDB file to FASTA?
Yes. Each chain becomes a record read from its coordinates, or from SEQRES when the file has no coordinates. Chain selection and other options are in PDB to FASTA.
Is my sequence data uploaded?
No. Files are read and converted in your browser, including Word and PDF text extraction.
Related tools

CSV to FASTA
Convert CSV or TSV sequence tables to FASTA with automatic delimiter detection and column mapping.

TXT to FASTA converter
Paste or upload DNA, RNA, or protein text and download FASTA. Sequence names, wrapped lines, and tables are detected automatically, and nothing leaves your browser.

GenBank Feature Extractor
Extract CDS, mRNA, gene, rRNA, and tRNA features from GenBank records into FASTA.

FASTA to CSV converter
Paste or upload FASTA and download a CSV or TSV table for spreadsheets, or a plain text file with one sequence per line.

FASTA to FASTQ Converter
Convert FASTA sequences to FASTQ with configurable mock Phred quality scores.

FASTQ to FASTA converter
Convert FASTQ sequencing reads to FASTA while preserving headers and sequence case by default.

GenBank to FASTA Converter
Convert GenBank records to FASTA by extracting primary sequences, CDS, or translations.

DNA to Protein Converter
Translate DNA sequences into proteins across reading frames and genetic code tables.

DNA to RNA converter
Transcribe DNA sequences to RNA by replacing thymine with uracil.

Protein to DNA converter
Reverse translate protein sequences into DNA with codon optimization and GC-content controls.