FASTA converter icon

FASTA converter

Docs

Paste or upload a sequence file in any common format and download FASTA. The format is detected automatically, and nothing leaves your browser.

Input

Output

Configure inputs to begin

Set options on the left, then click “Convert”.

How to convert a file to FASTA format?

Paste a sequence file or upload it, and the FASTA converter detects the format and writes a .fasta file with one record per sequence. GenBank, EMBL, FASTQ, Clustal, Stockholm, PHYLIP, NEXUS, MSF, PDB, mmCIF, CSV, and plain text all convert without choosing a format first, and nothing is uploaded to a server.

A GenBank record keeps its versioned accession and definition line as the FASTA header, while the feature table and position numbers are dropped:

Text
LOCUS       DEMO0001                  60 bp    DNA     linear   SYN 01-JAN-2024
DEFINITION  Synthetic construct with a short open reading frame.
ACCESSION   DEMO0001
VERSION     DEMO0001.1
FEATURES             Location/Qualifiers
     CDS             4..36
                     /gene="demoA"
ORIGIN
        1 ccgatggctt cgaaaggcga agagctgttc taataggtca ttgaccgcat gcatgcaaat
//
Text
>DEMO0001.1 Synthetic construct with a short open reading frame.
CCGATGGCTTCGAAAGGCGAAGAGCTGTTCTAATAGGTCATTGACCGCATGCATGCAAAT

Alignment formats become aligned FASTA when Preserve alignment gaps is on. This Clustal input, converted with gaps kept and no line wrapping:

Text
CLUSTAL W (1.83) multiple sequence alignment

KRAS_HUMAN      MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
HRAS_HUMAN      MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
RHEB_HUMAN      MPQSKSRKIAILGYRSVGKSSLTIQFVEGQFVDSYDPTIENTF
Text
>KRAS_HUMAN
MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
>HRAS_HUMAN
MTEYK---LVVVGAGGVGKSALTIQLIQNHFVDEYDPTIEDSY
>RHEB_HUMAN
MPQSKSRKIAILGYRSVGKSSLTIQFVEGQFVDSYDPTIENTF

With the default settings the six gap characters are removed instead, giving ungapped sequences, and a notice says how many were dropped.

Supported formats

The format is read from the content, not the file extension, so a GenBank record saved as .txt still converts as GenBank.

FormatRecognized byFASTA header comes from
GenBank (.gb, .gbk, .gbff)LOCUS and ORIGIN linesVERSION (or ACCESSION, then LOCUS) plus DEFINITION
EMBL (.embl)ID and SQ linesFirst AC accession plus DE description
FASTQ (.fastq, .fq)@ header, sequence, +, and a quality line of equal lengthThe read header; quality lines are dropped
Clustal (.aln)A first line starting with CLUSTAL, MUSCLE, or MAFFTSequence names; interleaved blocks are joined
Stockholm (.sto)# STOCKHOLM on the first lineSequence names; #= markup lines are skipped
PHYLIP (.phy)A first line with sequence and site countsSequence names; sequential and interleaved layouts
NEXUS (.nex)#NEXUS and a matrix blockTaxon names, including quoted names with spaces
MSF (.msf)An MSF: line and a // dividerSequence names
PDB (.pdb)ATOM or SEQRES recordsStructure ID and chain, one record per chain
mmCIF (.cif)A data_ block with _atom_site coordinatesEntry ID and chain from the first model
CSV or TSVA delimiter on nearly every row and a column of sequencesThe first column that isn't the sequence, or a column named id or name
FASTA or plain textAnything elseExisting > headers, or name lines found in the text

Plain text is read with the same rules as TXT to FASTA, which also explains how name lines, wrapped lines, and lists are handled. Word (.docx) and PDF uploads are converted to text in the browser first.

Settings

SettingDescription
Sequence typeAuto-detect (default), DNA, RNA, Nucleotide (DNA or RNA), or Protein. A specific type removes letters outside that alphabet and reports them. A U in DNA mode or a T in RNA mode skips that sequence with an explanation.
Sequence namesUse names from input (default) keeps the headers listed in the table above. Numbered (seq_1, seq_2, …) replaces them.
Name prefixShown for Numbered names. The text before the number. Default: seq.
Line wrapping80 characters per line (standard) (default), 60 characters per line, No wrapping (single line), or Custom characters per line.
Custom line lengthShown for custom wrapping. 0 turns wrapping off; an invalid value uses 80. Default: 80.

Under Advanced settings:

SettingDescription
Preserve alignment gapsKeeps - and . so converted alignments stay column-aligned. Default: off.
Preserve stop codonsKeeps a terminal *. Default: off.
Case formatUppercase (default), lowercase, or Preserve original. Stockholm insert columns are lowercase, so keep the original case when that distinction matters.
ValidationLenient (clean and report) (default), Strict (reject anything that needs cleanup), or Off.
Keep descriptionsShown for Numbered names. Writes the original header after the new name, for example >seq_1 DEMO0001.1 Synthetic construct with a short open reading frame. Default: off.

Results

The preview shows the FASTA text under a summary line that names the detected format, such as 1 sequence · 60 residues · DNA · read as GenBank. Notices above it list anything that changed the sequences, such as removed gaps or stop markers. Files holds the .fasta file, named after the uploaded file, and a run.log with the detected format, statistics, and cleanup counts.

What is FASTA format?

FASTA is a plain-text format for nucleotide and protein sequences: a header line starting with >, then the sequence as one-letter codes. It comes from the FASTA similarity-search program by Pearson and Lipman (1988), and its simplicity made it the default input for BLAST, aligners, structure predictors, and most sequence tools.

Richer formats hold more than FASTA can. GenBank and EMBL carry annotation and feature tables, FASTQ carries per-base quality scores, and Clustal, Stockholm, PHYLIP, and NEXUS encode alignments for specific programs. Converting to FASTA keeps the identifiers and sequences and drops the rest, which is usually what the next tool needs. When the annotation matters, extract it first: GenBank to FASTA writes coding sequences or protein translations, and the GenBank Feature Extractor pulls genes and other features.

Which converter should I use?

GoalTool
Convert any sequence file to FASTA without choosing a formatFASTA converter
Turn notes, pasted text, or a Word file into FASTATXT to FASTA
Extract CDS or translated proteins from GenBankGenBank to FASTA
Filter FASTQ reads by quality while convertingFASTQ to FASTA
Choose CSV columns and header prefixesCSV to FASTA
Choose chains or use SEQRES from a PDB filePDB to FASTA
Turn FASTA into a spreadsheet or plain textFASTA to CSV

Converted sequences can go straight into multiple sequence alignment with Clustal Omega or be divided with the FASTA splitter.

FAQ

How do I convert GenBank to FASTA?

Paste the GenBank record or upload the .gb file. Each record becomes a FASTA entry named by its versioned accession, followed by the definition line.

How do I convert a Clustal alignment to FASTA?

Paste the .aln file and turn on Preserve alignment gaps to keep the alignment. Interleaved blocks are joined per sequence, and the conservation line is ignored.

Can I convert FASTQ to FASTA online?

Yes. Headers are kept and the + and quality lines are dropped. For quality filtering during conversion, use FASTQ to FASTA.

Can I convert a PDB file to FASTA?

Yes. Each chain becomes a record read from its coordinates, or from SEQRES when the file has no coordinates. Chain selection and other options are in PDB to FASTA.

Is my sequence data uploaded?

No. Files are read and converted in your browser, including Word and PDF text extraction.

Table of contents

Related tools

CSV to FASTA

CSV to FASTA

Convert CSV or TSV sequence tables to FASTA with automatic delimiter detection and column mapping.

format-conversionprotein+4
TXT to FASTA converter

TXT to FASTA converter

Paste or upload DNA, RNA, or protein text and download FASTA. Sequence names, wrapped lines, and tables are detected automatically, and nothing leaves your browser.

format-conversionprotein+3
GenBank Feature Extractor

GenBank Feature Extractor

Extract CDS, mRNA, gene, rRNA, and tRNA features from GenBank records into FASTA.

sequence-manipulationDNA+4
FASTA to CSV converter

FASTA to CSV converter

Paste or upload FASTA and download a CSV or TSV table for spreadsheets, or a plain text file with one sequence per line.

format-conversionprotein+3
FASTA to FASTQ Converter

FASTA to FASTQ Converter

Convert FASTA sequences to FASTQ with configurable mock Phred quality scores.

format-conversionDNA+3
FASTQ to FASTA converter

FASTQ to FASTA converter

Convert FASTQ sequencing reads to FASTA while preserving headers and sequence case by default.

format-conversionDNA+3
GenBank to FASTA Converter

GenBank to FASTA Converter

Convert GenBank records to FASTA by extracting primary sequences, CDS, or translations.

format-conversionDNA+3
DNA to Protein Converter

DNA to Protein Converter

Translate DNA sequences into proteins across reading frames and genetic code tables.

format-conversionDNA+1
DNA to RNA converter

DNA to RNA converter

Transcribe DNA sequences to RNA by replacing thymine with uracil.

format-conversionDNA+1
Protein to DNA converter

Protein to DNA converter

Reverse translate protein sequences into DNA with codon optimization and GC-content controls.

format-conversionprotein+1