
Convert CSV and TSV files containing sequence data to FASTA format with flexible column mapping and automatic delimiter detection

Reverse translate protein sequences to possible DNA sequences

Convert TXT or plain text sequences into FASTA format files for DNA, RNA, and protein workflows with cleanup, validation, and downloads

Extract sequence features (CDS, mRNA, gene, etc.) from GenBank files in FASTA format with support for spliced features

Convert DNA sequences to RNA (transcription) - replaces T with U

Convert FASTA sequence files to FASTQ format with mock quality scores

Convert standard FASTQ reads to FASTA with validation, IUPAC nucleotide support, average-quality filtering, and downloadable summaries

Convert GenBank files to FASTA format

Convert single-letter amino acid codes to three-letter codes

Convert Protein Data Bank files to Crystallographic Information File format

Convert CSV and TSV files containing sequence data to FASTA format with flexible column mapping and automatic delimiter detection

Reverse translate protein sequences to possible DNA sequences

Convert TXT or plain text sequences into FASTA format files for DNA, RNA, and protein workflows with cleanup, validation, and downloads

Extract sequence features (CDS, mRNA, gene, etc.) from GenBank files in FASTA format with support for spliced features

Convert DNA sequences to RNA (transcription) - replaces T with U

Convert FASTA sequence files to FASTQ format with mock quality scores

Convert standard FASTQ reads to FASTA with validation, IUPAC nucleotide support, average-quality filtering, and downloadable summaries

Convert GenBank files to FASTA format

Convert single-letter amino acid codes to three-letter codes

Convert Protein Data Bank files to Crystallographic Information File format
Configure inputs to begin
Set options on the left, then click “Convert”.
The conversion of DNA to protein is a fundamental process in all living organisms, essential for the structure, function, and regulation of the body's tissues and organs. This process, also known as gene expression, is the mechanism by which the genetic information stored in DNA is used to synthesize functional proteins.
The central dogma of molecular biology describes the flow of genetic information within a biological system. First proposed by Francis Crick in 1958, this principle states that genetic information flows from DNA to RNA to protein. This unidirectional flow is a cornerstone of molecular biology and is often summarized as "DNA makes RNA, and RNA makes protein".
There are three key processes involved in the central dogma:
While this is the general flow of genetic information, there are some exceptions. For instance, in retroviruses like HIV, reverse transcription can occur, where RNA is used as a template to synthesize DNA.
The conversion of the genetic instructions in DNA into a functional protein is a two-step process: transcription and translation.
Transcription is the process of creating an RNA copy of a gene's DNA sequence. This process is catalyzed by an enzyme called RNA polymerase. Transcription can be broken down into three main stages:
In eukaryotic cells, the initial mRNA transcript, called pre-mRNA, undergoes further processing. This includes splicing, where non-coding regions (introns) are removed, and the remaining coding regions (exons) are joined together. A protective cap and tail are also added to the ends of the mRNA molecule.
Translation is the process where the genetic information encoded in mRNA is used to synthesize a protein. This complex process occurs in the cytoplasm on ribosomes and involves another type of RNA molecule called transfer RNA (tRNA). Like transcription, translation has three main stages:
The genetic code is the set of rules by which information encoded in genetic material (DNA or RNA sequences) is translated into proteins. The code is read in groups of three nucleotides called codons. There are 64 possible codons, with 61 of them coding for the 20 different amino acids used to build proteins. The remaining three codons are stop codons.
A key feature of the genetic code is its degeneracy, meaning that some amino acids are specified by more than one codon. This redundancy can help to protect against mutations, as a change in a single nucleotide may not always result in a different amino acid.
The reading frame is also crucial. Since codons are read in threes, the sequence of amino acids is determined by where the reading of the mRNA begins. A shift in the reading frame can result in a completely different and often non-functional protein. The start codon AUG establishes the reading frame for protein synthesis.
The genetic code is typically represented as a codon table, which shows the correspondence between mRNA codons and their respective amino acids. To use a DNA to protein converter, a DNA sequence is first transcribed into its complementary mRNA sequence (with T replaced by U). Then, the mRNA sequence is read in triplets to determine the amino acid sequence based on the standard genetic code table.
Standard Genetic Code (RNA Codon Table)
| Codon | Amino Acid | Codon | Amino Acid | Codon | Amino Acid | Codon | Amino Acid |
|---|---|---|---|---|---|---|---|
| UUU | Phe | UCU | Ser | UAU | Tyr | UGU | Cys |
| UUC | Phe | UCC | Ser | UAC | Tyr | UGC | Cys |
| UUA | Leu | UCA | Ser | UAA | STOP | UGA | STOP |
| UUG | Leu | UCG | Ser | UAG | STOP | UGG | Trp |
| CUU | Leu | CCU | Pro | CAU | His | CGU | Arg |
| CUC | Leu | CCC | Pro | CAC | His | CGC | Arg |
| CUA | Leu | CCA | Pro | CAA | Gln | CGA | Arg |
| CUG | Leu | CCG | Pro | CAG | Gln | CGG | Arg |
| AUU | Ile |
| ACU |
| Thr |
| AAU |
| Asn |
| AGU |
| Ser |
| AUC | Ile | ACC | Thr | AAC | Asn | AGC | Ser |
| AUA | Ile | ACA | Thr | AAA | Lys | AGA | Arg |
| AUG | Met | ACG | Thr | AAG | Lys | AGG | Arg |
| GUU | Val | GCU | Ala | GAU | Asp | GGU | Gly |
| GUC | Val | GCC | Ala | GAC | Asp | GGC | Gly |
| GUA | Val | GCA | Ala | GAA | Glu | GGA | Gly |
| GUG | Val | GCG | Ala | GAG | Glu | GGG | Gly |