[Statistics]
AI drug discovery statistics in 2026 include more than 500 FDA submissions with AI components from 2016 to 2023, 75 AI-discovered molecules entered the clinic by 2023, 241 million AlphaFold DB predictions, and measurable virtual-screening and ADMET benchmarks.
Dr. Matic Broz Computational chemist
[Statistics]
AlphaFold DB v6 contains 241,070,489 predicted protein structures, the database has reached more than 3.4 million users across 190 countries, and AlphaFold 2 achieved 0.96 angstrom median backbone accuracy in CASP14.
Dr. Matic Broz Computational chemist
[Statistics]
About 86% to 93% of drug candidates that enter clinical trials fail before approval, depending on the dataset and counting method. Phase 2 is the largest hurdle, with about 71% failing to advance in the 2011-2020 BIO benchmark.
Dr. Matic Broz Computational chemist
[Statistics]
Drug development cost estimates range from 161 million to 4.54 billion US dollars. Recent transparent studies put the typical approved new drug near 700 million to 1.3 billion US dollars after failure and capital adjustments.
Dr. Matic Broz Computational chemist
[Statistics]
Drug discovery trends in 2026 include 46 FDA CDER novel approvals in 2025, 104 EMA human medicine recommendations, 218,766 interventional drug or biological studies in ClinicalTrials.gov, over 500 FDA submissions with AI components from 2016 to 2023, and strong growth in precision and modality-specific discovery.
Dr. Matic Broz Computational chemist
[Statistics]
Humans and chimpanzees share about 98.8% of directly comparable DNA sequence, about 96% when insertions and deletions are included in the classic 2005 genome comparison, and lower strict one-to-one fractions in complete-assembly analyses.
Dr. Matic Broz Computational chemist
[Statistics]
The human genome is about 3.1 billion base pairs per haploid copy, equal to about 3.1 million kilobases. File size ranges from about 797 MB in 2bit format to 938 MB for compressed hg38 FASTA and roughly 3 GB as plain sequence text.
Dr. Matic Broz Computational chemist
[Statistics]
About 8% of the human genome is viral DNA in the standard sense: inherited endogenous retrovirus sequences. That is about 250 million bases in a 3.1-billion-base reference genome.
Dr. Matic Broz Computational chemist
[Statistics]
The human proteome is about 19,400 proteins if counted as one reference protein per protein-coding gene, 172,117 distinct GENCODE translations, and likely millions of proteoforms depending on the definition.
Dr. Matic Broz Computational chemist
[Statistics]
Humans share about 17-25% of protein-coding genes with bananas, not 50-60% of DNA as widely claimed. Whole-genome alignment would be less than 1%. The shared genes are ancient housekeeping genes conserved for 1.5 billion years.
Dr. Matic Broz Computational chemist
[Statistics]
The genome of the fork fern Tmesipteris oblanceolata spans 160 billion base pairs, making it the largest known eukaryotic genome. The bacterial symbiont Nasuia deltocephalinicola has just 112,000 base pairs. The human genome sits in between at roughly 3.2 billion base pairs.
Dr. Matic Broz Computational chemist
[Statistics]
The typical protein is about 300–400 amino acids long, weighing 30–50 kDa. Human proteins average ~375 amino acids, bacterial proteins ~267, and the extremes span from ~20 amino acids to 45,212.
Dr. Matic Broz Computational chemist
[Statistics]
Depending on how function is defined, between 8% and 80% of the human genome may be functional. By the stricter evolutionary definition, about 8–15% of DNA is under selective constraint, leaving the remaining 85–92% without evidence of function.
Dr. Matic Broz Computational chemist
[Statistics]
Modern whole-genome sequencing can generate human genome data in hours to about a day. The first public human genome took 13 years, and consumer or clinical reports can still take days to weeks.
Dr. Matic Broz Computational chemist
[Statistics]
Arginine is the most basic standard amino acid by side-chain pKa, while aspartic acid is the most acidic. At pH 7.4, arginine is about 99.999% positively charged and aspartate is about 99.98% negatively charged.
Dr. Matic Broz Computational chemist
[Statistics]
The Protein Data Bank contains 256,006 released structures as of June 30, 2026. RCSB PDB also lists 1,062,058 computed structure models, while wwPDB recorded 4.72 billion PDB data downloads and views in 2025.
Dr. Matic Broz Computational chemist
[Statistics]
Up-to-date E. coli statistics covering the K-12 MG1655 reference genome, current gene and protein counts, cell size, growth rate, ribosomes, mutation rate, public genome assemblies, pathogenic groups, and U.S. STEC foodborne burden estimates.
Dr. Matic Broz Computational chemist
[Statistics]
Humans have about 19,000 to 20,000 protein-coding genes. ProteinIQ estimates the number at 19,442; the broader human reference annotation has 78,733 gene entries.
Dr. Matic Broz Computational chemist
[Statistics]
Leucine is the most common standard amino acid in UniProtKB/Swiss-Prot release 2026_02 at 9.65%. Glycine is the answer for collagen because collagen has glycine at every third residue.
Dr. Matic Broz Computational chemist
[Statistics]
The smallest protein depends on the definition: TAL/pri is an 11-amino-acid natural functional translated product, chignolin is a 10-amino-acid designed folding peptide, and insulin is a small 51-amino-acid mature human protein but not the smallest human entry.
Dr. Matic Broz Computational chemist
[Protein analysis]
Chai-2 is an AI model for de novo antibody and protein binder design. Learn how it works, its performance benchmarks, access options, and alternatives for computational binder design.
Dr. Matic Broz Computational chemist
[Statistics]
Human DNA is about 2.06 meters long in one diploid cell, 6.27-6.37 billion base pairs per cell, and roughly 6.2-14.4 billion kilometers across the nucleated cells in one body.
Dr. Matic Broz Computational chemist
[Statistics]
PKZILLA-1 is the largest known protein at 45,212 amino acids and 4.7 MDa. Titin remains the largest human protein, while dystrophin is much smaller but is encoded by an unusually large gene.
Dr. Matic Broz Computational chemist
[Statistics]
UniRef100 contains 475 million protein sequence clusters, GENCODE v50 annotates 19,442 human protein-coding genes and 172,117 distinct translations, the HUPO HPP reference proteome has 94% confident detection, and cell-level molecule counts range from 42 million in yeast to 10 trillion in a human cell.
Dr. Matic Broz Computational chemist