Use case
Maximum likelihood phylogenetics
Align homologous protein sequences, infer independent IQ-TREE and RAxML-NG trees, and compare models, topology, and branch support.
Inputs
1 required
Methods
3 connected
- 01MAFFT
- 02IQ-TREE
- 03RAxML-NG
Align homologous proteins with MAFFT, then infer independent maximum-likelihood trees with IQ-TREE and RAxML-NG from the same alignment.
Use this templateWhat is maximum likelihood phylogenetics?
Maximum likelihood phylogenetics is an approach that searches for the phylogenetic tree, branch lengths, and model parameters under which an observed sequence alignment has the highest likelihood. Because evaluating every possible topology is impractical for most datasets, programs use heuristic searches to find high-likelihood trees. The result is a best-scoring tree under the selected substitution model, accompanied by likelihood values and, when requested, branch-support estimates.
The alignment and evolutionary model define the evidence being evaluated. Protein analyses commonly compare amino-acid substitution models and may model unequal rates among sites. Partitioning can represent genes or domains separately, but unnecessary complexity can destabilize estimation. Model selection does not repair paralogs, poor taxon sampling, alignment errors, recombination, or nonhomologous columns.
Bootstrap and approximate likelihood support quantify repeatability under defined resampling or local tests; they are not probabilities that a biological claim is true. Compare independent searches, inspect model reports and warnings, review long branches and unstable taxa, and retain every tree rather than reporting only a polished figure.
When to use maximum likelihood phylogenetics
- Best fit. Estimating a best tree efficiently, comparing substitution models, and calculating bootstrap or likelihood-based branch support
- Required evidence. At least three homologous protein sequences, a reviewed MSA, defensible taxon sampling, and a suitable evolutionary model
Benefits of maximum likelihood phylogenetics
- Statistical inference. Provides an explicit statistical criterion for comparing trees under a defined evolutionary model.
- Evolutionary models. Scales to practical protein-family datasets with heuristic tree searches and efficient support calculations.
- Reusable evidence. Produces reusable Newick trees, model reports, likelihoods, and support values for downstream review.
Primary limitations
- Conditional result. The highest-likelihood topology is conditional on the alignment, taxon sampling, partitions, and substitution model.
- Data sensitivity. Heuristic searches do not guarantee that every analysis finds the global optimum.
- Interpretive limit. High support cannot rescue paralogy, recombination, contamination, alignment error, or biased sampling.
Maximum likelihood phylogenetics methods
IQ-TREE combines heuristic maximum-likelihood search with ModelFinder and several support procedures. RAxML-NG emphasizes scalable tree search and bootstrapping. FastTree provides approximate maximum-likelihood estimation for large alignments when speed is the priority. Their raw likelihoods are only directly comparable when the data, model, and likelihood conventions match.
A robust analysis normally uses multiple starts or independent programs, records the chosen model and partitions, and keeps support estimation distinct from best-tree search. Rooting is a separate interpretive step based on an outgroup, midpoint convention, or external evolutionary evidence.
Maximum likelihood phylogenetics applications
Maximum likelihood phylogenetics is used to study protein-family relationships, gene duplication and loss, orthology, pathogen or lineage history, ancestral hypotheses, and the evolutionary context of sequence or functional change. The taxon sample and locus determine which of those claims the tree can support.
Keep the inferred tree connected to sequence provenance, alignment decisions, models, support procedures, and independent biological evidence. A gene or protein tree is not automatically a species tree, and topological proximity does not independently establish direct ancestry or shared function.
How to run maximum likelihood phylogenetics online
Use the connected workflow to preserve the input sequences, reviewed alignment, method settings, native reports, trees, and warnings. Complete every external step explicitly rather than presenting a partial workflow as end-to-end inference.
- Define the evolutionary question. State the taxa, gene or protein region, rooting plan, and biological claim the tree should evaluate.
- Curate homologous sequences. Check identifiers, orthology, duplicates, fragments, domain boundaries, contamination, and taxon coverage.
- Build and review the MSA. Align with MAFFT, then inspect gap-rich, ambiguous, low-complexity, and poorly supported regions before inference.
- Run independent ML searches. Run IQ-TREE and RAxML-NG with documented models, search settings, random seeds, and branch-support procedures.
- Compare and export evidence. Compare likelihoods, topology, branch lengths, support, warnings, and outliers; export the alignment, reports, and Newick trees.
How to interpret maximum likelihood phylogenetics results
Read topology, branch length, and support separately. A clade describes an inferred grouping; branch length represents expected substitutions under the model; a support label summarizes a specified test. None independently supplies dates, direct ancestry, species identity, or functional equivalence.
Investigate discordance between IQ-TREE and RAxML-NG, unusually long branches, unstable placements, and support changes after alignment or taxon sensitivity checks. Report unresolved relationships as uncertainty and preserve the unrooted result when the rooting evidence is weak.
How maximum likelihood phylogenetics works
Align homologous proteins with MAFFT, then infer independent maximum-likelihood trees with IQ-TREE and RAxML-NG from the same alignment.
- Define the evolutionary question. State the taxa, gene or protein region, rooting plan, and biological claim the tree should evaluate.
- Curate homologous sequences. Check identifiers, orthology, duplicates, fragments, domain boundaries, contamination, and taxon coverage.
- Build and review the MSA. Align with MAFFT, then inspect gap-rich, ambiguous, low-complexity, and poorly supported regions before inference.
- Run independent ML searches. Run IQ-TREE and RAxML-NG with documented models, search settings, random seeds, and branch-support procedures.
- Compare and export evidence. Compare likelihoods, topology, branch lengths, support, warnings, and outliers; export the alignment, reports, and Newick trees.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Phylogenetic evidence.
FASTATSVCSVNEXUSHomologous protein FASTA records with stable identifiers, taxon metadata, justified boundaries, and enough taxonomic coverage for the question.
Outputs
- Trees and diagnostics.
NewickNEXUSCSVTSVTXTReviewed MSA, best-scoring Newick trees, model-selection reports, likelihoods, branch lengths, support values, logs, and provenance.
Tools for maximum likelihood phylogenetics
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

IQ-TREE
Infer maximum-likelihood trees with model selection and branch support

RAxML-NG
Run maximum-likelihood tree searches and bootstrap analysis

MAFFT
Create configurable protein, DNA, or RNA multiple-sequence alignments

FastTree
Estimate approximately maximum-likelihood trees for large alignments

MUSCLE5
Generate conventional or ensemble multiple-sequence alignments

Clustal Omega
Create scalable multiple-sequence alignments for homologous sequences

HMMER
Find homologs and inspect family membership with profile HMMs

MMseqs2
Search and cluster large protein or nucleotide sequence collections

StringZilla v5
Calculate pairwise similarity matrices and identify anomalous sequences

GenBank Feature Extractor
Extract annotated genes or proteins from GenBank records

CSV to FASTA
Convert sequence tables into consistently identified FASTA records

UniProt Download
Retrieve reviewed protein records and sequence metadata from UniProt
Other sequence analysis workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Frequently asked questions
No. It is the tree with the highest likelihood found under a specified alignment, model, taxon sample, and search procedure. Alternative data choices or model assumptions may support another topology, so sensitivity checks and independent evidence remain important.
Homologous protein FASTA records with stable identifiers, taxon metadata, justified boundaries, and enough taxonomic coverage for the question.
No. Support summarizes evidence under a particular alignment, model, sampling procedure, and taxon set. It does not correct paralogy, recombination, contamination, alignment error, biased taxon sampling, or model misspecification.
Only with a documented, reproducible rule and sensitivity analysis. Aggressive trimming can discard genuine signal, while retaining nonhomologous or highly uncertain columns can distort inference. Keep both the original and analyzed alignment.
Retain the exact MSA, sequence provenance, excluded columns and taxa, model and partitions, software versions, seeds, search replicates, support settings, logs, and every reported Newick tree.
A complete maximum likelihood phylogenetics project is normally quote-based because sequence retrieval, orthology review, alignment curation, model selection, compute, diagnostics, interpretation, and figures vary by dataset. Harvard’s FY26 bioinformatics core lists $180–$265 per consulting hour; MSU lists $84–$110 per hour and an eight-hour minimum for custom analysis.
Total cost depends on taxon and sequence count, alignment length, recombination or paralogy checks, partitioning, bootstrap or MCMC effort, replicate runs, convergence troubleshooting, sensitivity analysis, tree annotation, and whether the deliverable includes biological interpretation rather than files alone.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with supported workflow runs estimated in credits before submission. Done-for-you analysis is scoped separately. Bayesian MCMC currently requires an external engine, so external compute and specialist review may add to that project cost.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.