Use case

Maximum likelihood phylogenetics

Align homologous protein sequences, infer independent IQ-TREE and RAxML-NG trees, and compare models, topology, and branch support.

Maximum likelihood phylogeneticsRead-only preview

Inputs

1 required

Methods

3 connected

  1. 01MAFFT
  2. 02IQ-TREE
  3. 03RAxML-NG

Align homologous proteins with MAFFT, then infer independent maximum-likelihood trees with IQ-TREE and RAxML-NG from the same alignment.

Use this template

What is maximum likelihood phylogenetics?

Maximum likelihood phylogenetics is an approach that searches for the phylogenetic tree, branch lengths, and model parameters under which an observed sequence alignment has the highest likelihood. Because evaluating every possible topology is impractical for most datasets, programs use heuristic searches to find high-likelihood trees. The result is a best-scoring tree under the selected substitution model, accompanied by likelihood values and, when requested, branch-support estimates.

The alignment and evolutionary model define the evidence being evaluated. Protein analyses commonly compare amino-acid substitution models and may model unequal rates among sites. Partitioning can represent genes or domains separately, but unnecessary complexity can destabilize estimation. Model selection does not repair paralogs, poor taxon sampling, alignment errors, recombination, or nonhomologous columns.

Bootstrap and approximate likelihood support quantify repeatability under defined resampling or local tests; they are not probabilities that a biological claim is true. Compare independent searches, inspect model reports and warnings, review long branches and unstable taxa, and retain every tree rather than reporting only a polished figure.

When to use maximum likelihood phylogenetics

  • Best fit. Estimating a best tree efficiently, comparing substitution models, and calculating bootstrap or likelihood-based branch support
  • Required evidence. At least three homologous protein sequences, a reviewed MSA, defensible taxon sampling, and a suitable evolutionary model

Benefits of maximum likelihood phylogenetics

  • Statistical inference. Provides an explicit statistical criterion for comparing trees under a defined evolutionary model.
  • Evolutionary models. Scales to practical protein-family datasets with heuristic tree searches and efficient support calculations.
  • Reusable evidence. Produces reusable Newick trees, model reports, likelihoods, and support values for downstream review.

Primary limitations

  • Conditional result. The highest-likelihood topology is conditional on the alignment, taxon sampling, partitions, and substitution model.
  • Data sensitivity. Heuristic searches do not guarantee that every analysis finds the global optimum.
  • Interpretive limit. High support cannot rescue paralogy, recombination, contamination, alignment error, or biased sampling.

Maximum likelihood phylogenetics methods

IQ-TREE combines heuristic maximum-likelihood search with ModelFinder and several support procedures. RAxML-NG emphasizes scalable tree search and bootstrapping. FastTree provides approximate maximum-likelihood estimation for large alignments when speed is the priority. Their raw likelihoods are only directly comparable when the data, model, and likelihood conventions match.

A robust analysis normally uses multiple starts or independent programs, records the chosen model and partitions, and keeps support estimation distinct from best-tree search. Rooting is a separate interpretive step based on an outgroup, midpoint convention, or external evolutionary evidence.

Maximum likelihood phylogenetics applications

Maximum likelihood phylogenetics is used to study protein-family relationships, gene duplication and loss, orthology, pathogen or lineage history, ancestral hypotheses, and the evolutionary context of sequence or functional change. The taxon sample and locus determine which of those claims the tree can support.

Keep the inferred tree connected to sequence provenance, alignment decisions, models, support procedures, and independent biological evidence. A gene or protein tree is not automatically a species tree, and topological proximity does not independently establish direct ancestry or shared function.

How to run maximum likelihood phylogenetics online

Use the connected workflow to preserve the input sequences, reviewed alignment, method settings, native reports, trees, and warnings. Complete every external step explicitly rather than presenting a partial workflow as end-to-end inference.

  1. Define the evolutionary question. State the taxa, gene or protein region, rooting plan, and biological claim the tree should evaluate.
  2. Curate homologous sequences. Check identifiers, orthology, duplicates, fragments, domain boundaries, contamination, and taxon coverage.
  3. Build and review the MSA. Align with MAFFT, then inspect gap-rich, ambiguous, low-complexity, and poorly supported regions before inference.
  4. Run independent ML searches. Run IQ-TREE and RAxML-NG with documented models, search settings, random seeds, and branch-support procedures.
  5. Compare and export evidence. Compare likelihoods, topology, branch lengths, support, warnings, and outliers; export the alignment, reports, and Newick trees.

How to interpret maximum likelihood phylogenetics results

Read topology, branch length, and support separately. A clade describes an inferred grouping; branch length represents expected substitutions under the model; a support label summarizes a specified test. None independently supplies dates, direct ancestry, species identity, or functional equivalence.

Investigate discordance between IQ-TREE and RAxML-NG, unusually long branches, unstable placements, and support changes after alignment or taxon sensitivity checks. Report unresolved relationships as uncertainty and preserve the unrooted result when the rooting evidence is weak.

How maximum likelihood phylogenetics works

Align homologous proteins with MAFFT, then infer independent maximum-likelihood trees with IQ-TREE and RAxML-NG from the same alignment.

  1. Define the evolutionary question. State the taxa, gene or protein region, rooting plan, and biological claim the tree should evaluate.
  2. Curate homologous sequences. Check identifiers, orthology, duplicates, fragments, domain boundaries, contamination, and taxon coverage.
  3. Build and review the MSA. Align with MAFFT, then inspect gap-rich, ambiguous, low-complexity, and poorly supported regions before inference.
  4. Run independent ML searches. Run IQ-TREE and RAxML-NG with documented models, search settings, random seeds, and branch-support procedures.
  5. Compare and export evidence. Compare likelihoods, topology, branch lengths, support, warnings, and outliers; export the alignment, reports, and Newick trees.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Phylogenetic evidence. FASTA TSV CSV NEXUS Homologous protein FASTA records with stable identifiers, taxon metadata, justified boundaries, and enough taxonomic coverage for the question.

Outputs

  • Trees and diagnostics. Newick NEXUS CSV TSV TXT Reviewed MSA, best-scoring Newick trees, model-selection reports, likelihoods, branch lengths, support values, logs, and provenance.

Other sequence analysis workflows

Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open workflow