What is phylogenetic analysis?

Phylogenetic analysis is the study of evolutionary relationships among sequences, genes, proteins, organisms, or other biological entities using a phylogenetic tree. For molecular analysis, researchers curate homologous sequences, align corresponding positions, choose an evolutionary model, infer one or more trees, and evaluate uncertainty. The resulting topology, branch lengths, and support values are model-based hypotheses about shared history, not direct observations of ancestry.

Maximum-likelihood phylogenetics searches for a high-likelihood tree under a substitution model and commonly reports a best tree with bootstrap or likelihood-based support. Bayesian phylogenetics combines likelihoods and priors, then uses MCMC to approximate a posterior distribution of trees and parameters. These methods answer related questions but produce different uncertainty summaries.

Every inference depends on sequence selection, orthology, taxon coverage, alignment quality, partitions, model assumptions, and rooting. Recombination, horizontal transfer, gene duplication, contamination, long-branch attraction, and missing taxa can make a gene or protein tree differ from the underlying species history. Preserve these limits when annotating or comparing clades.

When to use phylogenetic analysis

  • Reconstruct evolutionary relationships. Compare homologous proteins, genes, taxa, variants, or lineages with an explicit tree model.
  • Test family and orthology hypotheses. Review duplications, lineage-specific groups, candidate orthologs, and anomalous records.
  • Add evolutionary context. Interpret sequence change, domains, functions, pathogens, or experiments alongside topology and support.

Benefits of phylogenetic analysis

  • Evolutionary context. Organizes homologous sequences into explicit, testable hypotheses about shared history.
  • Inspectible uncertainty. Associates topology with branch lengths, support, model reports, and sensitivity analyses.
  • Reusable outputs. Produces alignments and Newick trees for annotation, visualization, reconciliation, and downstream analysis.

Primary limitations

  • Conditional inference. The result depends on sampling, alignment, models, partitions, rooting, and search or sampling behavior.
  • Different biological histories. Recombination, horizontal transfer, duplication, and incomplete lineage sorting can separate gene and species trees.
  • Support is not proof. High branch support cannot correct biased or contaminated inputs or establish function and direct ancestry.

Types of phylogenetic analysis

These searched use cases separate best-tree likelihood inference from posterior sampling with explicit priors and MCMC diagnostics.

Maximum likelihood phylogenetics

Searches for the tree and model parameters that make an observed sequence alignment most probable.

Best for: Efficient best-tree estimation, model testing, bootstrapping, and large alignments
Requires: A reviewed homologous sequence alignment and an appropriate substitution model

Bayesian phylogenetics

Uses priors, a likelihood model, and MCMC sampling to estimate a posterior distribution of trees and parameters.

Best for: Posterior tree uncertainty, model-rich inference, divergence analysis, and explicit prior assumptions
Requires: A reviewed alignment, justified priors and models, independent MCMC runs, and convergence diagnostics

Maximum-likelihood and Bayesian phylogenetics

Choose maximum likelihood when the primary output is an efficiently estimated best tree with model testing and branch support. Choose Bayesian inference when the scientific question requires posterior distributions, explicit priors, clock or demographic models, or direct propagation of tree and parameter uncertainty.

The approaches are complementary rather than interchangeable. An independent ML tree can reveal topology conflicts in a Bayesian project, while posterior sampling can expose uncertainty hidden by one best ML tree. Support values must retain their method labels because bootstrap proportions and posterior probabilities have different meanings.

How to run phylogenetic analysis online

A defensible tree begins with the biological sampling design and a reviewed alignment. Algorithm choice cannot correct a dataset that mixes paralogs, nonhomologous regions, contaminated records, or unsuitable taxa.

  1. Define the question. Specify taxa, locus or protein region, expected evolutionary process, rooting evidence, and intended claim.
  2. Curate sequences. Review provenance, orthology, duplicates, fragments, boundaries, contamination, and taxon coverage.
  3. Build the alignment. Align homologs and inspect uncertain, gap-rich, repetitive, low-complexity, and nonhomologous regions.
  4. Choose and run inference. Use maximum likelihood or an external Bayesian engine with documented models, settings, seeds, and replicates.
  5. Validate and report. Compare topologies, support, convergence where relevant, outliers, sensitivity analyses, and independent evidence.

Phylogenetic analysis applications

Protein phylogenetics supports family classification, orthology and paralogy review, evolutionary annotation, domain-history analysis, pathogen and lineage comparison, ancestral hypotheses, target selection, and experimental design. Different loci may have different histories, so a protein tree should not be labeled a species tree without a justified reconciliation model.

Trees can also reveal suspicious sequences, unexpected clustering, long branches, or taxonomic gaps that require input correction. These quality-control signals are useful even when the final biological relationship remains unresolved.

How to interpret a phylogenetic tree

Topology represents nested relationships; branch lengths represent change under the fitted model; branch support describes method-specific repeatability or posterior evidence. The left-to-right order of tips can rotate around nodes without changing the topology, and two neighboring tips are not necessarily direct ancestors.

Rooting determines the direction of interpretation and requires outgroup or other evolutionary evidence. Treat short internal branches, method disagreement, low support, poor MCMC diagnostics, unstable taxa, and sensitivity to alignment or taxon selection as uncertainty that belongs in the conclusion.

How phylogenetic analysis works

The hub workflow runs IQ-TREE, RAxML-NG, and FastTree from one reviewed protein alignment so their native trees and support reports can be compared.

  1. Define the question. Specify the homologous region, taxon sample, rooting evidence, and intended evolutionary claim.
  2. Review the alignment. Check orthology, identifiers, fragments, boundaries, gaps, uncertain regions, and taxon coverage.
  3. Run three ML methods. Submit one aligned FASTA to IQ-TREE, RAxML-NG, and FastTree.
  4. Compare results. Review topology, branch lengths, support, likelihood context, long branches, and unstable placements.
  5. Export evidence. Choose a justified result while preserving alternative trees and the exact alignment.

Inputs and outputs

Check formats before running, then inspect and download the result from every workflow step.

Inputs

  • Sequence and sampling evidence. FASTA CSV TSV NEXUS Homologous protein sequences or a reviewed MSA with stable identifiers, taxon metadata, and documented boundaries and exclusions.

Outputs

  • Phylogenetic evidence. Newick NEXUS CSV TSV TXT Method-native trees, branch lengths, support values, likelihood and model reports, logs, alignments, and provenance.

Featured phylogenetic analysis workflow

Preserve three method-native trees to compare topology, branch lengths, support, likelihood reports, and warnings.

Phylogenetic analysis method panelRead-only preview

Inputs

1 required

Methods

3 connected

  1. 01IQ-TREE
  2. 02RAxML-NG
  3. 03FastTree

Preserve three method-native trees to compare topology, branch lengths, support, likelihood reports, and warnings.

Use this template

Frequently asked questions

Start with a workflow you can inspect and edit

Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.

Open method panel