Use case
Bayesian phylogenetics
Prepare a reviewed protein alignment, plan priors and independent MCMC runs, and create an IQ-TREE cross-check before external Bayesian inference.
Inputs
1 required
Methods
2 connected
- 01MUSCLE5
- 02IQ-TREE · Independent ML Cross-check
Create a protein MSA with MUSCLE5 and an independent IQ-TREE maximum-likelihood cross-check before configuring Bayesian MCMC externally.
Use this templateWhat is bayesian phylogenetics?
Bayesian phylogenetics is an approach that combines a likelihood model, prior distributions, and sequence data to estimate a posterior distribution over phylogenetic trees and model parameters. Markov chain Monte Carlo sampling approximates that posterior by visiting states in proportion to their posterior probability. The result is a sample of trees and parameters that must be checked for convergence and summarized, not a single definitive tree.
Bayesian analysis makes assumptions explicit through priors on topology, branch lengths, substitution parameters, clocks, or demographic processes. Those priors can be informative or weakly informative, but they still affect the posterior when the data provide limited information. Model, partition, calibration, and prior sensitivity should therefore be planned before interpreting clade probabilities or dates.
Independent chains must mix adequately and converge on the same posterior region. Review traces, effective sample sizes, split frequencies or comparable diagnostics, burn-in, replicate agreement, and sensitivity runs. ProteinIQ prepares the alignment and an independent maximum-likelihood cross-check; Bayesian MCMC inference itself currently runs in an external program such as MrBayes, BEAST, or RevBayes.
When to use bayesian phylogenetics
- Best fit. Estimating posterior tree uncertainty, testing model-rich evolutionary hypotheses, and incorporating justified prior information
- Required evidence. A reviewed homologous alignment, documented models and priors, an external Bayesian engine, independent MCMC runs, and convergence diagnostics
Benefits of bayesian phylogenetics
- Statistical inference. Represents uncertainty as a posterior distribution over trees and parameters rather than only one best topology.
- Evolutionary models. Supports rich evolutionary models and the explicit inclusion of scientifically justified prior information.
- Reusable evidence. Provides posterior samples that can propagate uncertainty into clade, parameter, or divergence-time summaries.
Primary limitations
- Conditional result. Poor mixing or apparent convergence to different posterior regions can invalidate summaries despite a completed run.
- Data sensitivity. Posterior results can be sensitive to priors, calibrations, partitions, alignment choices, and model misspecification.
- Interpretive limit. MCMC may require substantial external compute and specialist diagnosis; ProteinIQ does not currently run the Bayesian engine.
Bayesian phylogenetics methods
MrBayes commonly supports Bayesian tree and model estimation, while BEAST focuses strongly on time-scaled phylogenies and phylodynamics; RevBayes provides a flexible graphical-model framework. Program choice follows the scientific model, not the desired appearance of the tree. ProteinIQ does not currently execute these Bayesian engines.
The ProteinIQ workflow prepares a traceable MSA with MUSCLE5 and runs IQ-TREE as an independent maximum-likelihood cross-check. Export the reviewed alignment and metadata to the chosen Bayesian program, then return posterior tree samples, consensus trees, logs, and diagnostics for integrated interpretation.
Bayesian phylogenetics applications
Bayesian phylogenetics is used to study protein-family relationships, gene duplication and loss, orthology, pathogen or lineage history, ancestral hypotheses, and the evolutionary context of sequence or functional change. The taxon sample and locus determine which of those claims the tree can support.
Keep the inferred tree connected to sequence provenance, alignment decisions, models, support procedures, and independent biological evidence. A gene or protein tree is not automatically a species tree, and topological proximity does not independently establish direct ancestry or shared function.
How to run bayesian phylogenetics online
Use the connected workflow to preserve the input sequences, reviewed alignment, method settings, native reports, trees, and warnings. Complete every external step explicitly rather than presenting a partial workflow as end-to-end inference.
- Define the Bayesian model. Specify the evolutionary question, partitions, likelihood model, priors, clock or tree model, calibrations, and planned sensitivity analyses.
- Curate homologous sequences. Check orthology, sequence provenance, taxon sampling, duplicates, fragments, contamination, and domain boundaries.
- Prepare and review the MSA. Run MUSCLE5, inspect uncertain columns, and generate an IQ-TREE result as an independent topology and branch-support cross-check.
- Run external MCMC. Export the reviewed alignment to MrBayes, BEAST, RevBayes, or another external engine and run multiple independently seeded chains.
- Diagnose and summarize. Check mixing, effective sample sizes, burn-in, replicate agreement, prior sensitivity, and consensus-tree stability before reporting.
How to interpret bayesian phylogenetics results
A posterior clade probability is conditional on the data, likelihood model, priors, and adequate MCMC sampling. It is not interchangeable with bootstrap support. A narrow posterior can still be misleading when the model is misspecified, calibrations are inappropriate, or different chains have not reached the same distribution.
Inspect trace stationarity, effective sample size, replicate agreement, topology frequencies, branch-length distributions, and sensitivity to priors or excluded columns. Report unresolved relationships, multimodality, weak calibration support, and conflicts with the ML cross-check instead of hiding them in a consensus tree.
How bayesian phylogenetics works
Create a protein MSA with MUSCLE5 and an independent IQ-TREE maximum-likelihood cross-check before configuring Bayesian MCMC externally.
- Define the Bayesian model. Specify the evolutionary question, partitions, likelihood model, priors, clock or tree model, calibrations, and planned sensitivity analyses.
- Curate homologous sequences. Check orthology, sequence provenance, taxon sampling, duplicates, fragments, contamination, and domain boundaries.
- Prepare and review the MSA. Run MUSCLE5, inspect uncertain columns, and generate an IQ-TREE result as an independent topology and branch-support cross-check.
- Run external MCMC. Export the reviewed alignment to MrBayes, BEAST, RevBayes, or another external engine and run multiple independently seeded chains.
- Diagnose and summarize. Check mixing, effective sample sizes, burn-in, replicate agreement, prior sensitivity, and consensus-tree stability before reporting.
Inputs and outputs
Check formats before running, then inspect and download the result from every workflow step.
Inputs
- Phylogenetic evidence.
FASTATSVCSVNEXUSHomologous protein FASTA records plus taxon metadata, partition definitions, prior rationale, calibration evidence where relevant, and an external Bayesian-engine configuration.
Outputs
- Trees and diagnostics.
NewickNEXUSCSVTSVTXTProteinIQ returns the reviewed MSA and ML cross-check; the external engine returns posterior tree samples, parameter traces, diagnostics, and consensus summaries.
Tools for bayesian phylogenetics
Use these methods to prepare inputs, run the core analysis, inspect outputs, and validate the evidence described in this workflow.

MUSCLE5
Generate conventional or ensemble multiple-sequence alignments

MAFFT
Create configurable protein, DNA, or RNA multiple-sequence alignments

IQ-TREE
Infer maximum-likelihood trees with model selection and branch support

Clustal Omega
Create scalable multiple-sequence alignments for homologous sequences

RAxML-NG
Run maximum-likelihood tree searches and bootstrap analysis

FastTree
Estimate approximately maximum-likelihood trees for large alignments

HMMER
Find homologs and inspect family membership with profile HMMs

MMseqs2
Search and cluster large protein or nucleotide sequence collections

UniProt Download
Retrieve reviewed protein records and sequence metadata from UniProt

GenBank to FASTA Converter
Convert GenBank records into traceable FASTA inputs

GenBank Feature Extractor
Extract annotated genes or proteins from GenBank records

CSV to FASTA
Convert sequence tables into consistently identified FASTA records
Other sequence analysis workflows
Compare related approaches based on the molecular system, available evidence, required inputs, and decision you need to support.
Frequently asked questions
Not currently. ProteinIQ can prepare and review the alignment and run an independent IQ-TREE maximum-likelihood cross-check. Run Bayesian MCMC in an external engine such as MrBayes, BEAST, or RevBayes, then retain its posterior samples, traces, consensus trees, and convergence diagnostics.
Homologous protein FASTA records plus taxon metadata, partition definitions, prior rationale, calibration evidence where relevant, and an external Bayesian-engine configuration.
No. Support summarizes evidence under a particular alignment, model, sampling procedure, and taxon set. It does not correct paralogy, recombination, contamination, alignment error, biased taxon sampling, or model misspecification.
Only with a documented, reproducible rule and sensitivity analysis. Aggressive trimming can discard genuine signal, while retaining nonhomologous or highly uncertain columns can distort inference. Keep both the original and analyzed alignment.
Retain the reviewed MSA, priors and their rationale, partitions, calibrations, software versions, seeds, chain lengths, sampling interval, burn-in rule, independent-run diagnostics, effective sample sizes, posterior tree samples, and sensitivity analyses.
A complete bayesian phylogenetics project is normally quote-based because sequence retrieval, orthology review, alignment curation, model selection, compute, diagnostics, interpretation, and figures vary by dataset. Harvard’s FY26 bioinformatics core lists $180–$265 per consulting hour; MSU lists $84–$110 per hour and an eight-hour minimum for custom analysis.
Total cost depends on taxon and sequence count, alignment length, recombination or paralogy checks, partitioning, bootstrap or MCMC effort, replicate runs, convergence troubleshooting, sensitivity analysis, tree annotation, and whether the deliverable includes biological interpretation rather than files alone.
ProteinIQ self-service starts at $29 per month for academic Plus and $99 per month for commercial Pro, with supported workflow runs estimated in credits before submission. Done-for-you analysis is scoped separately. Bayesian MCMC currently requires an external engine, so external compute and specialist review may add to that project cost.
Start with a workflow you can inspect and edit
Add your inputs, review the settings, and keep every structure, score, table, and file connected to the step that produced it.