Simple Phylogenetics Workflow for DNA BarcodesOne great application of DNA barcodes is the ability to generate accurate and relevant phylogenetic trees for ecological and evolutionary analyses. There are lots of ways to do this, but not all of them may be necessary or relevant to your end goals. What do you need to know before you get started? This post provides a simple road map that you can follow to decide whether and how to construct a phylogeny using DNA barcode data for your research. Whether you’re trying to resolve relationships among a handful of closely related species or you’ve assembled a broad sampling across families, DNA-barcode data provide a fast, cost-effective starting point for phylogenetic inference. Below is a simple workflow—along with key tools and resources—to take you from data mining all the way through a time-calibrated tree. RESOURCES FOR PHYLOGENETICS WITH DNA BARCODES
4. CONSTRAIN TREE (Optional)
-Constraints can be specified to restrict the possible number of relationships among taxa that the phylogenetics software will explore -Constraining trees is generally a good idea for trees built from DNA barcodes, particularly if your taxon set includes taxa from divergent lineages (e.g. different families, orders, ect.) -Implementation of constraints depends on the program used to estimate the tree (read the manual) 5. CALIBRATE TREE (Optional) -There are two main ways of time calibrating a tree -Assign node ages using fossils and infer calibration during phylogenetic analysis -Rescale tree after phylogenetic analysis (Phylocom's Bladj) -Great tutorial on time calibration available from Tracy Heath (http://phyloworks.org/workshops/DivTime_BEAST2_tutorial_FBD.pdf) 6. ESTIMATE TREE -Several different “flavors” of analysis (each with their own assumptions) including Parsimony, Maximum Likelihood, and Bayesian tree estimation -For Parsimony implementation use TNT (http://www.zmuc.dk/public/phylogeny/tnt/) -For Maximum Likelihood use RAxML(https://cme.h-its.org/exelixis/software.html) -For Bayesian use MrBayes (http://nbisweden.github.io/MrBayes/), BEAST (http://beast.community), BEAST2 (http://www.beast2.org) or RevBayes (https://revbayes.github.io) -Generally, people resist Parsimony at this point and prefer Maximum Likelihood or Bayesian tree estimation -You always have the option of using multiple approaches COMPUTING POWER -While many of these programs will run on local machines just fine for small sets of taxa, if you are doing analyses with hundreds of species or just want to do things faster, use CIPRES for free phylogenetics supercomputing (http://www.phylo.org) -Constraints can be specified to restrict the possible number of relationships among taxa that the phylogenetics software will explore -Constraining trees is generally a good idea for trees built from DNA barcodes, particularly if your taxon set includes taxa from divergent lineages (e.g. different families, orders, ect.) -Implementation of constraints depends on the program used to estimate the tree (read the manual) 4. CALIBRATE TREE -There are two main ways of time calibrating a tree -Assign node ages using fossils and infer calibration during phylogenetic analysis -Rescale tree after phylogenetic analysis (Phylocom's Bladj) -Great tutorial on time calibration available from Tracy Heath (http://phyloworks.org/workshops/DivTime_BEAST2_tutorial_FBD.pdf) 5. ESTIMATE TREE -Several different “flavors” of analysis (each with their own assumptions) including Parsimony, Maximum Likelihood, and Bayesian tree estimation -For Parsimony implementation use TNT (http://www.zmuc.dk/public/phylogeny/tnt/) -For Maximum Likelihood use RAxML(https://cme.h-its.org/exelixis/software.html) -For Bayesian use MrBayes (http://nbisweden.github.io/MrBayes/), BEAST (http://beast.community), BEAST2 (http://www.beast2.org) or RevBayes (https://revbayes.github.io) Most researchers now favor ML or Bayesian approaches on barcode data, but you can always compare multiple methods for consistency. Bonus: Computing Resources If your dataset is small, most of these programs will run comfortably on a desktop or laptop. For larger datasets (hundreds of species) or to speed up analyses, take advantage of the free CIPRES Science Gateway: http://www.phylo.org; Brown University members should use the Oscar Supercomputer. Conclusion By following these six steps—research, align, model/partition, (optionally) constrain, (optionally) calibrate, and estimate—you’ll produce a robust phylogenetic hypothesis based on DNA barcodes. Each stage offers multiple software choices; pick the tools that best match your expertise and computing resources. Happy tree building!
0 Comments
Your comment will be posted after it is approved.
Leave a Reply. |
RSS Feed