Showing posts with label description. Show all posts
Showing posts with label description. Show all posts

Wednesday, November 30, 2016

Comparative methods in the genomic era

Every modern biologist should know about phylogenetic comparative methods, even if they are not familiar with the term. The idea is that we cannot compare biological traits without taking into account the evolutionary relations between the species/populations used as "sample points". This is because individuals are not independent samples of whatever variables we are trying to correlate (they are "pseudoreplicates").

However, not everybody is an up-to-date biologist, as we see in the figure below: birds with similar body masses would also have similar flight speeds not because these two traits are related, but because these birds are closely related. That is, not long ago they were a single species (with a single value for each trait). The solution to avoid these spurious correlations is, in abstract terms, to think about what trait values do we expect for their ancestors. Hopefully we will have this in mind in upcoming large-scale genomic analyses, where sophisticated statistical algorithms for ever-increasing amounts of data might obfuscate their biological appropriateness.

from doi:10.1063/1.4886855
Here I include a few links to good comments and papers discussing comparative methods, starting with Felsenstein's reaction to the paper above (which is about airplanes, by the way). I also include at the bottom the slides of an impromptu journal club I presented a few years ago about things we can model over a tree, with strong emphasis on phylogenetic contrasts.

Physicists and engineers decide how to analyze evolution (by Joe Felsenstein):
They make allometric plots of features of new airplane models, log-log plots over many orders of magnitude. The airplanes show allometry: did you know that a 20-foot-long airplane won’t have 100-foot-long wings? That you need more fuel to carry a bigger load?
But permit me a curmudgeonly point: This paper would have been rejected in any evolutionary biology journal. Most of its central citations to biological allometry are to 1980s papers on allometry that failed to take the the phylogeny of the organisms into account. The points plotted in those old papers are thus not independently sampled, a requirement of the statistics used. (More precisely, their error residuals are correlated).
Steps towards understanding comparative methods – Wainwright Lab:
Felsenstein elaborates on his method of calculating standardized contrasts (phylogenetically independent contrasts) to help overcome the non-independence of character traits.  These contrasts are basically the differences between trait values of species pairs weighted by the evolutionary change separating them; they are estimates of the rate of change over time. A common use of standardized contrasts is to look for correlation in this rate between two traits; if standardized contrasts of traits X and Y are compared in a regression analysis, a linear trend suggests correlated rates of evolution between the traits.
Strong phylogenetic inertia on genome size and transposable element content among 26 species of flies | Biology Letters doi:10.1098/rsbl.2016.0407:
To date, quantifying the importance of phylogenetic inertia in TE content distribution remains a key question as the dynamic of TE accumulation is still poorly understood. Here, we analysed the evolution of genome size and genomic TE content in 26 Drosophila, using a phylogenetic framework. We estimated genomic TE content using a de novo TE assembly approach, tested the correlation between TE content and genome size among closely related species and finally estimated the phylogenetic inertia.
(...)
Comparative analyses were performed using ape [21], nlme [22] and phytools [23] packages in R. Ancestral trait reconstruction of genome size was calculated using phylogenetic independent contrasts. We tested the phylogenetic signal using Pagel's λ [24]. Best-fitting model to the trait evolution and its covariance structure was tested among (i) absence of phylogenetic signal, (ii) neutral Brownian motion and (iii) constrained evolution Ornstein–Uhlenbeck (OU) models using generalized least squares (GLS) and selected according to minimum Akaike information criterion (AIC).
Controlling for non-independence in comparative analysis of patterns across populations within species | Philosophical Transactions of the Royal Society B: Biological Sciences doi:10.1098/rstb.2010.0311:
How do we quantify patterns (such as responses to local selection) sampled across multiple populations within a single species? Key to this question is the extent to which populations within species represent statistically independent data points in our analysis. Comparative analyses across species and higher taxa have long recognized the need to control for the non-independence of species data that arises through patterns of shared common ancestry among them (phylogenetic non-independence), as have quantitative genetic studies of individuals linked by a pedigree.
The Unsolved Challenge to Phylogenetic Correlation Tests for Categorical Characters | Systematic Biology doi:10.1093/sysbio/syu070:
When comparative biologists observe that animal species living in caves also tend to have reduced eyes, they may see such correlation as evidence that the traits are adaptively or functionally linked: for instance, selection to maintain eye function is relaxed when light is unavailable.
(...)
However, the last few decades have taught us that among-species correlative tests should take into account evolutionary relationships (Felsenstein 1985; Ridley 1989; Harvey and Pagel 1991). If phylogeny is not taken into account, an interpreted correlation may have a trivial explanation different from the biological relationship we claim. There is a correlation among species in the distribution of fur and bones in the middle ear—species with fur also have three bones in the middle ear, and vice versa. These two traits are characteristics of mammals, and absent outside the mammals. Using their shared distribution as evidence of an interesting biological relationship between fur and middle ear bones would be considered a mistake, however, for reasons understood long ago by Darwin (1872)

Saturday, December 26, 2009

New article describing priors on biomc2



Our new paper is out!
Leonardo de Oliveira Martins and Hirohisa Kishino (2010) Distribution of distances between topologies and its effect on detection of phylogenetic recombination. Annals of the Institute of Statistical Mathematics 62(1): 145--159. doi:10.1007/s10463-009-0259-8
Unfortunately I had to transfer the copyright to the ISM, but I am still allowed to maintain the accepted manuscript at my personal homepage [1]. In this article we describe in more details theoretical aspects that we couldn't explore before (because it would divert the audience's attention). The novelties are:
  • describing the topology distance parameter as an augmented variable. It seems contradictory that our method assumes that the number of recombination break-points is variable, and at the same time we claim that the number of parameters is constant. This is because we have a variable (we could think of it as an indicator variable) that tells if there is recombination or not. Actually it is not just an indicator variable since it doesn't simply say if there is recombination or not, but actually approximates the amount of recombination. I like the analogy with linear models (the infamous Y=B0 + B1 X1 + B2 X2 + ... + Bn Xn), where despite the number of variables is constant (n+1) those parameters too close to zero mean that the "effective" number of variables is smaller.
  • describing the mini-sampler procedure. There is a strategy in reversible-jump MCMC (rjMCMC) that allows the chain to walk a little bit before accepting/rejecting, that we refer to as the mini-sampler. As we just saw our model doesn't need a rjMCMC since the number of parameters is constant, but still we employ the same strategies, making sure that the moves respect the detailed balance of the chain. This mini-sampler is necessary because on the one hand each step changes the topology just a little, and on the other hand we want to be able to handle cases where neighboring topologies are very different.
  • showing the importance of the modified Poisson distribution as a prior for the distances. We explicitly work with two scenarios:
    1. forcing the penalty hyperparameter to a fixed value. This shows the relevance of the hierarchical modelling.
    2. using a simplified distance that can only detect presence/absence of recombination witout being able to quantify it. This is a model analogous to other procedures that don't explicitly take the distance into account, and is equivalent to the "indicator variable" described above.

  • description of the ensemble of mosaics and how to choose the most representative one. Each MCMC sample is one set of topologies (one per segment) that we call the mosaic structure. We then devise the calculation of a distance between these mosaics to quantify how similar two samples are, and find the centroid mosaic - the sample most similar to all other samples. It is worth noticing that
    1. this is done after the MCMC sampling is finished, and is not part of the Bayesian model per se
    2. this distance between samples is completely unrelated to the distance between topologies (that we call dSPR).

In the meanwhile I'm updating the software site and upgrading the program: nothing special, but I realized that many libraries were unnecessary and that biomc2.summarise was painfully slow for large topologies. Now it is only slow.

[1] the first time that you try to access any file hosted on https://corn.ab.a.u-tokyo.ac.jp/ your browser will complain about my self-signed certificate - some lobby is not happy about it. Please neglect the terrorist warnings.


Wednesday, July 9, 2008

Introduction - welcome message

Welcome to the biomc2 blog. I plan to use it as a discussion place for the software biomc2, as described in the PLoS paper doi:10.1371/journal.pone.0002651, but it is OK to discuss related issues, and further developments on the software.

It is intended as a flexible alternative to mail communication, but since this is a discussion place anything can change. I believe a blog would be less intrusive than mailing list, and it is somehow easier to keep track of the topics. You can follow new messages through the RSS feed (you will need a feed reader like google reader, mozilla plugin, etc. - I use the KDE akregator for linux).

The other homepage (http://corn.ab.a.u-tokyo.ac.jp/~leo/biomc2) will be the static version, with download area. (Update 2009 06 16: this homepage was superseded by http://www.biomcmc.org which is now the "aggregator" of information about the software)

Sincerely,
Leo