My Blog List

Showing posts with label Admixture. Show all posts
Showing posts with label Admixture. Show all posts

Sunday, November 11, 2012

The levels and dating of admixture in Belorusians


Following Dienekes's suggestion on using Pathan and Lithuanian samples as references for ROLLOFF analysis, i decided to undertake a second attempt of formal analysis of admixture and dating of admixture events in Belorusian samples which are available to me: the reference dataset of Belorusians from Behar et al.2011., and Belorusian samples collected by our project.


Below  you can glean the results of experiment which i deem less noisy in contrast to my previous attempt. 

valid snps: 746877
group 0 Lithuanian
group 1 Pathan
number admixed: 13 number of references: 2
numsnps: 746877  numindivs: 55
starting main loop. numsnps: 158101


Summary of fit:

Formula: wcorr ~ (C + A * exp(-m * dist/100))

Parameters:
   Estimate Std. Error t value Pr(>|t|)  
C 2.332e-04  3.029e-04   0.770  0.44165  
A 3.306e-02  1.227e-02   2.695  0.00728 **
m 1.169e+02  3.851e+01   3.037  0.00252 **
---
Signif. codes:  0 ‘***’ 0.001 ‘**’ 0.01 ‘*’ 0.05 ‘.’ 0.1 ‘ ’ 1

Residual standard error: 0.006508 on 493 degrees of freedom

Number of iterations to convergence: 0
Achieved convergence tolerance: 9.103e-06

mean (generations):  116.9416
jackknife (generations)   105.086+-52.591 
The date of admixture event in Belarusian_V sample with Belarusian and Pathan being reference populations appears to be very close to the date which was estimated by Dienekes for Lithuanian [Lithuanian_D;Pathan].


Inference of Admixture Parameters in Belarusians using Weighted Linkage Disequilibrium

On 1 November 2012 Po-Ru Loh, Mark Lipson, Nick Patterson, Priya Moorjani, Joseph K Pickrell, David Reich, Bonnie Berger announced and published their new paper, in which they introduced a new approach that harnesses the exponential decay of admixture-induced linkage disequilibrium (LD) as a function of genetic distance. They proposed a new weighted LD statistic that can be used to infer mixture proportions as well as dates with fewer constraints on reference populations than previous methods.


I haven't had enough time to investigate this method in full extent, but i used a software package ALDer which implements the weighted LD statistic for a quick & dirty experiment of dating the admixture events in Belarusian sample:

 

Sunday, October 14, 2012

ROLLOFF analysis of Poles, Belarusians and Russians from central regions of Russia

A month ago  the notorious Reich lab released an alpha version of ADMIXTOOOLS version 1.0. The alpha version package was developed for in-house use, so the operating routine is not always self-explanatory. The goof thing, however, is that ADMIXTOOLS package maintains full format compatibility with another very well-known EIGENSOFT software program developed by the same lab. This makes the learning curve of ADMIXTOOLS much steeper and flatter. 


The aforementioned package features 6 cool programs, among which i find most useful qp3Pop and rolloff. Due to limitations of this post, i am not going to discuss qp3pop in all details and for the purpose of my presentation it is suffice to say that this program implements three- population (f_3) test for treeness of populations from Reich et al. 2009
Rather i'd suggest reading ADMIXTOOLS supporting material, Dienekes' posts and Reich's paper to get an idea of what f_3 test is about.

ROLOFF method, however, needs a closer look.This is a method that measures time since admixture. It does so by looking at the linkage disequilibrium between SNPs due to admixture.  Now it is time to recall the standard definition of the linkage disequilibrium. The linkage disequilibrium (further -LD) is the nonrandom association between two alleles such that certain combinations are more likely to occur together. As two SNPs get farther apart, we expect there to be less admixture LD.The rate of decline of admixture LD is directly related to the number of generations since admixture, since that indicates how many recombinations have occurred between any two SNPs. In short: Rolloff fits an exponential curve to a plot of admixture LD vs. distance, and uses the rate of exponential decline to calculate the generations since admixture. Given that one generation is roughly equal to 29 years, one can convert the number of generations since admixture into years.

Dienekes has already tested ADMIXTOOLS programs with various worldwide populations, and among other he has carried out f_3 and rolloff analyses of Poles, Lithuanians, and Ukrainians.

Below are words from the man himself:

Using the aforementioned idea, I set out to see whether Lithuanians, who occupy the European end of the Europe-South Asia cline present such a signal of admixture LD. I used the Lithuanian_D sample from the Dodecad Project and the Balochi HGDP sample as reference populations (to calculate allele frequency differences), and the Behar et al. (2010) Lithuanians for admixture LD. There were only ~300k SNPs usuable in this set, but sufficient to detect the signal of admixture LD:
The admixture time estimate is 200.350 +/- 61.608 generations, or 5,810 +/- 1790 years. This is not very precise, probably because of the small number of SNPs and individuals used, but it certainly points to the Neolithic-to-Bronze Age for the occurrence of this admixture. The date is certainly reminiscent of the expansion of the Kurgan culture out of eastern Europe, or, the later Corded Ware culture of northern Europe.

So, it may well appear that at least some of the people participating in these groups of cultures, were indeed influenced by the Indo-Europeans as they expanded from their West Asian homeland. These intruders mixed with eastern Europeans who vacillated during the late Neolithic between a northern Europeoid pole akin to Mesolithic hunter gatherers from
Gotland and Iberia, and a widely dispersed Sardinian-like population that is in evidence at least in the Sweden-Italian Alps-Bulgaria triangle. The gradual appearance of non-mtDNA U related lineages in Siberia and Ukraine is most likely related to this phenomenon.

......
I have carried out rolloff analysis of my 25-strong Polish_D sample using Lithuanians and Pathans as references:

The signal is fairly distinct, and corresponds to 149.296 +/- 38.783 generations or 4330 +/- 1120 years. I am guessing that either the different reference population (Pathans vs. Balochi), or, more likely the increased number of target individuals (25 vs. 10) have contributed to the narrowing down of the uncertainty. It will be interesting to explore this signal further with more population pairs.

...

I have used the Yunusbayev et al. sample of Ukrainians, and estimated its admixture time using Lithuanians and Balochi as reference populations: The admixture time estimate is 191.078 +/- 35.079 generations, or 5,540 +/- 1,020 years. It seems very similar to that in Lithuanians, with a smaller standard error, perhaps on account of either the larger number of SNPs or larger number of individuals.

It is tempting to associate this admixture signal with the
Maikop culture which appeared at around this time. Assuming that North_European/West_Asian (or Lithuanian-like and Balochi-like) gene pools existed north and south of the Pontic-Caspian-Caucasus set of geographical barriers, then the Maikop culture which shows links to both the early Transcaucasian culture and those of Eastern Europe would have been an ideal candidate region for the admixture picked up by rolloff to have taken place. There are, of course, other possibilities.


As always, Dienekes' analysis spawned a lot of criticism on behalf of another genome blogger - Davidski from Eurogenes. In his latest post he argued  that it’s difficult to say what this experiment was testing exactly because Pathans aren’t pure West Asians and Lithuanians aren’t pure Mesolithic Europeans. He also claimed that Dienekes' interpretations are  wrong, because f3-statistics and rolloff tests are basically picking up (belated) signals of the Mesolithic and Neolithic peopling of Europe.

Since the aforementioned populations have the strongest presence in my dataset, and there is no consensus-opinion between genome bloggers on how to interpret the ADMIXTOOLS  i've decided to put my 5 cents and  to test ADMIXTOOLS.

For the purposes of this analysis, i created ad-hoc dataset, which includes 750 000 snps samples in 250 worldwide populations.  Next, i made 3*62 000 trios in the following form (X,Y; Z), where X and Y are two paired reference populations, and Z is one of three populations - central Russians, Poles and Belarusians. After that i carried out q3Pop analysis of those trios.

From the obtained results I have picked up only those with significant negative Z-score:

Poles:

X                   Y                   Z
Estonian    Jew-Iraqi    Polish    -0.002039    0.000179    -11.368
Jew-Iraqi    Estonian    Polish    -0.002039    0.000179    -11.368
Italian-North    Latvian    Polish    -0.001211    0.000109    -11.098
Latvian    Italian-North    Polish    -0.001211    0.000109    -11.098
Estonian    Italian-North    Polish    -0.001023    0.000093    -11.037
Italian-North    Estonian    Polish    -0.001023    0.000093    -11.037
Estonian    Jew-Iran    Polish    -0.001861    0.000172    -10.831
Jew-Iran    Estonian    Polish    -0.001861    0.000172    -10.831
Armenian    Estonian    Polish    -0.001425    0.000136    -10.505
Estonian    Armenian    Polish    -0.001425    0.000136    -10.505
Italian-South    Latvian    Polish    -0.001344    0.000129    -10.458
Latvian    Italian-South    Polish    -0.001344    0.000129    -10.458
Cypriot    Estonian    Polish    -0.001626    0.000161    -10.113
Estonian    Cypriot    Polish    -0.001626    0.000161    -10.113
 

Russians_Central

North_Amerind    Sardinian    Russian_Center    -0.004202    0.000479    -8.779
Sardinian    North_Amerind   Russian_Center    -0.004202    0.000479    -8.779
Basque    Ket    Russian_Center    -0.003771    0.000444    -8.493
Ket    Basque    Russian_Center    -0.003771    0.000444    -8.493
Karitiana    Sardinian    Russian_Center    -0.005947    0.000704    -8.453
Sardinian    Karitiana    Russian_Center    -0.005947    0.000704    -8.453
Pima    Sardinian    Russian_Center    -0.004908    0.000605    -8.117
Sardinian    Pima    Russian_Center    -0.004908    0.000605    -8.117
Ket    Sardinian    Russian_Center    -0.004295    0.000552    -7.786
Sardinian    Ket    Russian_Center    -0.004295    0.000552    -7.786
Lithuanian    Oroqen    Russian_Center    -0.00344    0.000445    -7.731
Oroqen    Lithuanian    Russian_Center    -0.00344    0.000445    -7.731
Basque    Karitiana    Russian_Center    -0.004629    0.000609    -7.596
Karitiana    Basque    Russian_Center    -0.004629    0.000609    -7.596
Basque    Pima    Russian_Center    -0.003711    0.000492    -7.545
Pima    Basque    Russian_Center    -0.003711    0.000492    -7.545
North_Amerind    Sardinian    Russian_Center    -0.003465    0.000462    -7.505
Sardinian    North_Amerind   Russian_Center    -0.003465    0.000462    -7.505
Basque    Nganassan    Russian_Center    -0.003574    0.000478    -7.471

Belarusians

Indian    Polish    Belarusian    -0.000736    0.000251    -2.935
Polish    Indian    Belarusian    -0.000736    0.000251    -2.935
Karitiana    Sardinian    Belarusian    -0.001278    0.000517    -2.471
Sardinian    Karitiana    Belarusian    -0.001278    0.000517    -2.471
Otzi    North_Amerind    Belarusian    -0.002556    0.001126    -2.271
Cirkassian    Polish    Belarusian    -0.000488    0.000231    -2.113
Polish    Cirkassian    Belarusian    -0.000488    0.000231    -2.113
Pima    Otzi    Belarusian    -0.002727    0.00137    -1.99
Pima    Sardinian    Belarusian    -0.000794    0.000431    -1.843
Sardinian    Pima    Belarusian    -0.000794    0.000431    -1.843
Otzi    Surui    Belarusian    -0.002938    0.001931    -1.522
Surui    Otzi    Belarusian    -0.002938    0.001931    -1.522

 Discussion



As it looks at first glance,the results of my ad-hoc experiment with 3qPop seems to be consistent with the findings in Patterson et al.2012 paper: "the most striking finding is a clear signal of admixture into northern Europe, with one ancestral population related to present day Basques and Sardinians, and the other related to present day populations of northeast Asia and the Americas. This likely reflects a history of admixture between Neolithic migrants and the indigenous Mesolithic population of Europe, consistent with recent analyses of ancient bones from Sweden and the sequencing of the genome of the Tyrolean ‘Iceman’".


Indeed, the admixture in Poles can be shown as the admixture between Neolithic + Mesolithic populations of Europe, Russians/Belarusians can be represented as admixture between the ancestral population of modern populations of NE-Asia/Amerinds and Neolithic populations of Europe.
However, more careful examination of results allows me to reveal the additional signals of admixtures in two of three target populations - Poles and Belarusians.
Although its is perfectly possible to treat Estonians and Latvians as modern day proxies for the NE-populations of  Mesolithic Europe, it is also obvious that these populations could have (at least in theory) the significant genetic legacy related to Baltic branch of Indo-European Corded Ware culture.  On other hand, the second component of admixture in Poles itself is a product of admixture between Near-East/Anatolian-like Neolithic populations and more recent genetic stratum, which is probably related to the massive migration of R1b people (the ancestors of 'Bell beakers') from NE-Asia to Western Europe.

Given that, i'd suggest to rewrite the components of admixtures in Poles in the following manner:

Pole=(Neolithic_populations of Europe)+"Bell Beakerish-like")+(Mesolithic_poplations)+"Corded_Ware" component) [1]


In Belarusians, the sources of signals of additional admixture are less clear and vague.
As was shown earlier, in terms of formal admixture analysis (f3 statistics), Belarusians could be represented as the admixture between Poles and Indian/Cirkassian. The first component of admixture is already known (see above [1]), the second one, according to results, must resemble the component, common to both Indian and Circkassian. From the history textbooks i've learned that the territory of modern Karachay-Cherkessia was occupied in the 1st millenium AD  by the Alans, or the Alani, who were a group of Sarmatian tribes, nomadic pastoralists of the 1st millennium AD who spoke an Eastern Iranian language which derived from Scytho-Sarmatian and which in turn evolved into modern Ossetian. The only currently known most recent ancestral population to modern Alans and modern Indians is Scytho-Sarmatian metapopulation.

Thus, we can re-write the admixture formula for Belarusians in the following manner

Belarusian=((Neolithic_populations of Europe)+"Bell Beakerish-like")+(Mesolithic_poplations)+"Corded_Ware" component)) + Scytho-Sarmatian-like

Now, after long discussion, it is time for the admixture_dating fun!



Admixture dating with ROLLOFF



To estimate the admixture date in Polish population, i used as reference populations Latvian and North-Italian


The admixture date is  119.670+-37.145 generations ago, which corresponds to 3470 +-1077 years before present, or 1510 +- 1077 AD.  The upper limits of our dating for the admixture event seem to overlap with the timescale of Unetice culture. The Bronze Age in Poland, as well as elsewhere in central Europe, begins with the innovative Unetice culture, in existence in Silesia and a part of Greater Poland during the first period of this era, that is from before 2200 to 1600 BC. This settled agricultural society's origins consisted of the conservative traditions inherited from the Corded Ware populations and dynamic elements of the Bell-Beaker people. Significantly, the Unetice people cultivated contacts with the highly developed cultures of the Carpathian Basin, through whom they had trade links with the cultures of early Greece. Their culture also echoed inspiring influence coming all the way from the most highly developed at that time civilizations of the Middle East.


To estimate admixture date in Belarusian population, i used as reference populations  Polish and Indian (note: i also lowered genetic distance threshold in ROLLOFF parameters to reduce noise from more recent admixtures)



As you can see, the signal of admixture is less detectable, and by virtue of that the margins of error in admixture dating are significantly higher than in previous example:   154.158+-87.024 generations ago (or, 4470 +-2523 years before present/2510 -+2523 years AD).




 


 

 


 




  

Thursday, September 20, 2012

The component maps of MDLP-World22 calculator


I would like to express my gratitude to the fellow members of ABF - Loxias and Wojewoda- for creating amazing maps of components and plots showing interrelationships between 22 components.

Since the readers of my blog might be interested in visual inspecting of components , i have decided to upload them on-line. 

The first batch of maps  includes "the component portraits", created by Loxias:

West-Asian


Atlantic-Mediterranean
East-Siberian
Amerind
Indian
Indo-Iranian
Indo-Tibetan
Paleo-Siberian
Pygmy
Sub-Saharan
Near-East
North-East-European
North-Siberian
North-Mesolithic-European
 The second batch includes the PCA scatter plot of components, created by Wojewoda. PCA scatter plots one PCA component vs. another PCA component of data obtained from ancestry coefficients  collected for each of MDLP World22 components: the spots are connected over time to produce a trajectory for each component.






Wednesday, September 19, 2012

Paint Me a Rainbow: Painting World 22 ancestral components

This update will be concerned with inter-related concepts of  "chromosome painting" and admixture. Modern genetics and personal genomics, particularly in the last 2-5 years, had paid very considerable attention to them, sometimes under the rubric of " determining ancestral origin of genomic segments".

Although the experiments met with moderate success, i wasn't satisfied with the results and decided to postpone the forthcoming experiments with chromosome painting to a future day. In so doing, i endorsed increasing  appreciation of the distinction between  population stratification's algorithms, implemented in LAMP and ADMIXTURE.

I have already discussed the differences between LAMP and Admixture, but in illustrating the idea of experiment, i need to turn to my previous explanation again:
 

1) ADMIXTURE software  implements a model-based approach to estimate ancestry coefficients as the parameters of a statistical model. It is also important to add that the model-based approach in ADMIXTURE is based  on the global ancestry paradigm (i.e the goal of  ADMIXTURE/STRUCTURE analysis is to estimate the proportion of ancestry from each contributing population, considered as an average over the individual's entire genome).

2) LAMP software is built upon an efficient dynamic-programming algorithm WINPOP that infers locus-specific ancestries.Genome is partitioned into chromosome segments of definite ancestral origin (overlapping, contiguous windows of SNPs) and likelihood model optimized over each window. The goal then is to find the segment boundaries and assign each segment's origin.
I understand the problem in terms of the ancestry assignment. My experience shows that methods based on  the inference of locus-specific ancestries are usually very accurate for ancestral deconvolution of genotype data that has consistently been shown to do better than more popular statistical and PCA-based methods, while being able:
1) to handle more than two ancestral populations
2) to model the paths of recombination between ancestral segments.
In June 2012, Jason Mezey Lab (Cornell University) released SupportMix - a machine learning algorithm for determining ancestral origin of genomic segments when analyzing individuals from a population with a recent or ancient history of admixture. As regards the accuracy of the software, the authors argued that SupportMix provides a robust tool for accurate and robust ancestral assignment by simultaneous analysis of a worldwide selection of ancestral populations. Such analyses will be critical for accurate assignment in the many world-wide admixed populations that are likely to have unexpected ancestry that reflects a richer history than known from anthropological or historical studies (from the provisional paper: Omberg et al.2012 "Inferring genome-wide patterns of admixture in Qataris using fifty-five ancestral populations"). The cited paper includes a number of other claims important for any analysis of genetic admixture: the accuracy of ancestry assignment was lower for more closely related populations but better than LAMP-ANC, a method that was been shown to consistently outperform other ancestry deconvolution methods.

This overly optimistic conclusion influenced my choice between LAMP-ANC and SupportMix in favor of latter. To be honest, i'm not the first genome blogger to use Supportmix - in July of 2012 Polako from Eurogenes carried out a loci-specific analysis of Finnish genomes using SupprotMix. I decided to repeat this experiment. There is, of course, a significant  difference between Polako's analaysis and my experiment. While Polako was using the modern populations as 'putative' donor populations, the final goal of my design was to imitate the results of the long-awaited update of 23andme's Ancestry Painting, which is is being updated to offer more detailed results based on approximately 20 world regions, drawn from both customer data and academic reference populations. In order to do so, i have used the dummy set of 22 simulated putative ancestral populations simulated from the allele frequencies of the World-22 calculator.

The experiment


SupportMix requires at least three input files. One file for each of the putative ancestral populations and one file containing the genetic information of the admixed individuals with additional requirement that he markers have to be phased. Each population should be represented by two files in Plink transposed format, a .tped file and a .tfam file.

The markers (80751 SNPs) were phased per each chromosome using default settings in BEAGLE software. While the original Plink format does not specify the order of the alleles in in the file, SupportMix works with phased data. Keeping that in mind, i have converted BEAGLE-phased dataset directly into Plink's tped format without pre-processing the dataset in Plink (hint: Kantele's beagle_to_tped script). Then i used UNIX text processing utilities to extract the genotypes from ancestral populations into the corresponding subsets ('references') and split the project dataset (93 individuals) into 9 subgroups.

Finally, i have interpolated the genetic map position of each SNP along chromosome using Rutgers genetic maps.

 SupportMix was run by specifying a configuration file with the default options: (window_size =400, generations_from_admixture_event=6).

Below are some screenshots of  SupportMiX output for the MDLP project participants:




Please note that each 'recipient' (i.e the project participant) is represented by two phased chromosomes, i.e V199_a and V199_b. The color legend of the components used in the analysis, has been attached to the right side of the plot.

The complete set (Chr.1-22 for all project participants) in tar.gz format (14.6 Mb) could be downloaded here.


SupportMix results: what to do next.


First of all, i encourage every participant of my project to compare their results to "chromosome painting" in MDLP World-22 calculator on John Olson's Gedmatch site. I have to contemplate the possibility that the painting on Gedmatch site might be different from that one produced by SupportMix. I would also suggest to compare SupportMix's paintings to other calculators' paintings and 23andme's Ancestry Finder, etc.

If you are familiar with basic techniques of image editing software, then it is a good idea to have your chromosomes cut&merged into the composite image (see example below):




Chromosome Painting (in 23andme's  style)



Chromosome Painting (an imitation of 23andme's Ancestry Finder)
 


 









Saturday, September 15, 2012

The reliability of 'ancestral components' in MDLP World-22 calculator: trying TreeMix on 22 ancestral components

In the recent paper of Pickrell and Pritchard outlined the most challenging problems in  the analysis of relationships between populations: 

Many aspects of historical relationships between populations  are reflected in genetic data. Inferring these relationships from dense genetic data remains a challenging and very difficult task. 

One such aspect is the (re)construction of ancestral population from precomputed allele frequencies.In the recent supervised Admixture analysis of the MDLP analysis (see previous post), i have  carried the analysis of dataset which consists of both extant populations and 'simulated un-admixed populations'. However, one can always question the accuracy of  the results of  such an experiment by posing a very simple question: is it possible to carry out the  simultaneous analysis of 22 putative ancestral populations while being independent of prior demographic information ( involving population splits, gene flow, and changes in population size). Accounting for demographic information is especially important in cases when we have only limited genetic data for the few ancestral populations that are known with greater certainty.

To the ultimate relief of genome bloggers, Joe Pickrell and Jonathan Pritchard have released a companion program to Structure/Admixture programs called TreeMix.   TreeMix uses large SNP data sets to estimate the relationships among populations  including both population splits and admixture events.  According to Pickrell and Pritchard, the new method provides a better representation of population histories than do standard tree-building methods when they  are applied to the worldwide populations.  In my opinion, the real advantage of TreeMix over traditional admixture programs is that it accounts for unknown ancestral allele frequencies. Indeed, in most cases  we do not know the ancestral values of allele frequencies, but instead only the values in sampled descendant populations.

I decided to give TreeMix a try on 22 putative ancestral ADMIXTURE components from the latest run. For this purpose, I used a conversion script, which was kindly provided by Dienekes Pontikos. This script takes ADMIXTURE P file and converts it into  a plink.treemix.gz file, which is ready for input into TreeMix.

 I ran TreeMix using following settings > treemix -i MDLPworld22.treemix.gz -root South-African -k 500   -o MDLP22world (consult TreeMix manual for explanation of program parameters).


After careful examination of the output tree, i was really surprised by how  remarkably similar the tree is to that one in the original paper of Pickrell & Pritchard:





In the next of experiment i have introduced to the tree a rudimentary model of migrations/gene flows by including -m 5 parameter. The model is, however, very rudimentary, because Pritchard and Pickrell have  modeled "migration between populations as occurring at single, instantaneous time points. This is, of course, a dramatic simplification of the migration process. This model will work best when gene flow between populations is restricted to a relatively short time period. Situations of continuous migration violate this assumption and lead to unclear results" (see the cited paper for the further discussion).


 Nevertheless, we may, though with all due caution, speculate about the hypothetical gene flow directions:

a) from East-South-Asian ancestral component -> to Tibetan ancestral component
b) from the split_point between Austronesian and Melanesian ancestral components -> to East-South-Asian ancestral component
c) from the split_point between East-Siberian and East-South-Asian -> to the split_point between  Indian and ((Austronesian;Melanesian) Tibetian)* ancestral component
d) and finally, from North-European-Mesolithic component to the split_point between West-Asian and (North-East-European; Atlantic-Mediterranean-Neolithic)* component.

The last point d) calls for a further explanation. As one can see from the attached population tree,North-European-Mesolithic component is unexpectedly shifted towards East-Asian populations.  Despite of all simplicity of test, i think that the result in consideration could probably support  a recent statement made by David Reich in the recent Reich et al. (2012) paper:
we took advantage of the fact that east/central Asian admixture  has affected northern Europeans to a greater extent than Sardinians (in our separate  manuscript in submission, we show that this is a result of the different amounts of central/east  Asian-related gene flow into these groups).
Cf. also Dienekes' comments 

If our finding is congruent Reich's hypothesis, then one could assume that he Mesolithic Europeans were Asian-shifted themselves, i.e the earliest episode of admixture could have occurred in the Mesolithic period.

UPDATE: 
The same scenario seems to be supported in (Patterson et al. 2012).


Another important question that could be raised in regard of the reliability of the inferred components is the question of 'component purity' (the notion of 'purity' is used here in rather speculative sense of Kant's Reinheit, and has nothing to do with 'racial purity'). Indeed,  a population tree (such as one presented above) could represent with the same degree of reliability both pure nested morphology of components  variable components of nested and non-nested morphology.

Fortunately for genome bloggers, we can try to tackle this problem by using three- and four- population tests for treeness from Reich et al. 2009. The 3-population test (Reich et al. 2009) allows one to detect the presence of admixture in a population X from two other populations A and B. The value
f3(X; A, B)

is negative when X does not appear to form a simple tree with A and B but appears to be a mixture of A and B (in Dienekes' laconic phrasing).  In order to calculate f3 statistics  i used the implementation of three-pop  TreeMix's  program threepop:


X                      A,B
Samoedic    North-Siberean,Atlantic_Mediterranean_Neolithic    0.00206352    0.000146163    14.118
North-Amerind    South-America_Amerind,South-African    0.00564354    0.00034763    16.2344
Samoedic    North-Siberean,North-East-European    0.00230317    0.000134738    17.0936
North-Amerind    South-America_Amerind,Melanesian    0.0063216    0.000345733    18.2846
Sub-Saharian    South-America_Amerind,South-African    0.0058646    0.000302046    19.4162
Sub-Saharian    Pygmy,South-America_Amerind    0.00420139    0.000216014    19.4496
Sub-Saharian    Pygmy,Paleo-Siberean    0.00426537    0.000217229    19.6354
Sub-Saharian    Pygmy,North-Siberean    0.00438367    0.000221767    19.767
North-Amerind    Pygmy,South-America_Amerind    0.0058802    0.000295087    19.927
Sub-Saharian    Pygmy,Mesomerican    0.00439921    0.000219885    20.0068
Sub-Saharian    Pygmy,Arctic-Amerind    0.00427798    0.000212131    20.1667
Samoedic    South-America_Amerind,South-African    0.00679588    0.000332084    20.4643
North-Amerind    South-America_Amerind,Austronesian    0.00658824    0.000314661    20.9375
Samoedic    North-Siberean,South-African    0.00504898    0.000240205    21.0195
North-Amerind    South-America_Amerind,Paleo-Siberean    0.00580508    0.000275451    21.0748
 



The full f3statistics file could be downloaded here.


After the careful examination of  f3statistics for 22 putative ancestral components, i haven't found negative values (negative value is a strong unambiguous signal of admixture). Thus, we could safely reject the hypothesis of mixture in ancestral components, and assume that each X components forms a simple tree with A and B components.











Behind the Curtains: MDLP World 22 showcase


Preliminary remarks

As you all may know, the MDLP  blog hasn't been updated since February 2012.
Half of year ago i promised myself that i would stop writing new posts on the MDLP blog before i'll finally get my scientific report on blog written.  Since I had to prioritize the completion of a scientific paper over the routine of blog posting,I was unable to continue updating the blog on a regular basis due to a lack of time, and had to make a change in how I conducted my research. So i decided to abstain from posting on the MDLP blog for a couple of month, being focused on more important matters. Despite of all limitations, i kept  working secretly on the MDLP project, collecting necessary data and performing different 'genomic' experiments in order to achieve my final goal (publishing of paper).  The results of secret experiments with new genomic samples and tools eventually leaked to the curious public, spawning immense interest in my project. After releasing a new version of my own modification of DIYDodecad calculator on Gedmatch.com, i was literally flooded by emails from Gedmatch.com users asking me  questions they wanted me to answer.

I understood the strategical mistake of releasing poorly documented data/analysis on Internet  and felt obliged  to explain details. Obviously, i will start new series of  the blog posts by covering the project feature the people most interested in, i.e the MDLP World22 calculator.
 
The population dataset of MDLP World22 calculator.

The reference population dataset of the calculator was assembled in PLINK by intersecting and thinning the samples from different data sources: HapMap 3 (the filtered dataset CEU,YRI,JPT,CHB), 1000genomes, Rasmussen et al. (2010)HGDP (Stanford) (all populations)Metspalu et al. (2011),Yunusbayev et al. (2011), Chaubey et al. (2010) etc. Furthermore i handpicked random 10 individuals from each European country panel in POPRES dataset, or the maximum number of individuals available otherwise, to select the POPRES European individuals to be included in our study. Finally, in order to evaluate the correlation between the modern and the ancient genetic diversity, i have also included ancient DNA genomic samples of Ötzi,(Keller et al.(2012)) Swedish Neolithic samples Gök4, Ajv52, Ajv70, Ire8, Ste7 (Skoglund et al. (2012)) and 2 La Braña individuals from the Mesolithic sites of the Iberian Peninsula (Sánchez-Quinto et al.(2012)). Then i added 90 samples of individuals-participants of our MDLP project.  After merging the aforementioned datasets and thinning the SNP set with PLINK command to exclude SNPs with missing rates greater than 1% and minor alleles, i filtered out duplicates, the individuals with high pairwise IBD-sharing  (estimated in Plink as as the average fraction of alleles shared between two individuals over all loci) and the individuals with kinship coefficient suggesting relatedness (kinship coefficients were estimated in KING software). Also i had to filter out individuals with more more tham 3 standard deviations from the population averages. Since kinship coefficient is robustly estimated by HWE (Hary-Weinberg expectations) among SNPs with the same underlying allele frequencies, SNPs showing strong deviation (p < 5.5 x10−8) from Hardy-Weinberg expectations were removed from the merged and filtered dataset. After that I filtered to keep the list of common SNPs present in Illumina/Affymetrix chips and performed  linkage disequilibrium based pruning using a window size of 50, a step of 5 and r^2 threshold of 0.3.

This complex sequence of consequent operations with the initial reference and project datasets yielded a final dataset which included 80751 SNPs  in 2516 individuals from 225 populations.


ADMIXTURE analysis


 As always, the final dataset in PLINK linked format was further processed in ADMIXTURE software. Sketching the plan for the design of ADMIXTURE test,  i had to face the difficult problem: as it has been shown in (Patterson et al.2006) the number of markers needed to resolve populations in ADMIXTURE analysis is inversely proportional to the genetic distance (Fst ) betweeen the populations. According to ADMIXTURE best practice, it is believed that 10,000 markers  are suffice to perform GWAS correction for continentally separated populations (for example, African, Asian, and European populations FST > .05) while more like 100,000 markers are necessary when the populations are within a continent (Europe, for instance, FST < 0.01).
To increase the accuracy of ADMIXTURE results i decided to use a method proposed by  Dienekes' for converting allele frequencies into 'synthetic individuals'(see also Zack's example). The idea is fairly simple: run an unsupervised ADMIXTURE analysis once to generate allele frequencies for your K ancestral components; then generate zombie populations using these allele frequencies; whenever you want to estimate admixture proportions in new samples run supervised ADMIXTURE analysis using the zombie populations.  Like any genome blogger engaged in the task of evaluating admixtures in samples, i must grapple with obvious question of the reliability of this approach.  Although i am aware of  methodological controversies in using simulated individuals, i would rather concur with Dienekes who considered "synthetic individuals" the best abstract proxies for the ancient ancestral populations. But my purpose is served if i can use the approach used by Dienekes and Zack to obtain meaningful results. To begin with, i routinely ran unsupervised ADMIXTURE  K=22 analysis (assuming 22 ancestral populations) which yielded the admixture proportions of individuals from these K populations, as well as the allele frequencies for all SNPs for each of 22 ancestral populations (below are conventional names for each of inferred components in order of appearance):

Pygmy
West-Asian
North-European-Mesolithic
Tibetan
Mesomerican
Arctic-Amerind
South-America_Amerind
Indian
North-Siberean
Atlantic_Mediterranean_Neolithic
Samoedic
Proto-Indo-Iranian
East-Siberean
North-East-European
South-African
North-Amerind
Sub-Saharian
East-South-Asian
Near_East
Melanesian
Paleo-Siberean
Austronesian
Therefore i took the allele frequencies which were computed earlier in unsupervised Admixture K=22 for the merged dataset, pooled them into PLINK and generated 10 "synthetic individuals per ancestral component) using PLINK command --simulate.  When the simulation had been  finished, i visualized the distance between simulated individuals using multi-dimensional scaling:



 

As a next step,i included simulated individuals you have as part of a new reference population (including 220 simulated individuals in 22 simulated populations).Then, I ran ADMIXTURE anew, this time in “supervised” mode for K = 22 (with simulated individuals being 'reference' individuals). The Admixture K=22 converged in 31 iterations (37773.1 sec) with final loglikelihood:-188032005.430318 (below are Fst divergences between estimated 'ancestral' populations):


The Fst distance/divergence matrix was used for inferring a most probable NJ-based topology of component distance tree (outgroup: South-African):



 The individual 'supervised' ADMIXTURE  results (in Excel spreadsheet) for the project participants have been uploaded to GoogleDocs (please note that the average results for reference populations is also available on special request).

MDLP World22 DIYcalculator

The output files of Admixture K=22 supervised run (average values of admixture coefficients in reference populations and FsT values)  were used for designing a new version of  the MDLP DIYcalculator, which is better known by its codename "World22" (online version is available in AdMix-Utilities section of Gedmatch under MDLP project). MDLP DIYcalculator itself is based on the code of Dodecad DIY calculator (c)ourtesy of Dienekes Pontikos and was developed as part of the Dodecad Ancestry Project. In its Gedmatch implementation MDLP 'World22' DIYcalculator is paired by MDLP 'World22' Oracle, also based on Dienekes' and Zack's code (Harappa/DodecadOracle). The 'Oracle' is designed to find in a single population mode your closest (closest in terms of similarity) population from MDLP ''Word22' admixture results. In a mixed mode, Oracle considers all pairs of populations, and for each one of them calculates the minimum Fst-weighted distance to the sample in consideration, and the admixture proportions that produce it.

Please notice: 'ancestral' populations (i.e 'simulated populations' from the previous step - see above) are labeled in Oracle results as (anc), while the 'real world' modern and ancient populations are marked as "derived".

If you have troubles with understanding/interpreting the results of Oracle and DIYcalculcator, please consult the corresponding topics on Dodecad and HarappaWorld blogs. It is not of avail to repeat in this blog everything they wrote in their own blogs.


What the heck are MDLP Word-22 components?

 One of those questions that i usually keep getting in emails is what do the various reference populations and ancestral components for my World K=12 and World-22 analyses mean. I've already provided hints to the answer  earlier, but - as old Chinese proverb says - one picture is worth ten thousand words. That's why i decided to display the admixture coefficients spatially on the globe surface. Following Francois Olivier, who proposed to use the graphical library of the statistical software R to display  spatial interpolates of the admixture coefficients (Q matrix) in two dimensions (where spatial coordinates are recorded as longitude and latitude), i created  2 contour maps per component.

Pygmy (modal in Biaka and Mbuti population)



West-Asian (bimodal component with peaks in Caucasian populations and south-western part of Iran, equal to Dienekes' Caucasian/Gedrosia component) 



 North-European-Mesolithic (local component with peaks in European Mesolithic samples of La_Brana and  modern North-European Saami population).


 Tibetan (Indo-Burmese) component (Himalay, Tibet)


Mesomerican (major genetic component in Native Americans from Mesoamerica)



North-Amerind (the 'native' component in North American Natives)




South-Amerind (the 'native' component in South American Natives)






  Atlantic-Mediterranean-Neolithic (the main genetic component  in Western and South-Western Europe)



  The rest contour maps for all components could be downloaded here.