My Blog List

Monday, November 14, 2011

Some achievements of the MDL project

Looking back at the project's goals posited  6 months ago, we would like to recapitulate some of the most important goals in order to analyzes how they have been achieved:

1.Performing the comprehensive Plink analysis, including the estimation of homozygous ROH (shared clusters and groups of homozygosity), possible Mendelian errors, extended LD-haplotypes (based on values of R2), shared IBD segments and IBS matrix (Plink format).

Although we performed all described types of Plink analysises and eve shared some results on the project's blog, we didn't consider these results worth of extensive coverage. And likewise, there was no interest in those analysises on behalf of the project's members.

Experiments with relatedness
Graphoanalytical approach to visualizing relatedness
IBD sharing
IBS similarity matrix in R

2.Phasing the genotype files, i.e establishing the haploid phase (this is a separate analysis demanding genotypes of your parents, so it will not be performed on a regular base) (Beagle or Merlin output format). 

We performed ad-hoc phasing of the genotypes in our project (MDLP) and, in order to assess possible discrepancies between phased and unphased data, we performed ADMIXTURE analysis (with 4 assumed clusters K=4) separately for original unphased dataset and BEAGLE-phased dataset.

Analyzing admixture in phased v.unphased dataset

3. Using AISconvert (based on HIRsearch) and Germline software to detect IBD segments.

Used only occassionally in combination with other analyses

Analyzing admixture in phased v.unphased dataset 
Grapho-analytical approach to the visualisation of IBD shared segments
IBD sharing


4.Using ADMIXTURE/STRUCTURE software for detecting admixture clusters and claculating allele frequencies.

We performed a plenty of ADMIXTURE and STRUCTURE runs (using different a priori number of assumed clusters under different models of  admixture). Discussions of ADMIXTURE results contibuted  the most signficant part to the MDLP's blog.
The allele frequencies, estimated in K=7 Admixture run, were provided for creating a custom modification of DIYDodecad's calculator (MDLP).

Analyzing admixture in phased v.unphased dataset
First results: Admixture unsupervised run
Admixture analysis: sorted after Baltic-Slavic component
Admixture results: Baltic-Slavic
Admixture analysis: the rest of groupings
The output of the PLINK and ADMIXTURE algorithms
Admixture clusters, Mclust and populations concordance
DIYDodecad calculator v2.0 for my BGA project (MDLP).
Root Means Square Comparison Excel 2007 Macro Enabled XLSM spreadsheet for the Magnus Ducatus Lithuaniae Project data

and many more ..

5. Creating MDS and  PCA plots
PCA plots (Eigensoft)

PCA plots for reference populations and project participants
MDS and PCA plots: for V157-V247
A close-up on "the core" of the MDL project

6.Creating RHHmapper schemes showing the location of rare heterozygous and homozygous genotypes

RHH mapper: results for V158-V165 and V201-V202








7 months of MDLP project: alpha phase is over

We would like to announce that the project has been active for more than 6 months, i.e fairly long to accomplish some of the posited goals. The main goal of the project's preliminary (alpha) phase was to collect a statistically reliable sample for obtaining the statistically significant results. The current dataset (MDLP v2) includes 531 unrelated individuals (379 males, 152 females) with 310652 SNPs, of those 183 individuals (48 Romanians, 2 Russians,17 Chuvashes,12 Uzbeks,16 Turks,18 Armenians,15 Lezgins,20 Georgians,19 Hungarians,8 Lithuanians,8 Belorussians) from Behar et all (2010) dataset.41 individuals from HGDP (25 Russians, 16 Adygei), 175 individuals from the 1000 Genomes Project (83 British, 92 Finns) and 62 individuals from Yunusbayev et all. (2011) paper (14 Mordovians, 16 Nogays,13 Bulgarians,19 Ukrainians).

 The ethnic distribution of the whole set would look as follows (ethnic groups in red need more participants/samples )


Belarussian 18
Adygei 16
Armenians 18
Aszkenazi 2
Bulgarians 13
Chuvashs 17
Finns 92
British 83
Georgians 20
Hungarian 20
Latvian 1
Lezgins 15
Lithuanians 27
Mordovians 14
Nogays 16
Ossetians 14
Norwegians 2
East Germans 7
Others  8
Poles 18
Romanians 14
Russians 36
Swedish 2
Turks 16
Ukrainians 30
Uzbeks 12

Another interesting characteristics of sample is that one of average inbreeding coefficient in each particular population,  based on the observed versus expected number of homozygous genotypes in given population.


FID F-coefficient
Lithuanian-average 0.0158738
Finn-average 0.01375742
GBR_Orkney-average 0.013074288
Lezgin-average 0.012808472
Belorussian-average 0.011024444
GBR_Cornwall-average 0.010527961
GBR_Kent-average 0.009641047
Georgian-average 0.00949285
Turk-average 0.0093435
Hungarian-average 0.007138795
Adygei-average 0.006826329
Romanian-average 0.006763092
Russian-average 0.006179208
Uzbek-average 0.005255747
Armenian-average 0.004329326
Chuvash-average 0.004147971



The following characteristic of  MDLP - an average number of shared IBD segments per population is especially valuable for evaluating the genomic structure of population. I've limited the results to Slavic populations only.



Poles 0.878788
Belarusians 0.722008
Ukrainaians 0.676113
Russians 0.561878
Lithuanians 0.548961


 

Friday, November 4, 2011

The major revision of MDLP: adding Yunusbayev et all.2011 data

Earlier this month i added Bulgarians, Ukrainians and Mordovians from the reference samples in Yunusbayev et all.2011 paper and calculated PCA loadings of new obtained

As promised earlier, i have wasted a couple of hours on plotting PCA loadings (from Eigensoft analysis of my MDL BGA project) in graphical interface of indispensable R-package "BiplotGUI". I've made a lot of efforts combining diffirent types of statistical tests into one meaningful illustration.

Below are results of my experiments with Biplot.





Another cool feature of BiplotGUI is that it fully supports rgl, which is a 3D visualization system based on OpenGL. It provides a medium to high level interface for use in R, currently modelled on classic R graphics, with extensions to allow for interaction and creation of 3D animations/movies.

I've not figured yet how it works, but the next time i will be doing PCA analysis, i'll try to make 3D rotating animations for members of my project. 


(Via) DODECAD:Comparing different ADMIXTURE runs using Zombies

Dienekes Pontikos of DODECAD BGA project  was using our MDLP calculator (based on DIYDodecad methodological paradigma) as proof concept for comparing/mapping the Eurasian components inferred from the Dodecad-dv3 dataset  (West Asian, West European, East European, Mediterranean, Northeast Asian, Southeast Asian) against MDLP components (Scandinavian,Volga_Region, Altaic, Celto_Germanic, Caucassian_Anatolian_Balkanic, Balto_Slavic, North_Atlantic).

This is what he did :

".. To compare components across different projects; there has been a proliferation of different ancestry projects since the launching of Dodecad nearly a year ago, and since all of them slightly different individuals/SNPs/terminology, it is quite useful to be able to gauge how one component from one project maps onto other components in other projects. As proof of concept, I took the MDLP calculator from the Magnus Ducatus Lituaniae Project and generated 50 zombies for each of its 7 ancestral components:
  1. Scandinavian
  2. Volga_Region
  3. Altaic
  4. Celto_Germanic
  5. Caucassian_Anatolian_Balkanic
  6. Balto_Slavic
  7. North_Atlantic
 I then inferred the ancestry of the MDLP zombies using Dodecad v3, and vice versa. Since Dodecad v3 also includes populations (e.g., Africans) not considered by MDLP, I did not try to map those onto MDLP.


I will comment on the MDLP-to-dv3 mapping:
  1. The MDLP "Scandinavian" component appears to be West/East European with a little Mediterranean and a little Northeast Asian
  2. The MDLP "Volga_Region" component appears to be East European with some Northeast Asian
  3. The MDLP "Altaic" component is West Asian+Northeast Asian+Southeast Asian. Note that in Dodecad v3, the Northeast Asian component peaks at Chukchi, Nganasan, and Koryak, and most other east Eurasian populations have much less of it
  4. The MDLP "Celto-Germanic" component is (surprisingly) Mediterranean-dominated. One possible interpretation is that in the context of MDLP this captures one aspect of the difference between Southwestern and Northeastern Europe -higher Mediterranean in the former-, whereas the...
  5. ... MDLP "North-Atlantic" component seems to be entirely West European, and is capturing a different aspect of east-west variation in Europe.
  6. The MDLP "Balto-Slavic" appears the reverse of the "Celto-Germanic" with lower Mediterranean and reversed East/West European
  7. Finally, the MDLP "Caucassian_Anatolian_Balkanic" component is predictably mainly West Asian, but with a little Mediterranean and Southwest Asian as well
A different way of comparing the different components is to include them all in a joint MDS plot, or calculate various types of distances between them (e.g., Fst).

For example, the first couple of dimensions are dominated by the African/Asian components of Dodecad v3 that are not present in MDLP.Notice, however, the position of "Altaic", right where one might expect to find it between West and East Eurasians.





It appears that the "North_Atlantic" component may be centered on a small number of related individuals."

PS. Our MDLP calculator has been occasionally used by various people, and thus results arecurrently being disseminated over the large number of  the different Internet communities of DNA-genealogy/molecular anthroplogy's hobbysts. We are going to collect the results made publicly available by those hobbysts and out them into one spreadsheet for the futher analysis.   

Thursday, October 6, 2011

Analyzing admixture in phased v.unphased dataset

Determination of haplotype phase is becoming increasingly important as we enter the era of large-scale sequencing because many of its applications, such as imputing low-frequency variants and characterizing the relationship between genetic variations in different populations. Haplotype phase can be generated through laboratory-based experimental methods, or it can be estimated using computational approaches. We assess the haplotype phasing method that is aravailable in BEAGLE software, focusing in particular on using its output in ADMIXTURE analysis.

For simplicity's sake we have selected individuals from "Balto-Slavic"  cluster (the cluster attribution of individuals were inferred from Dienekes's Mclust using 11 MDS dimension), which is the major cluster of our project. Here is an all inclusive list of IDs for selected participants of our project:

V158
V157
V160
V202
V169
V170
V171
V174
V176
V177
V180
V181
V188
V189
V196
V205
V208
V211
V215
V218
V220
V221
V222
V228
V225
V232
V236
V237
V235
V231
V244
V246
V238







We had  thinned the genotype data of selected individuals to c.100 000 SNPs, removing SNPs in strong LD and  low quality SNps.  After that we used  GERMLINE pipeline for phasing PLINK format data with BEAGLE and processing in GERMLINE (phasing was performed in a homologous populations), Then, in order to assess possible discrepancies between phased and unphased data, we performed ADMIXTURE analysis (with 4 assumed clusters K=4) separately for original unphased dataset and BEAGLE-phased dataset.

To our surprise, we haven't be able to find expected signficant differences between phased and unphased multi -SNP markers genotypes (the range of difference is c.1-5%).

Unphased data:
Phased data:


Spreadsheet with ADMIXTURE results can be found here.

Tuesday, October 4, 2011

Simulated SNP-populations of MDLP

Yesterday I had set out to repeat  "simulation" experiments with a SNP dataset of my project's dataset, using PLINK's simulation techniques first described (in terms of population genetics) by Dienekes (the analogous experiments were performed by Harappa DNA BGA project and Eurogenes BGA project).
Synthetic "ancestral" populations (Altaic, Anatolian-Balkanian, Balto-Slavic, North-Atlantic, Scandinavian, Volga-Uralic and Celto-Germanic) were simulated using standard PLINK's simulation routine, with each ""synthetic" population including 5 generated synthetic individuals:
plink --simulate wgas1.sim --make-bed --out sim11
plink --simulate wgas2.sim --make-bed --out sim12
etc. ..
In data simulation, we assumed that  each of 7 clusters defined by specific combination of  allele frequencies of c.100000 Snps (obtained from ADMIXTURE K=7 run under unsupervised model) represents one ancestral pupulation.

Since I was interested in PCA loadings of "ancestral" populations, i used Eigensoft for explicit modeling  differences between different components" along continuous axes of variation. The calculated PCA loadings were then visualized as interactive biplot in R-package BiplotGUI  using the following  Biplot's command:

> Biplots(Data = PCA[, -1], groups = [, 1])

Afterwards  i performed three statistical tests on imported PCA loadings: linear regression, circular regression and procrusted analysis.







MDLP modification of DIYDodecad calculator: additional instructions/ideas

I have come up with another idea of how to use the estimated frequencies of my project for inferring the origin of shared HIR segments. Suppose, for example, that a Lithuanian/Belorussian and a Norwegian share some HIR segment. This could be:

1)Balto-Slavic-like ancestry in the Norwegian individual
2)Scandinavian-like ancestry in the Lithuanian individual
3)third party ancestry in both individuals

Using byseg 500 50 mode in DIYDodecad, AncestryFinder file, and MDLP allele frequencies, i was able to predict the origin of the HIR segment by picking one of three scenarios:

1) if the Lithuanian sees an excess of Scandinavian, then he should pick the second scenario
2) if he sees nothing unremarkable, the first one
3) if an excess of some component (relatively) low in both Lithuanian and Norwegian (e.g., Caucassian-Anatolian), the third.


After analyzing my "Scandinavian" "matches" from AF's file, i can conclude that DIYDodecad analysis (in byseg mode) reveals the presence of all three possible scenarios:

Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
X Sweden Sweden United States United States 1 156.4 160 3.6 5.6
60%-Scenario nr.1 (Scandinavian-like ancestry in the Balto-Slavic individual)
X Denmark Denmark Denmark Denmark 1 87.9 94.3 6.4 6.3
Scenario nr.1 (Scandinavian-like ancestry in the Balto-Slavic individual)
X Norway Norway Norway Norway 1 62.1 65.4 3.3 5
Scenario nr.1 (Scandinavian-like ancestry in the Balto-Slavic individual)
Anonymous0454 Finland Finland Finland Finland 1 104.8 110.3 5.5 6.4
Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
X Sweden Not Provided Not Provided Not Provided 2 82.2 88.1 5.9 5
60%-Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
X Sweden Sweden United States United States 3 174.9 178.1 3.2 5
80%-Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
Anonymous0353 Denmark Denmark Denmark Denmark 3 68.4 73.3 4.9 8.1
50%-Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
Anonymous0245 Denmark Denmark Denmark Denmark 4 70.5 77.3 6.8 5.3


HLA-MHC group
-Scenario III - third party (Celto-Germanic) ancestry in both individuals

X Sweden Sweden United States United States 6 25.8 34.2 8.4 5.4
X Sweden Denmark United States United States 6 27.6 34.5 6.9 5.1
Anonymous0207 Sweden Sweden Not Provided Not Provided 6 25.6 34.1 8.5 5.3
X Norway Not Provided Belgium Not Provided 6 29.7 36.1 6.4 5.3
Anonymous0370 Sweden Sweden Not Provided Sweden 6 30.7 36.6 5.9 5.4
X Denmark Denmark Denmark Denmark 6 25.6 35.5 9.9 5.9
Anonymous0439 Denmark Not Provided United States Hungary 6 26.3 36 9.7 6
X Denmark Denmark Denmark Denmark 6 26 34 8 5.1
X Norway Norway Norway Norway 6 25.6 34.2 8.6 5.6

Chr14. group - 70%-Scenario nr.2 (Balto-Slavic-like ancestry in the Scandinavian individual)
Anonymous0001 Norway Norway Norway Norway 14 38 48.8 10.8 6.5
Anonymous0037 Sweden Norway United States United States 14 39.2 48 8.8 5.1

Scenario nr.1 (Scandinavian ancestry in the Balto-Slavic individual)
Anonymous0002 Denmark Denmark United States United States 15 43.9 51.3 7.4 6