I revised my earlier test about I1-L22 trying to figure Scandinavian and Finnish clades using TRMCA method based on 67 STR markers. The main reason for doing this is new available CTS2208 samples. It is really fascinating to see how CTS2208 divides L22 subclades into two brances, implying the Finnish "Bothnian" clade being older than the estimated age of 1850 years. Here are recent TMRCA estimates
L22 - 4100 BP
Z74 - 4100 BP (It is not credible to assume both clades being 4100 years old and L22 is likely older than predicted)
P109 - 3400 BP
CTS2208 - 2800 BP
L205 - 1400 BP
L287 - 1850 BP
L258 - 1700 BP
The logic goes that downstream clades can be older than the calculated TMRCA, at the maximum as old as the TMRCA of its nearest known upstream clade.
Here is also a tree figure. 67 STR markers are not enough to create a perfect tree, but it gives anyway certain idea of the close relation of the "Bothnian" and CTS2208.
tiistai 31. toukokuuta 2016
torstai 5. toukokuuta 2016
Comparison of Ice Age and modern Europeans, Ice Age remix
Thanks to the new study "The genetic history of Ice Age Europe" and the corresponding data we have now a lot more really old human samples. As a quick experiment I made some comparisons between those ancient samples, following the grouping presented in the study, and modern Europeans. Using dstat and selected third populations from America, Asia and Europe I try to infer the amount of common ancestry of selected Europeans and Karitians, Hans and Frenchmen insofar it goes to selected ancient samples.
The dstat formula was d(European population, Karitian/Han/French ; ancient sample group, Chimp)
06.05.16 20:05 There was a small error in El Mirón numbers, showing somewhat too low similarity for Europeans. Now corrected.
15.05.16 11.00 Added dstat-gtaphics (as above) regarding Northeast Europe:
16.05.16 18:45
Added GoyetQ116-1 to the first series of graphics.
The dstat formula was d(European population, Karitian/Han/French ; ancient sample group, Chimp)
06.05.16 20:05 There was a small error in El Mirón numbers, showing somewhat too low similarity for Europeans. Now corrected.
15.05.16 11.00 Added dstat-gtaphics (as above) regarding Northeast Europe:
16.05.16 18:45
Added GoyetQ116-1 to the first series of graphics.
keskiviikko 20. huhtikuuta 2016
Neolithic and Bronze Age Irish samples, compared to modern populations
It was worth of waiting for a few weeks to see these Irish samples, especially because I already expected that Irish insular samples could reveal new things about ancient people who lived in Northwest Europe. You see the original study here. There are four samples, three from Rathlin Island in Northern Ireland and one sample from Ballynahatty, which locates in Northern Ireland. Two of Rathlin samples are of low quality and don't work well with my database based on Estonian Biocentre's data. Maybe I'll download them later to the Lazaridis' database. The third Rathlin and Ballynahatty samples are however excellent.
Picking from the study
- Ballynahatty, a Neolithic woman (3343–3020 cal BC)
- Rathlin, in context of an early megalithic passage-like grave, an Early Bronze Age man from Rathlin Island (2026–1885 cal BC)
I was really excited when started to analyse Rathlin samples, because it was possible that it would reveal new knowledge about ancient people who lived in North Europe before eastern Bronze Age steppe migrations. I decided to compare them to present-day population instead of using ancient samples, to make results touchable. At first I tested which of modern populations are closest Rathlin and Ballynahatty samples and found that the Rathlin genome emphasized still Irish people. Ballynahatty sample was closest present-day Sardinians, representing typical Neolithic era.
After processing all this from fastq-files 1) I made two qpDstat comparisons to find out who of modern populations resembles best those ancient Irish samples in comparison with best fits of modern populations. In comparison with the Rathlin man I included also my project samples, mainly Finnish and Swedish individuals.
Rathlin and modern populations
Ballynahatty and modern populations
Inspired by the western origin of Saami people I made one comparison more using another database to get reliable results with the Saami sample introduced by Haak et al. 2015. It looks like, despite of the remarkable North Asian admixture, they have Rathlin like ancestry more than Eastern Finns, who have less North Siberian.
Saami between ancient samples, using the arrangement seen already in my previous post
FI15 is from Northern Karelia, FI12 is western Finnish, FI10 is from Finnish Lapland.
Finally, after downloading and testing DNA.LAND's admixture program, I made some admixture analyses. You can find and download the software from their site, here. This small program is based on allele frequencies and probably the method is Markov chain Monte Carlo. It is not based on original alleles and genetic drift, thus there is always a residual admixture. There are also other weaknesses, what kind of, it could be a new topic. Now I only say that in my opinion it has problems in composing kinship populations with different minor admixtures.
Two results using references downloaded from DNA.LAND
Rathlin1
CSAMERICA 0.00697236
KALASH 0.0165295
NEEUROPE 0.223957
NEUROPE 0.731415
PATHAN-SINDHI-BURUSHO 0.00252671
SWEUROPE 0.0185991
Ballynahatty
ITALY 0.0116662
SARDINIA 0.565326
SWEUROPE 0.423008
Two results using my Estonian-BC database as reference
Rathlin1
Bulgaria 0.0524726
Colombian 0.00995061
Ireland 0.213786
Kalash 0.00510618
Latvia 0.0102334
Lithuania 0.221711
Orcadian 0.140857
RU_Smolensk 0.0244413
Scotland 0.207837
Udmurtia 0.0262457
Welsh 0.0840646
Ballynahatty
Basque 0.0446958
Ireland 0.0547766
NorthItaly 0.0953578
Sardinian 0.569061
Scotland 0.0409351
Sicily 0.0670495
Spain 0.12727
Tuscany 0.000854723
1) I have changed my fastq-process. Although BWA is an excellent program in mapping reads, it's automatic trimming is not powerful enough and now I have rerun also all older samples using separate trimming program.
Picking from the study
- Ballynahatty, a Neolithic woman (3343–3020 cal BC)
- Rathlin, in context of an early megalithic passage-like grave, an Early Bronze Age man from Rathlin Island (2026–1885 cal BC)
I was really excited when started to analyse Rathlin samples, because it was possible that it would reveal new knowledge about ancient people who lived in North Europe before eastern Bronze Age steppe migrations. I decided to compare them to present-day population instead of using ancient samples, to make results touchable. At first I tested which of modern populations are closest Rathlin and Ballynahatty samples and found that the Rathlin genome emphasized still Irish people. Ballynahatty sample was closest present-day Sardinians, representing typical Neolithic era.
After processing all this from fastq-files 1) I made two qpDstat comparisons to find out who of modern populations resembles best those ancient Irish samples in comparison with best fits of modern populations. In comparison with the Rathlin man I included also my project samples, mainly Finnish and Swedish individuals.
Rathlin and modern populations
Ballynahatty and modern populations
Inspired by the western origin of Saami people I made one comparison more using another database to get reliable results with the Saami sample introduced by Haak et al. 2015. It looks like, despite of the remarkable North Asian admixture, they have Rathlin like ancestry more than Eastern Finns, who have less North Siberian.
Saami between ancient samples, using the arrangement seen already in my previous post
FI15 is from Northern Karelia, FI12 is western Finnish, FI10 is from Finnish Lapland.
Finally, after downloading and testing DNA.LAND's admixture program, I made some admixture analyses. You can find and download the software from their site, here. This small program is based on allele frequencies and probably the method is Markov chain Monte Carlo. It is not based on original alleles and genetic drift, thus there is always a residual admixture. There are also other weaknesses, what kind of, it could be a new topic. Now I only say that in my opinion it has problems in composing kinship populations with different minor admixtures.
Two results using references downloaded from DNA.LAND
Rathlin1
CSAMERICA 0.00697236
KALASH 0.0165295
NEEUROPE 0.223957
NEUROPE 0.731415
PATHAN-SINDHI-BURUSHO 0.00252671
SWEUROPE 0.0185991
Ballynahatty
ITALY 0.0116662
SARDINIA 0.565326
SWEUROPE 0.423008
Two results using my Estonian-BC database as reference
Rathlin1
Bulgaria 0.0524726
Colombian 0.00995061
Ireland 0.213786
Kalash 0.00510618
Latvia 0.0102334
Lithuania 0.221711
Orcadian 0.140857
RU_Smolensk 0.0244413
Scotland 0.207837
Udmurtia 0.0262457
Welsh 0.0840646
Ballynahatty
Basque 0.0446958
Ireland 0.0547766
NorthItaly 0.0953578
Sardinian 0.569061
Scotland 0.0409351
Sicily 0.0670495
Spain 0.12727
Tuscany 0.000854723
1) I have changed my fastq-process. Although BWA is an excellent program in mapping reads, it's automatic trimming is not powerful enough and now I have rerun also all older samples using separate trimming program.
keskiviikko 23. maaliskuuta 2016
Two-fold ancestry of Finnish people
It has been a common idea, especially among linguists, to say that Baltic Finnic languages came from the Volga region, from so called Volga river bend near Samara. It is a carefully cherished tradition in Finnish science, but any movement of people from there to Finland is still without genetic evidences. Now I am going to prove something which contradicts with this idea of the Volga origin of Finns, or at least gives a new view about it. I'll show a plausible genetic evidence of Volga-Saami connection using the Saami sample (Haak et al. 2015 and Lazaridis et al. 2014), which shows very high similarity with the ancient Eneolithic Samara sample (Mathiesson et al.).
The other half of my Finnish story tells about ancient Central-European influence in Finland. Around 20% of Finnish samples from the 1000genome project show Corded-Ware similarity comparable to Estonians and Lithuanians, and Western Finnish project samples show equally Corded-Ware similarity with Swedes, some even more, despite of the fact that they are much more "eastern" when compared to present-day Swedes.
This Finnish duality doesn't tell were and when the mixing occurred and so far I have not seen any genetic evidence about the Baltic Finnic origin. It looks very possible that genetically Baltic Finns were born somewhere in region from Estonia to White Sea, no matter what the origin of Baltic Finnish language could have been.
Saami results
Saamis are genetically closer for Eneolithic Samara people than Mordovians (Mordva) and Chuvashes. Worth noticing is that Mordovians, who live near Volga are not closer those ancient people living in Samara. Saami people live thousands kilometers and thousands years away from what was the suggested Volga home range. Siberian admixture of Chuvashes roughly equals to Saami Siberian. This statistic has however very limited use, because Saami people are not Central Europeans, but still the statistic shows them being comparable to Central Europeans when compared to ancient East European samples. What could be the best outcome?
Probably some readers can think that the Eneolithic Samara - Saami - Finnish genetic connection is only based on the amount of Siberian. It is not true and easily proved false. Chuvashes and Mansi people (and Komis, not included) with high Siberian admixture are far away from the Eneolithic Samara, definitely not comparable to the Saamis. Similarly those Finns being closest Eneolithic Samara have less Siberian than Russians living in Archangel and Pinega regions in Russia (look project results).
Only people in northernmost Europe beat Saami_WGA in comparison with Eneolithic Samara. Have to admit, this is a bit complicated question. Then let's look at another perspective of supposed Finnish ancestry, Corded Ware samples. It is less complicated.
Corded Ware results
Only Lithuanians beat the Finnish CW-group (20% of Finnish samples from the 1000g project after removing outliers) when the test is done using over half million SNPs. Even Lithuanians would be beaten with more homogeneous Finnish sample group. There is all variations from very CW-looking to only moderately CW-looking. They don't look like coming from Volga bend. Not really.
Then combining Saami and CW results and project members. To do this I have to use my smaller data base, based on Estonian Biocentre's data. The accuracy is somewhat poorer. Numbers show the difference between Eneolithic Samara and German Corded Ware affinities in Finland and in neighboring countries, as well as results for project members. Using Eneolithic Samara and CW samples the Siberian-like admixture becomes excluded and results show only affinities common for those two groups, even if tested populations or project members have extra Siberian admixture. It is important to understand that this table alone doesn't tell how much individuals and populations have those two ancient affinities (it tells only a ratio). To see the big picture you have to take into account also two previous tables showing how significant is the relation between ancient and modern populations.
Project results
The other half of my Finnish story tells about ancient Central-European influence in Finland. Around 20% of Finnish samples from the 1000genome project show Corded-Ware similarity comparable to Estonians and Lithuanians, and Western Finnish project samples show equally Corded-Ware similarity with Swedes, some even more, despite of the fact that they are much more "eastern" when compared to present-day Swedes.
This Finnish duality doesn't tell were and when the mixing occurred and so far I have not seen any genetic evidence about the Baltic Finnic origin. It looks very possible that genetically Baltic Finns were born somewhere in region from Estonia to White Sea, no matter what the origin of Baltic Finnish language could have been.
Saami results
Saamis are genetically closer for Eneolithic Samara people than Mordovians (Mordva) and Chuvashes. Worth noticing is that Mordovians, who live near Volga are not closer those ancient people living in Samara. Saami people live thousands kilometers and thousands years away from what was the suggested Volga home range. Siberian admixture of Chuvashes roughly equals to Saami Siberian. This statistic has however very limited use, because Saami people are not Central Europeans, but still the statistic shows them being comparable to Central Europeans when compared to ancient East European samples. What could be the best outcome?
Probably some readers can think that the Eneolithic Samara - Saami - Finnish genetic connection is only based on the amount of Siberian. It is not true and easily proved false. Chuvashes and Mansi people (and Komis, not included) with high Siberian admixture are far away from the Eneolithic Samara, definitely not comparable to the Saamis. Similarly those Finns being closest Eneolithic Samara have less Siberian than Russians living in Archangel and Pinega regions in Russia (look project results).
Only people in northernmost Europe beat Saami_WGA in comparison with Eneolithic Samara. Have to admit, this is a bit complicated question. Then let's look at another perspective of supposed Finnish ancestry, Corded Ware samples. It is less complicated.
Corded Ware results
Only Lithuanians beat the Finnish CW-group (20% of Finnish samples from the 1000g project after removing outliers) when the test is done using over half million SNPs. Even Lithuanians would be beaten with more homogeneous Finnish sample group. There is all variations from very CW-looking to only moderately CW-looking. They don't look like coming from Volga bend. Not really.
Then combining Saami and CW results and project members. To do this I have to use my smaller data base, based on Estonian Biocentre's data. The accuracy is somewhat poorer. Numbers show the difference between Eneolithic Samara and German Corded Ware affinities in Finland and in neighboring countries, as well as results for project members. Using Eneolithic Samara and CW samples the Siberian-like admixture becomes excluded and results show only affinities common for those two groups, even if tested populations or project members have extra Siberian admixture. It is important to understand that this table alone doesn't tell how much individuals and populations have those two ancient affinities (it tells only a ratio). To see the big picture you have to take into account also two previous tables showing how significant is the relation between ancient and modern populations.
Project results
sunnuntai 13. maaliskuuta 2016
Continuing tests with ancient Brits, better material and final results
1. Hinxton2 is HI2 from the study "Iron Age and Anglo-Saxon genomes from East England reveal British migration history". HI2 Hinxton Male 170 BCE – 80 CE.
2. Rabrit3 is one of Roman Age samples from the study "Genomic signals of migration and continuity in Britain before the Anglo-Saxons". I can't identify which one it is of those six local samples from Driffield Terrace, because study authors don't tell connections between sample labels and sample data. Rabrit3 is processed using sample files ERR1043145, ERR1043146, ERR1043147.
3. Iabrit is M1489 from the same study (Genomic signals...). M1489 Iron Age Melton, age estimate between 210 BC and 40 AD.
4. Anglosaxon/anglosaxon2 is NO3423, again from the same study. NO3423 Anglo-Saxon Norton on Tees. Age estimate is unknown, but it is mentioned to be Anglo-Saxon.
All four samples are remastered using BWA-mem as described in my previous post. BWA-mem makes automatic trimming for reads and gives great results with minimum personal action and control, the process is fully automated. Before choosing BWA-mem I tested three additional softwares.
I have also standardized the sample selection in this test to ensure same SNP coverage for all samples and to avoid errors due to SNP qualification and differences in SNP counts. So each ancient sample is compared almost exactly similarly against modern populations. This is fundamental, because especially differences in the SNP count can cause severe biases to results.
Here are Dstat results:
result: CEU Mbuti_Pygmy iabrit Chimp.DG 0.4532 100.000 17762 6683 184450
result: CEU Mbuti_Pygmy anglosaxon Chimp.DG 0.4604 100.000 28390 10489 287312
result: French Mbuti_Pygmy iabrit Chimp.DG 0.4533 100.000 17714 6663 184450
result: French Mbuti_Pygmy anglosaxon Chimp.DG 0.4579 100.000 28268 10511 287312
result: FinnLocal Mbuti_Pygmy iabrit Chimp.DG 0.4491 100.000 17677 6721 184450
result: FinnLocal Mbuti_Pygmy anglosaxon Chimp.DG 0.4542 100.000 28218 10592 287312
result: FinnMostCW Mbuti_Pygmy iabrit Chimp.DG 0.4539 100.000 17757 6670 184450
result: FinnMostCW Mbuti_Pygmy anglosaxon Chimp.DG 0.4592 100.000 28352 10509 287312
result: IBS Mbuti_Pygmy iabrit Chimp.DG 0.4481 100.000 17602 6709 184450
result: IBS Mbuti_Pygmy anglosaxon Chimp.DG 0.4521 100.000 28066 10590 287312
result: Kent Mbuti_Pygmy iabrit Chimp.DG 0.4546 100.000 17771 6664 184450
result: Kent Mbuti_Pygmy anglosaxon Chimp.DG 0.4604 100.000 28375 10484 287312
result: Estonia Mbuti_Pygmy iabrit Chimp.DG 0.4536 100.000 17611 6620 183032
result: Estonia Mbuti_Pygmy anglosaxon Chimp.DG 0.4609 100.000 28197 10406 285211
result: Sardinian Mbuti_Pygmy iabrit Chimp.DG 0.4498 100.000 17634 6692 184450
result: Sardinian Mbuti_Pygmy anglosaxon Chimp.DG 0.4528 100.000 28105 10587 287311
result: Orcadian Mbuti_Pygmy iabrit Chimp.DG 0.4552 100.000 17766 6652 184450
result: Orcadian Mbuti_Pygmy anglosaxon Chimp.DG 0.4595 100.000 28345 10496 287311
result: TSI Mbuti_Pygmy iabrit Chimp.DG 0.4479 100.000 17597 6711 184450
result: TSI Mbuti_Pygmy anglosaxon Chimp.DG 0.4529 100.000 28076 10573 287312
result: North_Italian Mbuti_Pygmy iabrit Chimp.DG 0.4497 100.000 17653 6702 184450
result: North_Italian Mbuti_Pygmy anglosaxon Chimp.DG 0.4557 100.000 28173 10533 287311
result: Russian_Vologda Mbuti_Pygmy iabrit Chimp.DG 0.4483 100.000 17635 6718 184450
result: Russian_Vologda Mbuti_Pygmy anglosaxon Chimp.DG 0.4525 100.000 28129 10603 287311
result: CEU Mbuti_Pygmy hinxton2 Chimp.DG 0.4391 100.000 45931 17904 433006
result: CEU Mbuti_Pygmy rabrit3 Chimp.DG 0.4557 100.000 29993 11214 299676
result: French Mbuti_Pygmy hinxton2 Chimp.DG 0.4375 100.000 45806 17923 433006
result: French Mbuti_Pygmy rabrit3 Chimp.DG 0.4536 100.000 29889 11235 299676
result: FinnLocal Mbuti_Pygmy hinxton2 Chimp.DG 0.4327 100.000 45622 18064 433006
result: FinnLocal Mbuti_Pygmy rabrit3 Chimp.DG 0.4486 100.000 29784 11336 299676
result: FinnMostCW Mbuti_Pygmy hinxton2 Chimp.DG 0.4384 100.000 45896 17921 433006
result: FinnMostCW Mbuti_Pygmy rabrit3 Chimp.DG 0.4540 100.000 29933 11239 299676
result: IBS Mbuti_Pygmy hinxton2 Chimp.DG 0.4329 100.000 45492 18006 433006
result: IBS Mbuti_Pygmy rabrit3 Chimp.DG 0.4505 100.000 29732 11263 299676
result: Kent Mbuti_Pygmy hinxton2 Chimp.DG 0.4402 100.000 45988 17876 433006
result: Kent Mbuti_Pygmy rabrit3 Chimp.DG 0.4574 100.000 30041 11186 299676
result: Estonia Mbuti_Pygmy hinxton2 Chimp.DG 0.4394 100.000 45638 17775 430442
result: Estonia Mbuti_Pygmy rabrit3 Chimp.DG 0.4558 100.000 29766 11126 297499
result: Sardinian Mbuti_Pygmy hinxton2 Chimp.DG 0.4327 100.000 45516 18021 433005
result: Sardinian Mbuti_Pygmy rabrit3 Chimp.DG 0.4515 100.000 29768 11250 299675
result: Orcadian Mbuti_Pygmy hinxton2 Chimp.DG 0.4384 100.000 45892 17919 433005
result: Orcadian Mbuti_Pygmy rabrit3 Chimp.DG 0.4571 100.000 30016 11185 299675
result: TSI Mbuti_Pygmy hinxton2 Chimp.DG 0.4329 100.000 45487 18005 433006
result: TSI Mbuti_Pygmy rabrit3 Chimp.DG 0.4491 100.000 29704 11292 299676
result: North_Italian Mbuti_Pygmy hinxton2 Chimp.DG 0.4346 100.000 45645 17989 433005
result: North_Italian Mbuti_Pygmy rabrit3 Chimp.DG 0.4528 100.000 29838 11238 299675
result: Russian_Vologda Mbuti_Pygmy hinxton2 Chimp.DG 0.4337 100.000 45616 18017 433005
result: Russian_Vologda Mbuti_Pygmy rabrit3 Chimp.DG 0.4493 100.000 29754 11306 299675
And here are corresponding graphic maps:
Results differ somewhat from what I got earlier, obviously due to the stricter data preparation and more neutral outgroups.
Finally, I made also IBS-statistics using same data and a PCA-plot. It is however reasonable to state that due to the homozygosity error of ancient samples most homozygous modern populations get extra boost and give us too high results. This is typical for Balts, Irismen and Scots. I don't know about Basque homozygosity.
I was able to catch extra populations using Plink and --geno 0.01 option to standardize the SNP set as much as possible.
Creating PCA needs more samples to pick proper and all-inclusive components and is here done using another data set with less SNPs and more populations.
edit 17.3.2016 23:05
I read a comment on a Finnish history forum that using two outgroups, as I did in this post, is not recommended and can distort results. I gladly admit that this is true. But the reason for using two outgroups is very clear; I used this way to get big amount of results comparable instead of comparing only three populations. Using three target populations and one outgroup makes impossible to compare results from separate qpDstat runs, or make it at least painful. Of course the latter method, using three targets, gives better accuracy.
But no smoke here without fire, my tests using two outgroups looks reliable. In my previous results (above) the FinnMostCW group was very close to the Iron Age British sample, closer than the French sample group. I made a new test using same data, now using three target populations and one outgroup. It confirms my previous results:
0 FinnMostCW 16
1 French 28
2 iabrit 1
3 Chimp.DG 1
jackknife block size: 0.050
snps: 605676 indivs: 46
number of blocks for jackknife: 551
nrows, ncols: 46 605676
result: FinnMostCW French iabrit Chimp.DG 0.0020 1.039 9159 9122 184450
Indeed, I will have to come back to this question with larger data.
edit 18.3.2016 12:40
Here is another dpDstat result using three target populations. I am quite disappointed to the way some people react when they are not happy seeing some results. My only goal is to make objective tests using primarily European genetic data. My focus is not on Finnish results, neither I try to avoid making reliable results about Finns.
0 FinnMostCW 16
1 FinnLocal 15
2 iabrit 1
3 Chimp.DG 1
jackknife block size: 0.050
snps: 605676 indivs: 33
number of blocks for jackknife: 551
nrows, ncols: 33 605676
result: FinnMostCW FinnLocal iabrit Chimp.DG 0.0073 3.750 9120 8989 184450
lauantai 5. maaliskuuta 2016
Continuing tests with ancient Brits
Before going ahead with Roman Age samples I want to publish a PCA plot including all ancient Brits, excluding the Middle Eastern one. It look like on the main axis all Roman Age samples are very close present-day Brits and Irishmen. The Anglo-Saxon and Roman Age sample 7 are closer Swedes. All those samples turn on the second axis somewhat towards Basques. But the the Iron Age sample is clearly different, it locates just between France and England. Maybe she was from Bretagne/Brittany. I am not aware of the British history why just the Iron Age sample from Melton would look like this.
PDF
keskiviikko 24. helmikuuta 2016
Iron Age Briton and Anglo-Saxon genomes tested using dstat
After a long testing period I have now tools to process fastq-files and I can create PED and EIGENSTRAT samples from original scan results. The work flow is based on BWA and GATK, figured simply:
1. mapping fastq-files separately using BWA-mem
2. sorting and merging(samtools)
3. extracting mapped reads over certain map-quality (samtools)
4. dropping doubles (Picard tools)
5. recalculating base quality scores (GATK)
6. mapping genotypes (GATK), checking the base quality
7. updating RS-ids
8. extracting ped from vcf (vcftools)
9. converting ped to eigenstrat
Processing one sample takes on my laptop (i7/3,5Ghz/8threads used/32GB memory) 6-12 hours.
After checking all samples from the study release I was sure that I could find more information using qpDstat, which compares genetic drift rather than IBS, which was used in the original study. I considered this being possible because I have seen in my works that IBS gives often high results for unmixed and drifted populations and the result doesn't of necessity imply common ancestry in all magnitude. Mixed populations evidently become underestimated. In this meaning qpDstat beats IBS-statistics.
A few comment about results. I used Kent samples as a baseline in comparison to ancient Anglo-Saxon and Iron Age Briton, assuming that present-day Brits should be closest relatives for their ancestors. It was not true in all cases. Apparently Brits are more mixed than some other North Europeans.
I used my new Finnish grouping splitting Finns into two genetic patterns, one consisting of more local ancestry and another resembling German Corded Ware samples published last year. Both groups represent around 20% of the original 1000genomes Finnish sample set, after removing outliers. It looks like the local 1000g group differs particularly in this test from East Finnish samples, which I have gathered straight from volunteers. The difference between East and West Finland is explained more by the Iron Age British sample than the Anglo-Saxon sample. Anglo-Saxon shows high similarity with present-day Scandinavians and looks more widespread than Iron Age Briton everywhere in Northernmost Europe. Of course more Iron Age West European samples could tell more and perhaps confirm my results. Hopefully British researchers dig soon more Iron Age samples to fulfill my dreams.
I have two samples from my project members (ISX and LSX), added to figure better Finnish 1000gemone samples. Both project samples are from genealogists of Finnish speaking ancestry.
I have two databases, the smaller one holding 1 million SNPs, but only a few populations, and the larger one 0.25 million SNPs and around 3000 samples. The first one makes possible to use in this particular case around 200kSNPs, the latter gives 112368 SNPs (Anglo-Saxon) and 71478 SNPs( Iron Age sample).
edit 25.2.2016 23:25
Replacing Kent by an outgroup (Mbuti) we get absolute distances in reasonable accuracy. Closest to the Anglo-Saxon sample are
Sweden
Norway
Kent
sounds good.
And closest to the Iron Age Briton are
France
Norway
Welsh
followed by Ireland, FinnsMostCW and Kent.
edit 26.2.2016 13:40
Ranking of ancient genomes released last year by Reich Lab. Should be noted that some results are based on small amounts of SNPs. It is likely that Hungary_MBA and Germany Bronze Age get too high scores due to fewer SNPs.
The second column is for the calculated difference between Yoruba and Anglo-Saxon or Iron Age Briton, compared to the difference between ancient populations and Anglo-Saxon or Iron Age Briton. The third column is the SNP number.
Anglo-Saxon:
Hungary_MBA.SG 0.4435 23477
Remedello_BA.SG 0.4254 178071
Germany_Bronze_Age.SG 0.4209 26827
Bell_Beaker_Germany.SG 0.4137 136565
Sintashta_MBA_RISE.SG 0.4127 245264
Andronovo.SG 0.4016 285166
Corded_Ware_Estonia.SG 0.4002 154859
Bell_Beaker_Czech.SG 0.3981 190277
Iron Age Brit:
Hungary_MBA.SG 0.4531 15019
Germany_Bronze_Age.SG 0.4109 16253
Bell_Beaker_Germany.SG 0.4071 86809
Remedello_BA.SG 0.4061 114783
Sintashta_MBA_RISE.SG 0.3982 154956
Nordic_LBA.SG 0.3839 11704
Andronovo.SG 0.3764 182788
Maros.SG 0.3758 59040
1. mapping fastq-files separately using BWA-mem
2. sorting and merging(samtools)
3. extracting mapped reads over certain map-quality (samtools)
4. dropping doubles (Picard tools)
5. recalculating base quality scores (GATK)
6. mapping genotypes (GATK), checking the base quality
7. updating RS-ids
8. extracting ped from vcf (vcftools)
9. converting ped to eigenstrat
Processing one sample takes on my laptop (i7/3,5Ghz/8threads used/32GB memory) 6-12 hours.
After checking all samples from the study release I was sure that I could find more information using qpDstat, which compares genetic drift rather than IBS, which was used in the original study. I considered this being possible because I have seen in my works that IBS gives often high results for unmixed and drifted populations and the result doesn't of necessity imply common ancestry in all magnitude. Mixed populations evidently become underestimated. In this meaning qpDstat beats IBS-statistics.
A few comment about results. I used Kent samples as a baseline in comparison to ancient Anglo-Saxon and Iron Age Briton, assuming that present-day Brits should be closest relatives for their ancestors. It was not true in all cases. Apparently Brits are more mixed than some other North Europeans.
I used my new Finnish grouping splitting Finns into two genetic patterns, one consisting of more local ancestry and another resembling German Corded Ware samples published last year. Both groups represent around 20% of the original 1000genomes Finnish sample set, after removing outliers. It looks like the local 1000g group differs particularly in this test from East Finnish samples, which I have gathered straight from volunteers. The difference between East and West Finland is explained more by the Iron Age British sample than the Anglo-Saxon sample. Anglo-Saxon shows high similarity with present-day Scandinavians and looks more widespread than Iron Age Briton everywhere in Northernmost Europe. Of course more Iron Age West European samples could tell more and perhaps confirm my results. Hopefully British researchers dig soon more Iron Age samples to fulfill my dreams.
I have two samples from my project members (ISX and LSX), added to figure better Finnish 1000gemone samples. Both project samples are from genealogists of Finnish speaking ancestry.
I have two databases, the smaller one holding 1 million SNPs, but only a few populations, and the larger one 0.25 million SNPs and around 3000 samples. The first one makes possible to use in this particular case around 200kSNPs, the latter gives 112368 SNPs (Anglo-Saxon) and 71478 SNPs( Iron Age sample).
edit 25.2.2016 23:25
Replacing Kent by an outgroup (Mbuti) we get absolute distances in reasonable accuracy. Closest to the Anglo-Saxon sample are
Sweden
Norway
Kent
sounds good.
And closest to the Iron Age Briton are
France
Norway
Welsh
followed by Ireland, FinnsMostCW and Kent.
edit 26.2.2016 13:40
Ranking of ancient genomes released last year by Reich Lab. Should be noted that some results are based on small amounts of SNPs. It is likely that Hungary_MBA and Germany Bronze Age get too high scores due to fewer SNPs.
The second column is for the calculated difference between Yoruba and Anglo-Saxon or Iron Age Briton, compared to the difference between ancient populations and Anglo-Saxon or Iron Age Briton. The third column is the SNP number.
Anglo-Saxon:
Hungary_MBA.SG 0.4435 23477
Remedello_BA.SG 0.4254 178071
Germany_Bronze_Age.SG 0.4209 26827
Bell_Beaker_Germany.SG 0.4137 136565
Sintashta_MBA_RISE.SG 0.4127 245264
Andronovo.SG 0.4016 285166
Corded_Ware_Estonia.SG 0.4002 154859
Bell_Beaker_Czech.SG 0.3981 190277
Iron Age Brit:
Hungary_MBA.SG 0.4531 15019
Germany_Bronze_Age.SG 0.4109 16253
Bell_Beaker_Germany.SG 0.4071 86809
Remedello_BA.SG 0.4061 114783
Sintashta_MBA_RISE.SG 0.3982 154956
Nordic_LBA.SG 0.3839 11704
Andronovo.SG 0.3764 182788
Maros.SG 0.3758 59040
Tilaa:
Blogitekstit (Atom)
Some ancestral changes in Iron Age Estonia
QpAdm was designed to detect admixtures giving also probability and standard error statistics. Two kind of parameters are inputted: admixtu...
-
Testing ancient "steppe" samples on PCA together with modern ones revealed unexpected issues. Studies have included different set...
-
This is probably not a popular writing, but I am not a politician and I don't need to please anyone. Here is the issue: why the yDna d...
-
Unfortunately the Lamnidis et al. 2018 did not test the origin of the Siberian ancestry in present-day Finns and Russians. I reveal it now...




















