Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2020 Apr 2;11(1):1628.
doi: 10.1038/s41467-020-15464-w.

Ancestry deconvolution and partial polygenic score can improve susceptibility predictions in recently admixed individuals

Affiliations

Ancestry deconvolution and partial polygenic score can improve susceptibility predictions in recently admixed individuals

Davide Marnetto et al. Nat Commun. .

Abstract

Polygenic Scores (PSs) describe the genetic component of an individual's quantitative phenotype or their susceptibility to diseases with a genetic basis. Currently, PSs rely on population-dependent contributions of many associated alleles, with limited applicability to understudied populations and recently admixed individuals. Here we introduce a combination of local ancestry deconvolution and partial PS computation to account for the population-specific nature of the association signals in individuals with admixed ancestry. We demonstrate partial PS to be a proxy for the total PS and that a portion of the genome is enough to improve susceptibility predictions for the traits we test. By combining partial PSs from different populations, we are able to improve trait predictability in admixed individuals with some European ancestry. These results may extend the applicability of PSs to subjects with a complex history of admixture, where current methods cannot be applied.

PubMed Disclaimer

Conflict of interest statement

The authors declare no competing interests.

Figures

Fig. 1
Fig. 1. Schematic workflow.
A graphical representation of the workflow we adopted to obtain normalized PS and ancestry specific pPS. White boxes represent input data, the two key steps of ancestry deconvolution and partial PS computation have an orange background.
Fig. 2
Fig. 2. Population-wide Polygenic Scores (PS) and ancestry specific partial PS.
PS distributions for seven reference populations (pastel colors), three admixed populations (yellow) and their relative ancestry specific partial PS (red and blue). Reference population medians are represented with dashed lines. The width of the boxplots is proportional to the median size of the ancestry fraction used to compute each aspPS. Four different PS for different phenotypes are shown: (a) T2D, (b) breast cancer (c) height, (d) BMI. Significant differences with randomly assigned ancestral components are encoded as: *: p ≤ 0.05, **: p ≤ 0.005, ***: p ≤ 10−5, (one-sided Wilcoxon signed-rank test). Sample sizes and exact P-values are reported in Supplementary Data 1. For each distribution, the box represent the interquartile range (IQR = Q3Q1), the line across the box indicate the median, the whiskers extend to the most extreme data points within Q1−1.5IQR and Q3 + 1.5IQR, outliers are omitted. CEU: North-West Europeans from Utah; IBS: Iberians from Spain; TSI: Tuscans from Italy; CHB: Han from Beijing; YRI: Yoruba from Nigeria; LWK: Luhya from Kenya; GUMUZ: Gumuz from Ethiopia; EGYPT: Egyptians; ETHIOPIA: Amhara, Oromo, Wolayta and Ethiopian Somali from Ethiopia; ASW: African-Americans from South-West USA.
Fig. 3
Fig. 3. pPS predictivity.
We plugged in four trait prediction models pPS obtained with genomic subsets of variable sizes (the same resulting from local ancestry analysis in our admixed individuals), in a non-admixed sample set derived from EstBB. Each point represents the performance of a different subset of the genome, applied to all individuals in the population; on the horizontal axis is reported the fraction of genomic SNPs included in each subset, while the vertical axis represents its predictivity, expressed in R2 (for binary traits we used Nagelkerke R2). Red dots represent pPS not significantly improving the base model without PS (p > 0.05, likelihood ratio test). The dashed line represents the total PS predictivity. (a) Type 2 diabetes, (b) breast cancer, (c) height, (d) BMI.
Fig. 4
Fig. 4. Population-wide PS and aspPS in UKBB admixed individuals.
PS distributions for four reference populations (pastel colors), three admixed populations (yellow) and their relative ancestry specific partial PS (red, blue, green). Reference population medians are represented with dashed lines. The width of the boxplots is proportional to the median size of the ancestry fraction used to compute each aspPS. Two different PS are shown: (a) height and (d) BMI. Significant differences with randomly assigned ancestral components are encoded as: *: p ≤ 0.05, **: p ≤ 0.005, ***: p ≤ 10−5 (one-sided Wilcoxon signed-rank test). Sample sizes and exact P-values are reported in Supplementary Data 1. For each distribution, the box represent the interquartile range (IQR = Q3Q1), the line across the box indicate the median, the whiskers extend to the most extreme data points within Q1−1.5IQR and Q3 + 1.5IQR, outliers are omitted. (c) PS bias, defined as mean PS difference not explained by trait difference, is compared with FST against the reference population. All populations extracted from UKBB are represented, showing for UK EURAFR the fraction of European ancestry, UK EUR is the reference population for UKBB-based PSs, while UK EAS is the reference for BBJ-based PSs. EUR indicates european descent, EAS east asian descent, AFR african descent; combinations indicate admixed samples. FAREUR indicates Europeans far from the UKBB core.
Fig. 5
Fig. 5. Predictivity in admixed genomes.
Each plot shows the improvement in R2 when adding a PS to the base, non-genetic model. The line color depicts which PS configuration has been used: traditional total PS (PSUKBB or PSBBJ according to the Biobank of origin), partial ancestry specific PSs or combined ancestry specific PS. Dots represent the realized R2 improvement in each set without resampling, while bars represent standard deviation derived from n=5000 bootstrap replications. a Added R2 for height in UKBB samples with admixed African and European ancestry, no casPS was available. b Added R2 for height in UKBB samples with admixed East Asian and European ancestry. c Added R2 for BMI in UKBB samples with admixed African and European ancestry; no casPS was available. d Added R2 for BMI in UKBB samples with admixed East Asian and European ancestry. EUR indicates european descent, EAS east asian descent, AFR african descent; combinations indicate admixed samples.

References

    1. Wray NR, Goddard ME, Visscher PM. Prediction of individual genetic risk to disease from genome-wide association studies. Genome Res. 2007;17:1520–8. doi: 10.1101/gr.6665407. - DOI - PMC - PubMed
    1. Wray NR, et al. Research review: polygenic methods and their application to psychiatric traits. J. Child Psychol. Psyc. 2014;55:1068–1087. doi: 10.1111/jcpp.12295. - DOI - PubMed
    1. Visscher PM, et al. 10 years of GWAS discovery: biology, function, and translation. Am. J. Hum. Genet. 2017;101:5–22. doi: 10.1016/j.ajhg.2017.06.005. - DOI - PMC - PubMed
    1. Buniello A, et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 2019;47:D1005–D1012. doi: 10.1093/nar/gky1120. - DOI - PMC - PubMed
    1. Mahajan A, et al. Genome-wide trans-ancestry meta-analysis provides insight into the genetic architecture of type 2 diabetes susceptibility. Nat. Genet. 2014;46:234–244. doi: 10.1038/ng.2897. - DOI - PMC - PubMed

Publication types