کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
15151 1382 2013 6 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Multi objective SNP selection using pareto optimality
موضوعات مرتبط
مهندسی و علوم پایه مهندسی شیمی بیو مهندسی (مهندسی زیستی)
پیش نمایش صفحه اول مقاله
Multi objective SNP selection using pareto optimality
چکیده انگلیسی

Biomarker discovery is a challenging task of bioinformatics especially when targeting high dimensional problems such as SNP (single nucleotide polymorphism) datasets. Various types of feature selection methods can be applied to accomplish this task. Typically, using features versus class labels of samples in the training dataset, these methods aim at selecting feature subsets with maximal classification accuracies. Although finding such class-discriminative features is crucial, selection of relevant SNPs for maximizing other properties that exist in the nature of population genetics such as the correlation between genetic diversity and geographical distance of ethnic groups can also be equally important. In this work, a methodology using a multi objective optimization technique called Pareto Optimal is utilized for selecting SNP subsets offering both high classification accuracy and correlation between genomic and geographical distances. In this method, discriminatory power of an SNP is determined using mutual information and its contribution to the genomic–geographical correlation is estimated using its loadings on principal components. Combining these objectives, the proposed method identifies SNP subsets that can better discriminate ethnic groups than those obtained with sole mutual information and yield higher correlation than those obtained with sole principal components on the Human Genome Diversity Project (HGDP) SNP dataset.

Figure optionsDownload as PowerPoint slideHighlights
► Twelve ethnic groups selected from HGDP dataset are used to demonstrate multi objective SNP selection.
► Pareto Optimal is used for selecting SNPs offering high accuracy and correlation of genomic and geographical distances.
► Pairwise genomic distances of ethnic groups are highly correlated with their geographical distances.
► Chromosome 11 was found to be the one that possessed SNPs yielding to both high correlation and accuracy values.

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Computational Biology and Chemistry - Volume 43, April 2013, Pages 23–28
نویسندگان
, , ,