کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
15143 1381 2014 7 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
Subgrouping Automata: Automatic sequence subgrouping using phylogenetic tree-based optimum subgrouping algorithm
ترجمه فارسی عنوان
زیرگروه اتوماتیک: زیرگروه سازی اتوماتیک دنباله با استفاده از الگوریتم زیرگروه بهینه سازی مبتنی بر درختی فیلوژنتیک
کلمات کلیدی
زیرگروه شدن، تبعیض نژادی پروتئین، گره زیرزمینی بهینه درخت فیلوژنتیک، تحلیل آماری
موضوعات مرتبط
مهندسی و علوم پایه مهندسی شیمی بیو مهندسی (مهندسی زیستی)
چکیده انگلیسی


• An algorithm is developed for automatic sequence subgrouping for given sequence set.
• The algorithm basically utilizes phylogenetic tree from multiple sequence alignment.
• The algorithm calculates all of pairwise sequence identities and statistical analysis.
• The algorithm automatically determines optimum subgrouping node in phylogenetic tree.
• The algorithm showed good performance for family- or superfamily-level sequence set.

Sequence subgrouping for a given sequence set can enable various informative tasks such as the functional discrimination of sequence subsets and the functional inference of unknown sequences. Because an identity threshold for sequence subgrouping may vary according to the given sequence set, it is highly desirable to construct a robust subgrouping algorithm which automatically identifies an optimal identity threshold and generates subgroups for a given sequence set. To meet this end, an automatic sequence subgrouping method, named ‘Subgrouping Automata’ was constructed. Firstly, tree analysis module analyzes the structure of tree and calculates the all possible subgroups in each node. Sequence similarity analysis module calculates average sequence similarity for all subgroups in each node. Representative sequence generation module finds a representative sequence using profile analysis and self-scoring for each subgroup. For all nodes, average sequence similarities are calculated and ‘Subgrouping Automata’ searches a node showing statistically maximum sequence similarity increase using Student's t-value. A node showing the maximum t-value, which gives the most significant differences in average sequence similarity between two adjacent nodes, is determined as an optimum subgrouping node in the phylogenetic tree. Further analysis showed that the optimum subgrouping node from SA prevents under-subgrouping and over-subgrouping.

Figure optionsDownload as PowerPoint slide

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Computational Biology and Chemistry - Volume 48, February 2014, Pages 64–70
نویسندگان
, , , , , , ,