کد مقاله کد نشریه سال انتشار مقاله انگلیسی نسخه تمام متن
2822659 1161306 2011 7 صفحه PDF دانلود رایگان
عنوان انگلیسی مقاله ISI
BIGpre: A Quality Assessment Package for Next-Generation Sequencing Data
موضوعات مرتبط
علوم زیستی و بیوفناوری بیوشیمی، ژنتیک و زیست شناسی مولکولی ژنتیک
پیش نمایش صفحه اول مقاله
BIGpre: A Quality Assessment Package for Next-Generation Sequencing Data
چکیده انگلیسی

The emergence of next-generation sequencing (NGS) technologies has significantly improved sequencing throughput and reduced costs. However, the short read length, duplicate reads and massive volume of data make the data processing much more difficult and complicated than the first-generation sequencing technology. Although there are some software packages developed to assess the data quality, those packages either are not easily available to users or require bioinformatics skills and computer resources. Moreover, almost all the quality assessment software currently available didn’t taken into account the sequencing errors when dealing with the duplicate assessment in NGS data. Here, we present a new user-friendly quality assessment software package called BIGpre, which works for both Illumina and 454 platforms. BIGpre contains all the functions of other quality assessment software, such as the correlation between forward and reverse reads, read GC-content distribution, and base Ns quality. More importantly, BIGpre incorporates associated programs to detect and remove duplicate reads after taking sequencing errors into account and trimming low quality reads from raw data as well. BIGpre is primarily written in Perl and integrates graphical capability from the statistics package R. This package produces both tabular and graphical summaries of data quality for sequencing datasets from Illumina and 454 platforms. Processing hundreds of millions reads within minutes, this package provides immediate diagnostic information for user to manipulate sequencing data for downstream analyses. BIGpre is freely available at http://bigpre.sourceforge.net/.

ناشر
Database: Elsevier - ScienceDirect (ساینس دایرکت)
Journal: Genomics, Proteomics & Bioinformatics - Volume 9, Issue 6, December 2011, Pages 238–244
نویسندگان
, , , , , , ,