BMC Bioinformatics | |
Light-weight reference-based compression of FASTQ data | |
Methodology Article | |
Yanli Yang1  Zexuan Zhu1  Yongpeng Zhang1  Linsen Li1  Shan He2  Xiao Yang3  | |
[1] College of Computer Science and Software Engineering, Shenzhen University, 518060, Shenzhen, China;School of Computer Science, University of Birmingham, B15 2TT, Birmingham, UK;The Broad Institute, 02142, Cambridge, MA, USA; | |
关键词: Quality Score; Compression Ratio; Encode Scheme; Next Generation Sequencing Data; FASTQ Format; | |
DOI : 10.1186/s12859-015-0628-7 | |
received in 2015-03-09, accepted in 2015-05-27, 发布年份 2015 | |
来源: Springer | |
【 摘 要 】
BackgroundThe exponential growth of next generation sequencing (NGS) data has posed big challenges to data storage, management and archive. Data compression is one of the effective solutions, where reference-based compression strategies can typically achieve superior compression ratios compared to the ones not relying on any reference.ResultsThis paper presents a lossless light-weight reference-based compression algorithm namely LW-FQZip to compress FASTQ data. The three components of any given input, i.e., metadata, short reads and quality score strings, are first parsed into three data streams in which the redundancy information are identified and eliminated independently. Particularly, well-designed incremental and run-length-limited encoding schemes are utilized to compress the metadata and quality score streams, respectively. To handle the short reads, LW-FQZip uses a novel light-weight mapping model to fast map them against external reference sequence(s) and produce concise alignment results for storage. The three processed data streams are then packed together with some general purpose compression algorithms like LZMA. LW-FQZip was evaluated on eight real-world NGS data sets and achieved compression ratios in the range of 0.111-0.201. This is comparable or superior to other state-of-the-art lossless NGS data compression algorithms.ConclusionsLW-FQZip is a program that enables efficient lossless FASTQ data compression. It contributes to the state of art applications for NGS data storage and transmission. LW-FQZip is freely available online at: http://csse.szu.edu.cn/staff/zhuzx/LWFQZip.
【 授权许可】
Unknown
© Zhang et al. 2015. This is an Open Access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly credited. The Creative Commons Public Domain Dedication waiver (http://creativecommons.org/publicdomain/zero/1.0/) applies to the data made available in this article, unless otherwise stated.
【 预 览 】
Files | Size | Format | View |
---|---|---|---|
RO202311102203155ZK.pdf | 962KB | download |
【 参考文献 】
- [1]
- [2]
- [3]
- [4]
- [5]
- [6]
- [7]
- [8]
- [9]
- [10]
- [11]
- [12]
- [13]
- [14]
- [15]
- [16]
- [17]
- [18]
- [19]
- [20]
- [21]
- [22]
- [23]
- [24]
- [25]
- [26]
- [27]
- [28]
- [29]
- [30]
- [31]