BMC Bioinformatics | |
MetaCRS: unsupervised clustering of contigs with the recursive strategy of reducing metagenomic dataset’s complexity | |
Lijun Guo1  Zhongjun Jiang1  Xiaobo Li2  | |
[1] College of Information Science and Technology, Ningbo University;College of Mathematics and Computer Science, Zhejiang Normal University; | |
关键词: Metagenomics; Unsupervised clustering; Contigs; Recursive strategy; Complexity of metagenomic samples; | |
DOI : 10.1186/s12859-021-04227-z | |
来源: DOAJ |
【 摘 要 】
Abstract Background Metagenomics technology can directly extract microbial genetic material from the environmental samples to obtain their sequencing reads, which can be further assembled into contigs through assembly tools. Clustering methods of contigs are subsequently applied to recover complete genomes from environmental samples. The main problems with current clustering methods are that they cannot recover more high-quality genes from complex environments. Firstly, there are multiple strains under the same species, resulting in assembly of chimeras. Secondly, different strains under the same species are difficult to be classified. Thirdly, it is difficult to determine the number of strains during the clustering process. Results In view of the shortcomings of current clustering methods, we propose an unsupervised clustering method which can improve the ability to recover genes from complex environments and a new method for selecting the number of sample’s strains in clustering process. The sequence composition characteristics (tetranucleotide frequency) and co-abundance are combined to train the probability model for clustering. A new recursive method that can continuously reduce the complexity of the samples is proposed to improve the ability to recover genes from complex environments. The new clustering method was tested on both simulated and real metagenomic datasets, and compared with five state-of-the-art methods including CONCOCT, Maxbin2.0, MetaBAT, MyCC and COCACOLA. In terms of the number and quality of recovered genes from metagenomic datasets, the results show that our proposed method is more effective. Conclusions A new contigs clustering method is proposed, which can recover more high-quality genes from complex environmental samples.
【 授权许可】
Unknown