| Proteomes | |
| A Preliminary Metagenome Analysis Based on a Combination of Protein Domains | |
| Yoji Igarashi1  Daisuke Mori1  Kazutoshi Yoshitake1  Susumu Mitsuyama1  Shuichi Asakawa1  Yoshizumi Ishino2  Hiroaki Ono3  Takashi Gojobori4  Takanori Kobayashi5  Shugo Watabe6  Yukiko Taniuchi7  Tomoko Sakami7  Tsuyoshi Watanabe7  Akira Kuwata7  | |
| [1] Department of Aquatic Bioscience, Graduate School of Agricultural and Life Sciences, The University of Tokyo, Bunkyo, Tokyo 113-8657, Japan;Graduate School of Bioresorce and Bioenvironmental Sciences, Kyushu University, Fukuoka, Fukuoka 812-0053, Japan;Japan Software Management Co, Ltd., Yokohama, Kanagawa 221-0056, Japan;King Abdullah University of Science and Technology, Thuwal 23955, Saudi Arabia;National Research Institute of Fisheries Science, Japan Fisheries Research and Education Agency, Yokohama, Kanagawa 236-8648, Japan;School of Marine Biosciences, Kitasato University, Sagamihara, Kanagawa 252-0373, Japan;Tohoku National Fisheries Research Institute, Japan Fisheries Research and Education Agency, Shiogama, Miyagi 985-0001, Japan; | |
| 关键词: protein domain; correlation coefficient; phylogenetic analysis; metagenomics; environmental DNA; | |
| DOI : 10.3390/proteomes7020019 | |
| 来源: DOAJ | |
【 摘 要 】
Metagenomic data have mainly been addressed by showing the composition of organisms based on a small part of a well-examined genomic sequence, such as ribosomal RNA genes and mitochondrial DNAs. On the contrary, whole metagenomic data obtained by the shotgun sequence method have not often been fully analyzed through a homology search because the genomic data in databases for living organisms on earth are insufficient. In order to complement the results obtained through homology-search-based methods with shotgun metagenomes data, we focused on the composition of protein domains deduced from the sequences of genomes and metagenomes, and we utilized them in characterizing genomes and metagenomes, respectively. First, we compared the relationships based on similarities in the protein domain composition with the relationships based on sequence similarities. We searched for protein domains of 325 bacterial species produced using the Pfam database. Next, the correlation coefficients of protein domain compositions between every pair of bacteria were examined. Every pairwise genetic distance was also calculated from 16S rRNA or DNA gyrase subunit B. We compared the results of these methods and found a moderate correlation between them. Essentially, the same results were obtained when we used partial random 100 bp DNA sequences of the bacterial genomes, which simulated raw sequence data obtained from short-read next-generation sequences. Then, we applied the method for analyzing the actual environmental data obtained by shotgun sequencing. We found that the transition of the microbial phase occurred because the seasonal change in water temperature was shown by the method. These results showed the usability of the method in characterizing metagenomic data based on protein domain compositions.
【 授权许可】
Unknown