Article Contents
ARTICLE   Open Access     Cite

Dix-seq: An integrated pipeline for fast amplicon data analysis

    Show all affliationsShow less
More Information
  • DownLoad: Full size image
    1. An integrated pipeline for optimized amplicon analysis has been developed.

      Dix-seq offers a one-step process with a single parameter sheet file to complete the entire pipeline.

      The modular design of Dix-seq supports custom analysis.

      Retrospective scripts make reproducing results or debugging errors easy.

  • Rapid advancements in sequencing technologies in the past decade have driven the widespread adoption of amplicon metagenome. However, current amplicon data analysis software/pipelines often require manual intervention spanning multiple steps, necessitating a clear understanding of parameters and hindering inexperienced users from automating their workflows. Here, we introduce Dix-seq, a fully containerized tool for rapid, automated, and scalable amplicon data analysis. With one single command, Dix-seq can process raw amplicon sequences down to various statistical and visualization results, generate html-based reports, and retrospective logfiles. Dix-seq utilizes a single parameter sheet file to drastically simplify its command line interface, making it much more approachable by inexperienced users while improving study reproducibility. The modular design of Dix-seq enables rapid adoption of new methods and databases into its software frame. Currently, more than 21 algorithms, software, and third-party procedures have been integrated into eight modules in Dix-seq, while more are coming down the line. This approach also allows experienced users to fine-tune the workflow, facilitating customized analysis. Benchmarks performed on datasets from real-world case studies demonstrated Dix-seq’s capabilities in generating publish-ready figures integrated with statistical information and extracting biologically meaningful patterns. Furthermore, it remained highly effective at detecting variance upon simulated sequencing depth drop, the results remained robust down to a depth of 11000 and 1000 in all and certain fronts, such as phylogenetic diversity and Pearson correlation, respectively. In summary, Dix-seq is a convenient yet highly customizable tool for amplicon data analysis, making it an ideal choice for both entry-level and experienced users.
  • 加载中
  • [1] Turnbaugh P., Ley R., Hamady M., et al. (2007). The Human Microbiome Project. Nature 449:804−810 DOI:10.1038/nature06244

    View in Article CrossRef Google Scholar Scopus

    [2] Gilbert J.A., Jansson J.K. and Knight R. (2014). The Earth Microbiome Project: Successes and aspirations. BMC Biol. 12:69 DOI:10.1186/s12915-014-0069-1

    View in Article CrossRef Google Scholar

    [3] Blaser M., Bork P., Fraser C. et al. (2013). The microbiome explored: recent insights and future challenges. Nat. Rev. Microbiol. 11:213−217 DOI:10.1038/nrmicro2973

    View in Article CrossRef Google Scholar

    [4] Ruppert K.M., Kline R.J. and Rahman M.S. (2019). Past, present, and future perspectives of environmental DNA (eDNA). metabarcoding: A systematic review in methods, monitoring, and applications of global eDNA. Glob. Ecol. Conserv. 17:e00547 DOI:10.1016/j.gecco.2019.e00547

    View in Article CrossRef Google Scholar Scopus

    [5] Zhou L. (2023). Microbial taxonomy with DNA sequence data as nomenclatural type: How far should we go. The Innovation Life 1:100017 DOI:10.59717/j.xinn-life.2023.100017

    View in Article CrossRef Google Scholar

    [6] Jin J., Liu X. and Shiroguchi K. (2024). Long journey of 16S rRNA‐amplicon sequencing toward cell‐based functional bacterial microbiota characterization. IMetaOmics 1:imo2.9 DOI:10.1002/imo2.9

    View in Article CrossRef Google Scholar

    [7] Zhang Z., Jiao N. and Zhang Y. (2024). Global ocean microbiome catalogue: A gateway to unlock marine microbial resources and ecosystem functions. The Innovation Life 2:100099 DOI:10.59717/j.xinn-life.2024.100099

    View in Article CrossRef Google Scholar

    [8] Li C., Gillings M., Zhang C., et al. (2024). Ecology and risks of the global plastisphere as a newly expanding microbial habitat. The Innovation 5:100543 DOI:10.1016/j.xinn.2023.100543

    View in Article CrossRef Google Scholar Scopus

    [9] Ewels P., Magnusson M., Lundin S., et al. (2016). MultiQC: Summarize analysis results for multiple tools and samples in a single report. Bioinformatics 32:3047−3048 DOI:10.1093/bioinformatics/btw354

    View in Article CrossRef Google Scholar Scopus

    [10] Bolger A.M., Lohse M. and Usadel B. (2014). Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 30:2114−2120 DOI:10.1093/bioinformatics/btu170

    View in Article CrossRef Google Scholar

    [11] Martin, M. (2011). Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 17:10 DOI:10.14806/ej.17.1.200

    View in Article CrossRef Google Scholar

    [12] Magoč T. and Salzberg S.L. (2011). FLASH: Fast length adjustment of short reads to improve genome assemblies. Bioinformatics 27:2957−2963 DOI:10.1093/bioinformatics/btr507

    View in Article CrossRef Google Scholar Scopus

    [13] Edgar R. (2016). UNOISE2: Improved error-correction for Illumina 16S and ITS amplicon sequencing. BioRxiv:081257. DOI:10.1101/081257

    View in Article Google Scholar

    [14] Callahan B.J., McMurdie P.J., Rosen M.J., et al. (2016). DADA2: High-resolution sample inference from Illumina amplicon data. Nat. Methods 13:581−583 DOI:10.1038/nmeth.3869

    View in Article CrossRef Google Scholar Scopus

    [15] Amir A., McDonald D., Navas-Molina J., et al. (2017). Deblur rapidly resolves single‐nucleotide community sequence patterns. MSystems 2:e00191−16 DOI:10.1128/mSystems.00191-16

    View in Article CrossRef Google Scholar

    [16] Edgar R. (2013). UPARSE: Highly accurate OTU sequences from microbial amplicon reads. Nat. Methods 10(10):996−998 DOI:10.1038/nmeth.2604

    View in Article CrossRef Google Scholar Scopus

    [17] Edgar R. (2010). Search and clustering orders of magnitude faster than BLAST. Bioinformatics 26:2460−2461 DOI:10.1093/bioinformatics/btq461

    View in Article CrossRef Google Scholar Scopus

    [18] Kopylova E., Noé L. and Touzet H. (2012). SortMeRNA: Fast and accurate filtering of ribosomal RNAs in metatranscriptomic data. Bioinformatics 28:3211−3217 DOI:10.1093/bioinformatics/bts611

    View in Article CrossRef Google Scholar Scopus

    [19] Edgar R. (2016). SINTAX: A simple non-Bayesian taxonomy classifier for 16S and ITS sequences. BioRxiv:074161. DOI:10.1101/074161

    View in Article Google Scholar

    [20] Wang Q., Garrity G.M., Tiedje J.M., et al. (2007). Naïve Bayesian classifier for rapid assignment of rRNA sequences into the new bacterial taxonomy. Appl. Environ. Microbiol. 73:5261−5267 DOI:10.1128/AEM.00062-07

    View in Article CrossRef Google Scholar

    [21] Bolyen E., Rideout J.R., Dillon M.R., et al. (2019). Reproducible, interactive, scalable and extensible microbiome data science using QIIME 2. Nat. Biotechnol. 37:852−857 DOI:10.1038/s41587-019-0209-9

    View in Article CrossRef Google Scholar Scopus

    [22] Wood D.E., Lu J. and Langmead B. (2019). Improved metagenomic analysis with Kraken 2. Genome Biol. 20:257 DOI:10.1186/s13059-019-1891-0

    View in Article CrossRef Google Scholar Scopus

    [23] Dixon P. (2003). VEGAN, a package of R functions for community ecology. J. Veg. Sci. 14:927−930 DOI:10.1111/j.1654-1103.2003.tb02228.x

    View in Article CrossRef Google Scholar Scopus

    [24] McMurdie P.J. and Holmes S. (2013). phyloseq: An R Package for Reproducible Interactive Analysis and Graphics of Microbiome Census Data. PLoS ONE 8:e61217 DOI:10.1371/journal.pone.0061217

    View in Article CrossRef Google Scholar

    [25] Lu Y., Zhou G. and Ewald J., et al. (2023). MicrobiomeAnalyst 2.0: Comprehensive statistical, functional and integrative analysis of microbiome data. Nucleic Acids Res. 51 :W310–W318. DOI:10.1093/nar/gkad407

    View in Article Google Scholar

    [26] Liu C., Cui Y., Li, X. and Yao M. (2021). Microeco: An R package for data mining in microbial community ecology. FEMS Microbiol. Ecol. 97:fiaa255 DOI:10.1093/femsec/fiaa255

    View in Article CrossRef Google Scholar Scopus

    [27] Liu B., Huang L., Liu Z., et al. (2022). EasyMicroPlot: An efficient and convenient R package in microbiome downstream analysis and visualization for clinical study. Front. Genet. 12:803627 DOI:10.3389/fgene.2021.803627

    View in Article CrossRef Google Scholar

    [28] Wen T., Xie P., Yang S., et al. (2022). ggClusterNet: An R package for microbiome network analysis and modularity‐based multiple network layouts. IMeta 1:imt2.32 DOI:10.1002/imt2.32

    View in Article CrossRef Google Scholar

    [29] Schloss P.D., Westcott S.L., Ryabin T., et al. (2009). Introducing mothur: Open-Source, Platform-Independent, Community-Supported Software for Describing and Comparing Microbial Communities. Appl. Environ. Microbiol. 75:7537−7541 DOI:10.1128/AEM.01541-09

    View in Article CrossRef Google Scholar

    [30] Rognes T., Flouri T., Nichols B., et al. (2016). VSEARCH: A versatile open source tool for metagenomics. PeerJ. 4:e2584 DOI:10.7717/peerj.2584

    View in Article CrossRef Google Scholar Scopus

    [31] Caporaso J.G., Kuczynski J., Stombaugh J., et al. (2010). QIIME allows analysis of high-throughput community sequencing data. Nat. Methods 7:335−6 DOI:10.1038/nmeth.f.303

    View in Article CrossRef Google Scholar Scopus

    [32] Comeau A.M., Douglas G.M. and Langille, M.G.I. (2017). Microbiome Helper: A custom and streamlined workflow for microbiome research. MSystems 2:e00127−16 DOI:10.1128/mSystems.00127-16

    View in Article CrossRef Google Scholar

    [33] Liu Y., Chen L., Ma T., et al. (2023). EasyAmplicon: An easy‐to‐use, open‐source, reproducible, and community‐based pipeline for amplicon data analysis in microbiome research. IMeta 2:imt2.83 DOI:10.1002/imt2.83

    View in Article CrossRef Google Scholar

    [34] Weissbecker C., Schnabel B. and Heintz-Buschart A. (2020). Dadasnake, a Snakemake implementation of DADA2 to process amplicon sequencing data for microbial ecology. Gigascience 9:giaa135 DOI:10.1093/gigascience/giaa135

    View in Article CrossRef Google Scholar

    [35] Weinstein M.M., Prem A., Jin M., et al. (2019). FIGARO: An efficient and objective tool for optimizing microbiome rRNA gene trimming parameters. BioRxiv:610394. DOI:10.1101/610394

    View in Article Google Scholar

    [36] Gonzalez A., Navas-Molina J.A., Kosciolek T., et al. (2018). Qiita: Rapid, web-enabled microbiome meta-analysis. Nat. Methods 15:796−798 DOI:10.1038/s41592-018-0141-9

    View in Article CrossRef Google Scholar

    [37] Keegan K.P., Glass E.M. and Meyer, F. (2016). MG-RAST, a metagenomics service for analysis of microbial community structure and function. Methods Mol Biol. 1399:207−233 DOI:10.1007/978-1-4939-3369-3_13

    View in Article CrossRef Google Scholar Scopus

    [38] Abueg L.A.L., Afgan E., Allart O., et al. (2024). The Galaxy platform for accessible, reproducible, and collaborative data analyses: 2024 update. Nucleic Acids Res. 52(W1):W83−W94 DOI:10.1093/nar/gkae410

    View in Article CrossRef Google Scholar Scopus

    [39] Shi W., Qi H., Sun Q., et al. (2019). gcMeta: A Global Catalogue of Metagenomics platform to support the archiving, standardization and analysis of microbiome data. Nucleic Acids Res. 47:D637−D648 DOI:10.1093/nar/gky1008

    View in Article CrossRef Google Scholar

    [40] Chen Y., Li J., Zhang Y., et al. (2022). Parallel‐Meta Suite: Interactive and rapid microbiome data analysis on multiple platforms. IMeta 1:imt2.1 DOI:10.1002/imt2.1

    View in Article CrossRef Google Scholar

    [41] Peng X. and Dorman K.S. (2021). AmpliCI: A high-resolution model-based approach for denoising Illumina amplicon data. Bioinformatics 36:5151−5158 DOI:10.1093/bioinformatics/btaa648

    View in Article CrossRef Google Scholar Scopus

    [42] Nilsen T., Snipen L.-G., Angell, I.L., et al. (2024). Swarm and UNOISE outperform DADA2 and Deblur for denoising high-diversity marine seafloor samples. ISME Communications 4:ycae071 DOI:10.1093/ismeco/ycae071

    View in Article CrossRef Google Scholar Scopus

    [43] Nearing J.T., Douglas G.M., Comeau A.M., et al. (2018). Denoising the Denoisers: An independent evaluation of microbiome sequence error-correction approaches. PeerJ 6:e5364 DOI:10.7717/peerj.5364

    View in Article CrossRef Google Scholar

    [44] Chiarello M., McCauley M., Villéger S., et al. (2022). Ranking the biases: The choice of OTUs vs. ASVs in 16S rRNA amplicon data analysis has stronger effects on diversity measures than rarefaction and OTU identity threshold. PLOS ONE 17 :e0264443. DOI:10.1371/journal.pone.0264443

    View in Article Google Scholar

    [45] Bokulich N.A., Kaehler B.D., Rideout J.R., et al. (2018). Optimizing taxonomic classification of marker-gene amplicon sequences with QIIME 2’s q2-feature-classifier plugin. Microbiome 6:90 DOI:10.1186/s40168-018-0470-z

    View in Article CrossRef Google Scholar

    [46] Edgar R.C. (2018). Accuracy of taxonomy prediction for 16S rRNA and fungal ITS sequences. PeerJ 6:e4652 DOI:10.7717/peerj.4652

    View in Article CrossRef Google Scholar Scopus

    [47] Zhu Q., Huang S., Gonzalez A., et al. (2022). Phylogeny-aware analysis of metagenome community ecology based on matched reference genomes while bypassing taxonomy. MSystems 7:e00167−22 DOI:10.1128/msystems.00167-22

    View in Article CrossRef Google Scholar

    [48] Douglas G.M., Maffei V.J., Zaneveld J.R., et al. (2020). PICRUSt2 for prediction of metagenome functions. Nat. Biotechnol. 38:685−688 DOI:10.1038/s41587-020-0548-6

    View in Article CrossRef Google Scholar Scopus

    [49] DeSantis T. Z., Hugenholtz P., Larsen N., et al. (2006). Greengenes, a chimera-checked 16S rRNA gene database and workbench compatible with ARB. Appl. Environ. Microbiol. 72:5069−5072 DOI:10.1128/AEM.03006-05

    View in Article CrossRef Google Scholar Scopus

    [50] Kõljalg U., Larsson K. and Abarenkov K. (2005). UNITE: A database providing web‐based methods for the molecular identification of ectomycorrhizal fungi. New Phytol. 166:1063−1068 DOI:10.1111/j.1469-8137.2005.01376.x

    View in Article CrossRef Google Scholar

    [51] Quast C., Pruesse E. and Yilmaz P., et al. (2012). The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 41:D590−D596 DOI:10.1093/nar/gks1219

    View in Article CrossRef Google Scholar

    [52] Escapa F.I., Huang Y., Chen T., et al. (2020). Construction of habitat-specific training sets to achieve species-level assignment in 16S rRNA gene datasets. Microbiome 8:65 DOI:10.1186/s40168-020-00841-w

    View in Article CrossRef Google Scholar Scopus

    [53] Dueholm M. K. D., Nierychlo M., Andersen K. S., et al. (2022). MiDAS 4: A global catalogue of full-length 16S rRNA gene sequences and taxonomy for studies of bacterial communities in wastewater treatment plants. Nat. Commun. 13:1908 DOI:10.1038/s41467-022-29438-7

    View in Article CrossRef Google Scholar Scopus

    [54] Wang Y., Wang K., Huang L., et al. (2020). Fine-scale succession patterns and assembly mechanisms of bacterial community of Litopenaeus vannamei larvae across the developmental cycle. Microbiome 8:106 DOI:10.1186/s40168-020-00879-w

    View in Article CrossRef Google Scholar

    [55] Love M.I., Huber W. and Anders S. (2014). Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 15:550 DOI:10.1186/s13059-014-0550-8

    View in Article CrossRef Google Scholar Scopus

    [56] Khleborodova A., Gamboa-Tuz S.D., Ramos M., et al. (2024). Lefser: Implementation of metagenomic biomarker discovery tool, LEfSe, in R. Bioinformatics 40 . DOI: 10.1093/bioinformatics/btae707

    View in Article Google Scholar

    [57] Wei Y., Jiang S., Tian L., et al. (2022). Benthic microbial biogeography along the continental shelf shaped by substrates from the Changjiang River plume. Acta Oceanol. Sin. 41:118−131 DOI:10.1007/s13131-021-1861-8

    View in Article CrossRef Google Scholar Scopus

    [58] Nguyen N.-P., Warnow T., Pop M., et al. (2016). A perspective on 16S rRNA operational taxonomic unit clustering using sequence similarity. Npj Biofilms Microbiomes 2:16004 DOI:10.1038/npjbiofilms.2016.4

    View in Article CrossRef Google Scholar Scopus

    [59] Johnson J.S., Spakowicz D.J., Hong B.-Y., et al. (2019). Evaluation of 16S rRNA gene sequencing for species and strain-level microbiome analysis. Nat. Commun. 10:5029 DOI:10.1038/s41467-019-13036-1

    View in Article CrossRef Google Scholar Scopus

    [60] Dong P., Guo H., Huang L., et al. (2023). Glucose addition improves the culture performance of Pacific white shrimp by regulating the assembly of Rhodobacteraceae taxa in gut bacterial community. Aquaculture 567:739254 DOI:10.1016/j.aquaculture.2023.739254

    View in Article CrossRef Google Scholar Scopus

    [61] Dong P., Guo H., Wang Y., et al. (2021). Gastrointestinal microbiota imbalance is triggered by the enrichment of Vibrio in subadult Litopenaeus vannamei with acute hepatopancreatic necrosis disease. Aquaculture 533:736199 DOI:10.1016/j.aquaculture.2020.736199

    View in Article CrossRef Google Scholar

  • Cite this article:

    Dong P., Chen Y., Wei Y., et al. (2025). Dix-seq: An integrated pipeline for fast amplicon data analysis. The Innovation Life 3:100120. https://doi.org/10.59717/j.xinn-life.2024.100120
    Dong P., Chen Y., Wei Y., et al. (2025). Dix-seq: An integrated pipeline for fast amplicon data analysis. The Innovation Life 3:100120. https://doi.org/10.59717/j.xinn-life.2024.100120

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(7)    

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(5834) PDF downloads(1707)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint