References#
The entries below are grouped by the role they play in the book: the theory the models rest on, the estimators the chapters run, the data and pipelines the examples consume, and the software that implements it all. Most readers need one or two entries from a section rather than the whole section.
Reading order for each worked example#
The 13-ASV example. Start with [C5] and [C6] for why counts must be treated as compositions, then [M1] for the graphical lasso itself. The example is those two ideas composed, on a dataset small enough to check the estimator’s output by eye.
The 300-ASV Atacama example. [M1] for the estimator, [M2] for the eBIC model selection that picks λ, [M6] and [M7] for the latent/low-rank decomposition, and [M4] for selection stability. For the data see [D1], and [D2] with [D3] for how the ASVs and their taxonomy were produced.
The shotgun metagenomics example. Compositional regression replaces network estimation here: [C2] for the log-contrast formulation, [C3] and [C7] for variable selection under the zero-sum constraint, and [C1] with [C8] for the proximal and robust formulations
q2-classoimplements. The profiles come from mOTUs [S5] and the study is [D6].
Core compositional theory#
Microbiome counts carry no absolute scale: sequencing depth is arbitrary, so only ratios
between features are meaningful. These references establish that, and give the log-ratio
and log-contrast machinery that lets ordinary multivariate methods be applied to data
living on the simplex. They underpin every transformation in the book — clr, mclr, and
the zero-sum constraint in the regression chapters.
[C5] — the foundational treatment of compositional data, and the reason the 13-ASV example transforms counts before estimating anything.
[C2] — introduces log-contrast models, the direct ancestor of the constrained regression used in the shotgun chapter.
[C6] — the accessible argument for why compositionality is not optional in microbiome analysis. Read this first if the others feel abstract.
[C3], [C7] — variable selection when covariates are compositional and coefficients must sum to zero.
[C1], [C8] — general log-contrast formulations and their robust variants, as implemented by
q2-classo.
Patrick L Combettes and Christian L Müller. Regression models for compositional data: general log-contrast formulations, proximal optimization, and microbiome data applications. Statistics in Biosciences, 13(2):217–242, 2021.
John Aitchison and John Bacon-Shone. Log contrast models for experiments with mixtures. Biometrika, 71(2):323–330, 1984.
Wei Lin, Pixu Shi, Rui Feng, and Hongzhe Li. Variable selection in regression with compositional covariates. Biometrika, 101(4):785–797, 2014.
Jacob Bien, Xiaohan Yan, Léo Simpson, and Christian L Müller. Tree-aggregated predictive modeling of microbiome data. Scientific Reports, 11(1):14505, 2021.
John Aitchison. The statistical analysis of compositional data. Journal of the Royal Statistical Society: Series B (Methodological), 44(2):139–160, 1982.
Gregory B Gloor, Jean M Macklaim, Vera Pawlowsky-Glahn, and Juan J Egozcue. Microbiome datasets are compositional: and this is not optional. Frontiers in microbiology, 8:2224, 2017.
Pixu Shi, Anru Zhang, and Hongzhe Li. Regression analysis for microbiome compositional data. The Annals of Applied Statistics, 10(2):1019–1040, 2016.
Aditya Mishra and Christian L Müller. Robust regression with compositional covariates. Computational Statistics & Data Analysis, 165:107315, 2022.
James T Morton, Jon Sanders, Robert A Quinn, Daniel McDonald, Antonio Gonzalez, Yoshiki Vázquez-Baeza, Jose A Navas-Molina, Se Jin Song, Jessica L Metcalf, Embriette R Hyde, and others. Balance trees reveal microbial niche differentiation. MSystems, 2(1):e00162–16, 2017.
Huang Lin and Shyamal Das Peddada. Analysis of compositions of microbiomes with bias correction. Nature communications, 11(1):3514, 2020.
James T Morton, Clarisse Marotz, Alex Washburne, Justin Silverman, Livia S Zaramela, Anna Edlund, Karsten Zengler, and Rob Knight. Establishing microbial composition measurement standards with reference frames. Nature communications, 10(1):2719, 2019.
High-dimensional methods used in this tutorial#
With 300 features and 54 samples the sample covariance is singular, so every estimate needs structure imposed on it. These references justify the modelling choices made in the worked examples: sparsity in the inverse covariance, a low-rank term for unobserved confounders, information criteria for choosing the penalty, and stability-based control of what ends up selected. The optimisation papers matter when a fit fails to converge rather than when it succeeds.
[M1] — introduces the graphical lasso, the estimator behind both the 13-ASV and 300-ASV network examples.
[M3] — the joint graphical lasso, for estimating several related networks at once (the multiple-graphical-lasso chapter).
[M2] — extended BIC for Gaussian graphical models. This is the γ-weighted criterion that selects λ = 0.8 on the Atacama data.
[M6] — latent-variable graphical model selection, the sparse-plus-low-rank decomposition used in the SLR chapters.
[M7] — applies that decomposition to microbiome data specifically, separating association from hidden environmental structure.
[M4] — stability selection, which controls which variables survive in the high-dimensional regime.
[M8] — robust PCA, the background for the low-rank and latent-component views of the data.
[M12] — SPIEC-EASI, the compositionally-aware network inference this book’s estimator is most directly compared against.
Jerome Friedman, Trevor Hastie, and Robert Tibshirani. Sparse inverse covariance estimation with the graphical lasso. Biostatistics, 9(3):432–441, 2008.
Rina Foygel and Mathias Drton. Extended bayesian information criteria for gaussian graphical models. Advances in neural information processing systems, 2010.
Patrick Danaher, Pei Wang, and Daniela M Witten. The joint graphical lasso for inverse covariance estimation across multiple classes. Journal of the Royal Statistical Society Series B: Statistical Methodology, 76(2):373–397, 2014.
Nicolai Meinshausen and Peter Bühlmann. Stability selection. Journal of the Royal Statistical Society Series B: Statistical Methodology, 72(4):417–473, 2010.
Grace Yoon, Irina Gaynanova, and Christian L Müller. Microbial networks in spring-semi-parametric rank-based correlation and partial correlation estimation for quantitative microbiome data. Frontiers in genetics, 10:516, 2019.
Venkat Chandrasekaran, Pablo A Parrilo, and Alan S Willsky. Latent variable graphical model selection via convex optimization. In 2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), 1610–1613. IEEE, 2010.
Zachary D Kurtz, Richard Bonneau, and Christian L Müller. Disentangling microbial associations from hidden environmental and technical factors via latent graphical models. BioRxiv, pages 2019–12, 2019.
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright. Robust principal component analysis? Journal of the ACM (JACM), 58(3):1–37, 2011.
Patrick L Combettes and Christian L Müller. Perspective maximum likelihood-type estimation via proximal decomposition. ElectronicJournalofStatistics, 2020.
Patrick L Combettes and Christian L Müller. Perspective functions: proximal calculus and applications in high-dimensional statistics. Journal of Mathematical Analysis and Applications, 457(2):1283–1306, 2018.
Arthur P Dempster. Covariance selection. Biometrics, pages 157–175, 1972.
Zachary D Kurtz, Christian L Müller, Emily R Miraldi, Dan R Littman, Martin J Blaser, and Richard A Bonneau. Sparse and compositionally robust inference of microbial ecological networks. PLoS computational biology, 11(5):e1004226, 2015.
Federico Tomasi, Veronica Tozzo, Saverio Salzo, and Alessandro Verri. Latent variable time-varying network inference. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2338–2346. 2018.
Shiqian Ma, Lingzhou Xue, and Hui Zou. Alternating direction methods for latent variable gaussian graphical model selection. Neural computation, 25(8):2172–2198, 2013.
Yangjing Zhang, Ning Zhang, Defeng Sun, and Kim-Chuan Toh. A proximal point dual newton algorithm for solving group graphical lasso problems. SIAM Journal on Optimization, 30(3):2197–2220, 2020.
Ning Zhang, Yangjing Zhang, Defeng Sun, and Kim-Chuan Toh. An efficient linearly convergent regularized proximal point algorithm for fused multiple graphical lasso problems. SIAM Journal on Mathematics of Data Science, 3(2):524–543, 2021.
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, Jonathan Eckstein, and others. Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends® in Machine learning, 3(1):1–122, 2011.
Daniela M Witten, Jerome H Friedman, and Noah Simon. New insights and faster computations for the graphical lasso. Journal of Computational and Graphical Statistics, 20(4):892–900, 2011.
Laurent Condat. A direct algorithm for 1-d total variation denoising. IEEE Signal Processing Letters, 20(11):1054–1057, 2013.
Xinghao Qiao, Shaojun Guo, and Gareth M James. Functional graphical models. Journal of the American Statistical Association, 114(525):211–222, 2019.
Saharon Rosset and Ji Zhu. Piecewise linear regularized solution paths. The Annals of Statistics, pages 1012–1030, 2007.
Brian R Gaines, Juhyun Kim, and Hua Zhou. Algorithms for fitting the constrained lasso. Journal of Computational and Graphical Statistics, 27(4):861–871, 2018.
Luis Briceño-Arias and Sergio López Rivera. A projected primal–dual method for solving constrained monotone inclusions. Journal of Optimization Theory and Applications, 180:907–924, 2019.
Patrick L Combettes and Jean-Christophe Pesquet. Primal-dual splitting algorithm for solving inclusions with mixtures of composite, lipschitzian, and parallel-sum type monotone operators. Set-Valued and variational analysis, 20(2):307–330, 2012.
Cameron Martino, James T Morton, Clarisse A Marotz, Luke R Thompson, Anupriya Tripathi, Rob Knight, and Karsten Zengler. A novel sparse compositional technique reveals microbial perturbations. MSystems, 4(1):10–1128, 2019.
Yoav Freund and Robert E Schapire. A decision-theoretic generalization of on-line learning and an application to boosting. Journal of computer and system sciences, 55(1):119–139, 1997.
Pierre Geurts, Damien Ernst, and Louis Wehenkel. Extremely randomized trees. Machine learning, 63:3–42, 2006.
Jerome H Friedman. Stochastic gradient boosting. Computational statistics & data analysis, 38(4):367–378, 2002.
Leo Breiman. Random forests. Machine learning, 45:5–32, 2001.
Corinna Cortes and Vladimir Vapnik. Support-vector networks. Machine learning, 20:273–297, 1995.
Naomi S Altman. An introduction to kernel and nearest-neighbor nonparametric regression. The American Statistician, 46(3):175–185, 1992.
Hui Zou and Trevor Hastie. Regularization and variable selection via the elastic net. Journal of the Royal Statistical Society Series B: Statistical Methodology, 67(2):301–320, 2005.
Arthur E Hoerl and Robert W Kennard. Ridge regression: biased estimation for nonorthogonal problems. Technometrics, 12(1):55–67, 1970.
Robert Tibshirani. Regression shrinkage and selection via the lasso. Journal of the Royal Statistical Society Series B: Statistical Methodology, 58(1):267–288, 1996.
Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman. The elements of statistical learning: data mining, inference, and prediction. Volume 2. Springer, 2009.
Microbiome data sources and pipelines#
These provide the datasets the examples run on, and the processing context that determines what a feature actually is. Read at least the denoising and taxonomy entries: the 300-ASV table did not arrive as a matrix, and choices made upstream — denoising, reference database, taxonomy assignment — shape every downstream network.
[D1] — the Atacama soil study, and the source of the 300-ASV dataset.
[D6] — the MAP preterm-infant study, and the source of the shotgun metagenomes profiled in Shotgun Metagenomics.
[D2] — DADA2, which produces the amplicon sequence variants that are the features in the 13-ASV and 300-ASV examples.
[D3] — the SILVA reference database behind the taxonomy used to label network nodes.
[D7] — how QIIME 2 assigns that taxonomy, and why classifier choice affects what a node is called.
Julia W Neilson, Katy Califf, Cesar Cardona, Audrey Copeland, Will Van Treuren, Karen L Josephson, Rob Knight, Jack A Gilbert, Jay Quade, J Gregory Caporaso, and others. Significant impacts of increasing aridity on the arid soil microbiome. MSystems, 2(3):e00195–16, 2017.
Benjamin J Callahan, Paul J McMurdie, Michael J Rosen, Andrew W Han, Amy Jo A Johnson, and Susan P Holmes. Dada2: high-resolution sample inference from illumina amplicon data. Nature methods, 13(7):581–583, 2016.
Christian Quast, Elmar Pruesse, Pelin Yilmaz, Jan Gerken, Timmy Schweer, Pablo Yarza, Jörg Peplies, and Frank Oliver Glöckner. The silva ribosomal rna gene database project: improved data processing and web-based tools. Nucleic acids research, 41(D1):D590–D596, 2012.
Sebastian Finger, Félix A Godoy, Geraldine Wittwer, Carlos P Aranda, Raúl Calderón, and Claudio D Miranda. Purification and characterization of indochrome type blue pigment produced by Pseudarthrobacter sp. 34lch1 isolated from Atacama desert. Journal of Industrial Microbiology and Biotechnology, 46(1):101–111, 2019. doi:10.1007/s10295-018-2088-3.
Lucas Horstmann, Daniel Lipus, Alexander Bartholomäus, Romulo Oses, Axel Kitte, Thomas Friedl, and Dirk Wagner. Microbial ecology of subsurface granitic bedrock: a humid-arid site comparison in Chile. ISME Communications, 5(1):ycaf199, 2025. doi:10.1093/ismeco/ycaf199.
Shiyu S. Bai-Tong, Megan S. Thoemmes, Kelly C. Weldon, and others. The impact of maternal asthma on the preterm infants' gut metabolome and microbiome (map study). Scientific Reports, 12(1):6437, 2022. doi:10.1038/s41598-022-10276-y.
Nicholas A Bokulich, Benjamin D Kaehler, Jai Ram Rideout, Matthew Dillon, Evan Bolyen, Rob Knight, Gavin A Huttley, and J Gregory Caporaso. Optimizing taxonomic classification of marker-gene amplicon sequences with qiime 2’s q2-feature-classifier plugin. Microbiome, 6(1):1–17, 2018.
Yoshiki Vázquez-Baeza, Antonio Gonzalez, Larry Smarr, Daniel McDonald, James T Morton, Jose A Navas-Molina, and Rob Knight. Bringing the dynamic microbiome to life with animations. Cell host & microbe, 21(1):7–10, 2017.
Patrick D Schloss, Sarah L Westcott, Thomas Ryabin, Justine R Hall, Martin Hartmann, Emily B Hollister, Ryan A Lesniewski, Brian B Oakley, Donovan H Parks, Courtney J Robinson, and others. Introducing mothur: open-source, platform-independent, community-supported software for describing and comparing microbial communities. Applied and environmental microbiology, 75(23):7537–7541, 2009.
Robert C Edgar. Search and clustering orders of magnitude faster than blast. Bioinformatics, 26(19):2460–2461, 2010.
Donovan H Parks, Gene W Tyson, Philip Hugenholtz, and Robert G Beiko. Stamp: statistical analysis of taxonomic and functional profiles. Bioinformatics, 30(21):3123–3124, 2014.
Antonio Gonzalez, Jose A Navas-Molina, Tomasz Kosciolek, Daniel McDonald, Yoshiki Vázquez-Baeza, Gail Ackermann, Jeff DeReus, Stefan Janssen, Austin D Swafford, Stephanie B Orchanian, and others. Qiita: rapid, web-enabled microbiome meta-analysis. Nature methods, 15(10):796–798, 2018.
Achal Dhariwal, Jasmine Chong, Salam Habib, Irah L King, Luis B Agellon, and Jianguo Xia. Microbiomeanalyst: a web-based tool for comprehensive statistical, visual and meta-analysis of microbiome data. Nucleic acids research, 45(W1):W180–W188, 2017.
Jeff Meilander, Chloe Herman, Andrew Manley, Georgia Augustine, Dawn Birdsell, Evan Bolyen, Kimberly R Celona, Hayden Coffey, Jill Cocking, Teddy Donoghue, Alexis Draves, Daryn Erickson, Marissa Foley, Liz Gehret, Johannah Hagen, Crystal Hepp, Parker Ingram, David John, Katarina Kadar, Paul Keim, Victoria Lloyd, Christina Osterink, Victoria Monsaint-Queeney, Diego Ramirez, Antonio Romero, Megan C Ruby, Jason W Sahl, Sydni Soloway, Nathan E Stone, Shannon Trottier, Kaleb Van Orden, Alexis Painter, Sam Wallace, Larissa Wilcox, Colin V Wood, Jaiden Yancey, and J Gregory Caporaso. Upcycling human excrement: the gut microbiome to soil microbiome axis. ISME Communications, 5(1):ycaf089, 2025. doi:10.1093/ismeco/ycaf089.
J Gregory Caporaso and Jeff Meilander. Upcycling human excrement: the gut microbiome to soil microbiome axis (supporting data). 2025. Zenodo record 15390940. doi:10.5281/zenodo.15390940.
Software and QIIME 2 plugins#
These are the tools the chapters invoke. They sit on top of the previous two sections: the plugins wrap the estimators from High-dimensional methods and apply them to data prepared according to Core compositional theory.
[S2] — GGLasso, the solver underneath
q2-gglasso, used in every network fit.[S3] — c-lasso, the solver underneath
q2-classo, used for constrained log-contrast regression.[S5] — mOTUs, the marker-gene profiler that produces the species-level table in the shotgun chapter.
[S1] — QIIME 2 itself, the framework both plugins are registered against and the source of the artifact and provenance model the book relies on.
[S4] —
q2-sample-classifier, for contrasting predictive classification with the interpretable models used here.[S6] — SCNIC, an alternative compositional network tool.
[S7], [S8] — scikit-learn, whose estimators appear in supporting analyses.
Evan Bolyen, Jai Ram Rideout, Matthew R Dillon, Nicholas A Bokulich, Christian C Abnet, Gabriel A Al-Ghalith, Harriet Alexander, Eric J Alm, Manimozhiyan Arumugam, Francesco Asnicar, and others. Reproducible, interactive, scalable and extensible microbiome data science using qiime 2. Nature biotechnology, 37(8):852–857, 2019.
Fabian Schaipp, Oleg Vlasovets, and Christian L. Müller. Gglasso - a python package for general graphical lasso computation. Journal of Open Source Software, 6(68):3865, 2021. URL: https://doi.org/10.21105/joss.03865, doi:10.21105/joss.03865.
Léo Simpson, Patrick L. Combettes, and Christian L. Müller. C-lasso - a python package for constrained sparse and robust regression and classification. Journal of Open Source Software, 6(57):2844, 2021. URL: https://doi.org/10.21105/joss.02844, doi:10.21105/joss.02844.
Nicholas A Bokulich, Matthew R Dillon, Evan Bolyen, Benjamin D Kaehler, Gavin A Huttley, and J Gregory Caporaso. Q2-sample-classifier: machine-learning tools for microbiome classification and regression. Journal of open research software, 2018.
Hans-Joachim Ruscheweyh, Alessio Milanese, Lucas Paoli, Nicolai Karcher, Quentin Clayssen, Marisa Isabell Keller, Jakob Wirbel, Peer Bork, Daniel R Mende, Georg Zeller, and Shinichi Sunagawa. Cultivation-independent genomes greatly expand taxonomic-profiling capabilities of motus across various environments. Microbiome, 10(1):212, 2022. doi:10.1186/s40168-022-01410-z.
Michael Shaffer, Kumar Thurimella, John D Sterrett, and Catherine A Lozupone. Scnic: sparse correlation network investigation for compositional data. Molecular Ecology Resources, 23(1):312–325, 2023.
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, and others. Scikit-learn: machine learning in python. the Journal of machine Learning research, 12:2825–2830, 2011.
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
Everything else#
Nothing should appear below this heading. It is a safety net: the sections above select
entries by keyword, so any reference added to references.bib without a keywords
field would silently vanish from this page rather than fail the build. If entries show up
here, tag them with one of compositional, methods, data or software and they will
move into place.