High-Dimensional Statistics with QIIME 2#
Two plugins extend the QIIME 2 microbiome multi-omics platform [S1] with methods for data where the number of features exceeds the number of samples:
q2-gglasso estimates microbial association networks by sparse inverse covariance estimation, including the sparse-plus-low-rank and multiple-group formulations [M1, S2].
q2-classo fits sparse log-contrast regression and classification models, with taxonomic aggregation through trac [C1, S3].
Both read and write QIIME 2 artifacts, so a network or a regression carries the same provenance record as any other step in a QIIME 2 analysis.
What the plugins do#
Estimate association networks under a penalty, where the sample covariance is singular and cannot be inverted directly.
Separate direct associations from shared latent drivers, by decomposing the precision matrix into a sparse component and a low-rank one.
Fit interpretable regression and classification models on compositional data, where only ratios between features are identified.
Aggregate features along a taxonomy, so that coefficients name clades rather than individual sequence variants.
How this book is organised#
Each action is demonstrated once on a dataset small enough to check by eye, then applied to a problem where the answer is not obvious:
Part |
Data |
What it establishes |
|---|---|---|
50 samples × 13 ASVs, Atacama soil |
every q2-gglasso action, output verifiable by inspection |
|
synthetic, known ground truth |
every q2-classo action, against a known answer |
|
300 ASVs and 100 mOTUs species |
model selection when \(p > n\) and the penalty decides the model |
|
— |
every parameter, and the failure modes worth knowing |
Start with Prerequisites & Installation, or read Why High-Dimensional Statistics? for the problem the plugins address.