SampleCollection
Data export, analysis and visualization functions are all contained within the
SampleCollection class.
A SampleCollection is returned whenever multiple Samples are returned via
the One Codex API using a model.
Usage
SampleCollection contains useful tools for data export, analysis,
visualization and statistics. See the following sections for more information:
import onecodex
ocx = onecodex.Api()
project = ocx.Project.get("d53ad03b010542e3")
samples = ocx.Samples.where(project=project)
type(samples) # SampleCollection
A SampleCollection can also be created manually from a list of samples:
from onecodex.models.collection import SampleCollection
sample_list = [
ocx.Samples.get("cee3b512605a43c6"),
ocx.Samples.get("01f703ac505e4a30")
]
samples = SampleCollection(sample_list)
# convert classification results to a Pandas DataFrame
samples.to_df()
filter
- SampleCollection.filter(filter_func: Callable[[Samples | Classifications], object]) Self
Return a new SampleCollection containing only samples meeting the filter criteria.
Will pass any kwargs (e.g., metric or skip_missing) used when instantiating the current class on to the new SampleCollection that is returned.
Parameters
- filter_funccallable
A function that will be evaluated on every object in the collection. Objects are kept when it returns a truthy value and dropped when it returns a falsy one, so it can return a field directly (e.g.
lambda s: s.tags) rather than a bool.
Returns
onecodex.models.SampleCollection containing only the objects filter_func returned a truthy value for.
Examples
Keep only the samples whose filename ends in
.fastq.gz:samples = ocx.Samples.all(limit=10) filtered = samples.filter(lambda s: s.filename.endswith('.fastq.gz'))
Note: Where fields support server-side filtering (such as
filename), useSamples.wheredirectly. The parameters ofSamples.whereandClassifications.wherelist every field that can be filtered server-side. See Querying for the query syntax:filtered = ocx.Samples.where(filename={'$endswith': '.fastq.gz'})
Custom metadata cannot be filtered using
where, sofilteris necessary:samples = ocx.Samples.where(project=project) filtered = samples.filter(lambda s: s.metadata.custom.get('subject') == '123')
As with
filename,wherecan do this for you, and much faster — it avoids downloading a results table for every sample:filtered = ocx.Samples.where(tax_ids=['357276'])
whereandfiltercan be chained, which is useful for narrowing down a large collection on the server first, then filtering the smaller result on custom metadata:from datetime import datetime, timedelta, timezone last_month = (datetime.now(timezone.utc) - timedelta(days=30)).isoformat() filtered = ocx.Samples.where(updated_at={'$gte': last_month}).filter( lambda s: s.metadata.custom.get('subject_id') == '123' )
to_otu
- SampleCollection.to_otu(biom_id: str | None = None, include_ranks: tuple[str] = ('superkingdom', 'kingdom', 'phylum', 'class', 'order', 'family', 'genus', 'species'), metric: Metric = auto)
Generate a BIOM-formatted data structure.
Parameters
- biom_idstring, optional
Optionally specify an id field for the generated v1 BIOM file.
- include_rankslist
A list of ranks to include in the taxonomy/OTU table. Uses onecodex.models.collection.CANONICAL_RANKS by default.
Returns
- otu_tableOrderedDict
A BIOM OTU table, returned as a Python OrderedDict (can be dumped to JSON)
to_df
- SampleCollection.to_df(analysis_type: str | AnalysisType = 'classification', **kwargs) pd.DataFrame
Transform Analyses of samples in a
SampleCollectioninto a tabular DataFrame.Parameters
- analysis_type{‘classification’, ‘functional’}, default=’classification’
The type of analysis to aggregate.
- **kwargs
Keyword arguments passed to the specific aggregation method.
Common Arguments:
metric (str | Metric): The metric to aggregate (default: Metric.Auto).
fill_missing (bool): Whether to fill missing values (default: True).
filler (Any): Value to use for filling missing values (default: 0).
If analysis_type=’classification’:
rank (Rank | str): Taxonomic rank to aggregate at (default: Rank.Auto).
top_n (int, optional): Return only the top N taxa by abundance.
threshold (float, optional): Filter taxa below this abundance threshold.
remove_zeros (bool): Remove taxa with zero abundance (default: True).
include_host (bool): Include host reads in the output (default: False).
table_format ({‘wide’, ‘long’}): The shape of the output DataFrame.
include_taxa_missing_rank (bool): Include taxa unspecified at the target rank.
If analysis_type=’functional’:
annotation (str | FunctionalAnnotations): The functional annotation database (default: Pathways).
taxa_stratified (bool): Whether to include taxonomic stratification (default: True).
Returns
- pd.DataFrame
A DataFrame containing the aggregated classification or functional results.
See Also
to_classification_df : Underlying method for classification extraction. to_functional_df : Underlying method for functional extraction.
to_classification_df
- SampleCollection.to_classification_df(rank: Rank | str = auto, top_n: int | None = None, threshold: float | None = None, remove_zeros: bool = True, include_host: bool = False, table_format: Literal['wide', 'long'] = 'wide', include_taxa_missing_rank: bool = False, fill_missing: bool = True, filler: Any = 0, metric: Metric | str = auto)
Generate a ClassificationsDataFrame, performing any specified transformations.
Collates the classification results of the samples in this collection into a single table of metric values, applies the requested filtering, and returns a new ClassificationsDataFrame.
Parameters
- rank{
Rank, str}, optional Analysis will be restricted to abundances of taxa at the specified level. Defaults to auto, which uses species for abundance metrics and for metagenomic (shotgun) analyses, and genus otherwise. See
Rankfor details.- top_ninteger, optional
Return only the N taxa with the highest metric values, summed across all samples.
- thresholdfloat, optional
Return only taxa whose metric value is at or above this threshold in one or more samples.
- remove_zerosbool, optional
Do not return taxa that have zero abundance in every sample. Defaults to True.
- include_hostbool, optional
Include host reads in the analysis. Defaults to False.
- table_format{‘wide’, ‘long’}, optional
If wide (the default), rows are classifications, cols are taxa, and elements are metric values. If long, rows are observations with three cols each: classification_id, tax_id, and a column named after the metric’s display name (e.g. “Relative Abundance”) holding the value.
- include_taxa_missing_rankbool, optional
Whether or not to include taxa that do not have a designated parent at rank (will be grouped into a “No <rank>” column). Defaults to False.
- fill_missingbool, optional
Fill np.nan values. Defaults to True.
- fillerfloat, optional
Value with which to fill np.nans. Defaults to 0.
- metric{
Metric, str}, optional The taxonomic abundance metric to use. Defaults to auto, which uses abundance_w_children when the samples have abundance estimates and normalized_readcount_w_children otherwise. See
Metricfor definitions.
Returns
ClassificationsDataFrame- rank{
to_functional_df
- SampleCollection.to_functional_df(annotation: FunctionalAnnotations = pathways, taxa_stratified: bool = True, metric: FunctionalAnnotationsMetric = coverage, fill_missing: bool = True, filler: Any = 0)
Generate a FunctionalDataFrame associated with functional analysis results.
Functional profiles are listed along the rows and functional annotations along the columns. When taxa_stratified=True, the columns are a MultiIndex of (feature_id, taxon_id). Otherwise, the columns are an index of feature_id.
Parameters
- annotation:class:onecodex.lib.enum.FunctionalAnnotations, str}, optional
Annotation data to return, defaults to pathways
- taxa_stratifiedbool, optional
Return taxonomically stratified data, defaults to True
- metric{onecodex.lib.enum.FunctionalAnnotationsMetric, str}, optional
Metric values to return {‘coverage’, ‘abundance’} for annotation==FunctionalAnnotations.Pathways or {‘rpk’, ‘cpm’} for other annotations, defaults to coverage
- fill_missingbool, optional
Fill np.nan values
- fillerfloat, optional
Value with which to fill np.nans