SampleCollection

Data export, analysis and visualization functions are all contained within the SampleCollection class.

A SampleCollection is returned whenever multiple Samples are returned via the One Codex API using a model.

Usage

SampleCollection contains useful tools for data export, analysis, visualization and statistics. See the following sections for more information:

import onecodex

ocx = onecodex.Api()

project = ocx.Project.get("d53ad03b010542e3")
samples = ocx.Samples.where(project=project)

type(samples) # SampleCollection

A SampleCollection can also be created manually from a list of samples:

from onecodex.models.collection import SampleCollection

sample_list = [
    ocx.Samples.get("cee3b512605a43c6"),
    ocx.Samples.get("01f703ac505e4a30")
]

samples = SampleCollection(sample_list)

# convert classification results to a Pandas DataFrame
samples.to_df()

filter

SampleCollection.filter(filter_func: Callable[[Samples | Classifications], object]) → Self

Return a new SampleCollection containing only samples meeting the filter criteria.

Will pass any kwargs (e.g., metric or skip_missing) used when instantiating the current class on to the new SampleCollection that is returned.

Parameters

filter_funccallable

A function that will be evaluated on every object in the collection. Objects are kept when it returns a truthy value and dropped when it returns a falsy one, so it can return a field directly (e.g. lambda s: s.tags) rather than a bool.

Returns

onecodex.models.SampleCollection containing only the objects filter_func returned a truthy value for.

Examples

Keep only the samples whose filename ends in .fastq.gz:

samples = ocx.Samples.all(limit=10)
filtered = samples.filter(lambda s: s.filename.endswith('.fastq.gz'))

Note: Where fields support server-side filtering (such as filename), use Samples.where directly. The parameters of Samples.where and Classifications.where list every field that can be filtered server-side. See Querying for the query syntax:

filtered = ocx.Samples.where(filename={'$endswith': '.fastq.gz'})

Custom metadata cannot be filtered using where, so filter is necessary:

samples = ocx.Samples.where(project=project)
filtered = samples.filter(lambda s: s.metadata.custom.get('subject') == '123')

As with filename, where can do this for you, and much faster — it avoids downloading a results table for every sample:

filtered = ocx.Samples.where(tax_ids=['357276'])

where and filter can be chained, which is useful for narrowing down a large collection on the server first, then filtering the smaller result on custom metadata:

from datetime import datetime, timedelta, timezone

last_month = (datetime.now(timezone.utc) - timedelta(days=30)).isoformat()
filtered = ocx.Samples.where(updated_at={'$gte': last_month}).filter(
    lambda s: s.metadata.custom.get('subject_id') == '123'
)

to_otu

SampleCollection.to_otu(biom_id: str | None = None, include_ranks: tuple[str] = ('superkingdom', 'kingdom', 'phylum', 'class', 'order', 'family', 'genus', 'species'), metric: Metric = auto)

Generate a BIOM-formatted data structure.

Parameters

biom_idstring, optional

Optionally specify an id field for the generated v1 BIOM file.

include_rankslist

A list of ranks to include in the taxonomy/OTU table. Uses onecodex.models.collection.CANONICAL_RANKS by default.

Returns

otu_tableOrderedDict

A BIOM OTU table, returned as a Python OrderedDict (can be dumped to JSON)

to_df

SampleCollection.to_df(analysis_type: str | AnalysisType = 'classification', **kwargs) → pd.DataFrame

Transform Analyses of samples in a SampleCollection into a tabular DataFrame.

Parameters

analysis_type{‘classification’, ‘functional’}, default=’classification’

The type of analysis to aggregate.

**kwargs

Keyword arguments passed to the specific aggregation method.

Common Arguments:

  • metric (str | Metric): The metric to aggregate (default: Metric.Auto).

  • fill_missing (bool): Whether to fill missing values (default: True).

  • filler (Any): Value to use for filling missing values (default: 0).

If analysis_type=’classification’:

  • rank (Rank | str): Taxonomic rank to aggregate at (default: Rank.Auto).

  • top_n (int, optional): Return only the top N taxa by abundance.

  • threshold (float, optional): Filter taxa below this abundance threshold.

  • remove_zeros (bool): Remove taxa with zero abundance (default: True).

  • include_host (bool): Include host reads in the output (default: False).

  • table_format ({‘wide’, ‘long’}): The shape of the output DataFrame.

  • include_taxa_missing_rank (bool): Include taxa unspecified at the target rank.

If analysis_type=’functional’:

  • annotation (str | FunctionalAnnotations): The functional annotation database (default: Pathways).

  • taxa_stratified (bool): Whether to include taxonomic stratification (default: True).

Returns

pd.DataFrame

A DataFrame containing the aggregated classification or functional results.

See Also

to_classification_df : Underlying method for classification extraction. to_functional_df : Underlying method for functional extraction.

to_classification_df

SampleCollection.to_classification_df(rank: Rank | str = auto, top_n: int | None = None, threshold: float | None = None, remove_zeros: bool = True, include_host: bool = False, table_format: Literal['wide', 'long'] = 'wide', include_taxa_missing_rank: bool = False, fill_missing: bool = True, filler: Any = 0, metric: Metric | str = auto)

Generate a ClassificationsDataFrame, performing any specified transformations.

Collates the classification results of the samples in this collection into a single table of metric values, applies the requested filtering, and returns a new ClassificationsDataFrame.

Parameters

rank{Rank, str}, optional

Analysis will be restricted to abundances of taxa at the specified level. Defaults to auto, which uses species for abundance metrics and for metagenomic (shotgun) analyses, and genus otherwise. See Rank for details.

top_ninteger, optional

Return only the N taxa with the highest metric values, summed across all samples.

thresholdfloat, optional

Return only taxa whose metric value is at or above this threshold in one or more samples.

remove_zerosbool, optional

Do not return taxa that have zero abundance in every sample. Defaults to True.

include_hostbool, optional

Include host reads in the analysis. Defaults to False.

table_format{‘wide’, ‘long’}, optional

If wide (the default), rows are classifications, cols are taxa, and elements are metric values. If long, rows are observations with three cols each: classification_id, tax_id, and a column named after the metric’s display name (e.g. “Relative Abundance”) holding the value.

include_taxa_missing_rankbool, optional

Whether or not to include taxa that do not have a designated parent at rank (will be grouped into a “No <rank>” column). Defaults to False.

fill_missingbool, optional

Fill np.nan values. Defaults to True.

fillerfloat, optional

Value with which to fill np.nans. Defaults to 0.

metric{Metric, str}, optional

The taxonomic abundance metric to use. Defaults to auto, which uses abundance_w_children when the samples have abundance estimates and normalized_readcount_w_children otherwise. See Metric for definitions.

Returns

ClassificationsDataFrame

to_functional_df

SampleCollection.to_functional_df(annotation: FunctionalAnnotations = pathways, taxa_stratified: bool = True, metric: FunctionalAnnotationsMetric = coverage, fill_missing: bool = True, filler: Any = 0)

Generate a FunctionalDataFrame associated with functional analysis results.

Functional profiles are listed along the rows and functional annotations along the columns. When taxa_stratified=True, the columns are a MultiIndex of (feature_id, taxon_id). Otherwise, the columns are an index of feature_id.

Parameters

annotation:class:onecodex.lib.enum.FunctionalAnnotations, str}, optional

Annotation data to return, defaults to pathways

taxa_stratifiedbool, optional

Return taxonomically stratified data, defaults to True

metric{onecodex.lib.enum.FunctionalAnnotationsMetric, str}, optional

Metric values to return {‘coverage’, ‘abundance’} for annotation==FunctionalAnnotations.Pathways or {‘rpk’, ‘cpm’} for other annotations, defaults to coverage

fill_missingbool, optional

Fill np.nan values

fillerfloat, optional

Value with which to fill np.nans