Skip to content

Analyze activity data

SweatStack stores normalized timeseries, so an analysis script can query weeks or months of activity in a single call, without building an ingestion pipeline first. This guide covers two patterns: a one-off Python script for ad-hoc exploration, and a longitudinal query for trend analysis across many activities.

What you'll build

  • A standalone Python script that fetches longitudinal cycling data and computes power and torque profiles
  • An understanding of when to use a per-activity query and when to use a longitudinal query

Setup

The script uses uv, so it declares its dependencies inline. No requirements.txt, no virtualenv to manage.

Create analysis.py with this header:

analysis.py
# /// script
# requires-python = ">=3.10"
# dependencies = [
#     "sweatstack[pandas]",
#     "numpy",
#     "matplotlib",
# ]
# ///

from datetime import date, timedelta

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
from sweatstack import Client

client = Client()

client.authenticate()

sweatstack.authenticate() opens the browser once to sign in. The SDK caches the credentials for later runs.

Run the script with uv run analysis.py. The first run installs the dependencies. Later runs are fast.

Longitudinal queries: many activities at once

For trends across many activities, a longitudinal query returns a single timeseries DataFrame spanning the requested window:

data = client.activities.longitudinal.data(
    sport=["cycling.road"],
    start=date.today() - timedelta(days=90),
    metrics=["power", "cadence"],
)

print(f"{data['activity_id'].nunique()} activities, {len(data)} samples")

The timestamp column is timezone-aware UTC, so standard pandas resampling works once you make it the index:

weekly_max_power = data.set_index("timestamp")["power"].resample("W").max()

That resamples on UTC week boundaries. To group by the athlete's local calendar (day, week), use the naive local timestamp_local column instead. See Timezones.

The default metrics are duration, power, heart_rate, and speed. Pass metrics=[...] to request any other metric from the data model.

Activity metadata: per-activity summary data

When you don't need the timeseries, only the per-activity summary (sport, distance, duration, start time), use activities.list():

activities = client.activities.list(
    sport=["cycling.road"],
    start=date.today() - timedelta(days=90),
    output="pandas",
)

The API returns JSON, and with output="pandas" the SDK turns it into a DataFrame. Smaller payloads, faster queries. Use this when your analysis is at the activity level, not at the sample level.

Putting longitudinal queries to work: power and torque profiles

A worked example with a longitudinal query: compare the power-cadence relationship of road and trainer cycling.

analysis.py
# ... setup and client.authenticate() above ...

road = client.activities.longitudinal.data(
    sport=["cycling.road"],
    start=date.today() - timedelta(days=90),
    metrics=["power", "cadence"],
)

trainer = client.activities.longitudinal.data(
    sport=["cycling+stationary"],
    start=date.today() - timedelta(days=90),
    metrics=["power", "cadence"],
)

def torque(power, cadence):
    angular_velocity = cadence * (2 * np.pi / 60)
    t = pd.Series(index=power.index, dtype="float64")
    mask = angular_velocity > 0
    t[mask] = power[mask] / angular_velocity[mask]
    return t

road["torque"] = torque(road["power"], road["cadence"])
trainer["torque"] = torque(trainer["power"], trainer["cadence"])

fig, (ax_power, ax_torque) = plt.subplots(2, 1, figsize=(10, 10))

for df, color, label in [
    (road, "red", "Road"),
    (trainer, "blue", "Trainer"),
]:
    ax_power.scatter(df["cadence"], df["power"], color=color, marker=".", alpha=0.5, label=label)
    ax_torque.scatter(df["cadence"], df["torque"], color=color, marker=".", alpha=0.5, label=label)

ax_power.set(xlabel="cadence [rpm]", ylabel="power [W]")
ax_torque.set(xlabel="cadence [rpm]", ylabel="torque [Nm]")
ax_power.legend()
ax_torque.legend()

plt.show()

Cycling torque profile

Tip: cache during iteration

While you iterate on the analysis itself, calling the API on every run is wasteful. Enable the SDK's local cache, and fix the start and end dates so the cache keys stay the same across runs:

import sweatstack
from sweatstack import Client

client = Client()

sweatstack.enable_cache()  # (1)!

data = client.activities.longitudinal.data(
    sport=["cycling.road"],
    start=date(2026, 1, 1),
    end=date(2026, 4, 1),  # fixed dates: the cache key stays the same
    metrics=["power", "cadence"],
)
  1. Caches longitudinal responses on disk, in the platform's cache directory. Pass path="..." for a different directory. sweatstack.clear_cache() empties it.

Avoid date.today() while caching. The window shifts each day, and every run adds a new cache entry.

Next steps

  • The data model covers every metric you can query, and how activities relate to traces, tests, and dailies.
  • For analysis in an app (a web UI, deployed somewhere), see the Streamlit app and FastAPI app guides.
  • For analysis of lab data (lactate tests, VO2max), see Tests and Metabolic profile.