Skip to content

Data output

What the SDK returns: Pydantic models, or frames in pandas, Polars or Arrow, and how to query them with DuckDB.

Choose models or a frame

Two kinds of methods return collections:

  • List methods like activities.list(), traces.list(), tests.list() and dailies.list() return Pydantic models by default.
  • Time-series methods like activities.data(), activities.mean_max() and activities.longitudinal.data() always return a frame. By default it's a frame of the library you installed: Polars, then pandas, then Arrow.

Pass output= to choose per call:

from sweatstack import Client

client = Client()
latest = client.activities.latest()

activities = client.activities.list()                       # list[ActivitySummary]
frame = client.activities.list(output="polars")             # polars.DataFrame
data = client.activities.data(latest.id, output="pandas")   # pandas.DataFrame
table = client.activities.data(latest.id, output="arrow")   # pyarrow.Table
raw = client.activities.data(latest.id, output="bytes")     # the parquet response
output List methods Time-series methods
"models" yes (default) no
"pandas" yes yes
"polars" yes yes
"arrow" yes yes
"bytes" no yes

If you pass an output a method can't produce, the SDK raises ValueError. Account, team, status and Portal methods (users.list(), profile.status(), ...) take no output and always return models.

Set the output once

If your code assumes one frame library, set it once instead of per call:

import sweatstack
from sweatstack import Client

sweatstack.set_output("polars")          # every client in this process
client = Client(output="pandas")         # or one client

The SDK resolves output in this order: the call's output=, then Client(output=), then sweatstack.set_output(), then the method's default. If a configured default is one a method can't produce, the method uses its own default.

Know the frame shape

No frame has an index. timestamp, the mean-max metric and the dailies date are ordinary columns. To use pandas time-based operations, set the index yourself:

from sweatstack import Client

client = Client()
data = client.activities.data(client.activities.latest().id, output="pandas")
per_minute = data.set_index("timestamp")["power"].resample("1min").mean()

The timestamp column is timezone-aware UTC. timestamp_local is the athlete's naive local wall-clock time. See Timezones.

Dtypes depend on the library:

  • pandas frames use float64 and nanosecond timestamps, so arithmetic doesn't overflow.
  • Polars and Arrow keep the compact types the API sends (Int16, Float32, dictionary-encoded strings). Polars upcasts Float16 to Float32.

From list methods, nested fields become dotted columns in pandas (summary.power.mean) and structs in Polars and Arrow:

from sweatstack import Client

client = Client()
frame = client.activities.list(output="polars")
power = frame.unnest("summary").unnest("power").select("id", "mean", "max")

Page through long lists

limit is how many items you get, not the API's page size. The SDK requests pages until it has limit items or there are no more. offset skips that many of the newest matching items.

from sweatstack import Client

client = Client()
season = client.activities.list(sport="cycling", limit=1000)

Query with DuckDB

DuckDB queries Arrow tables, Polars frames and parquet files directly. Its Python API reads in-memory frames through pyarrow, so install sweatstack[arrow] for the first two:

Source Code Needs
Arrow table table = ...longitudinal.data(..., output="arrow"), then duckdb.sql("select ... from table") sweatstack[arrow]
Polars frame frame = ...longitudinal.data(..., output="polars"), then duckdb.sql("select ... from frame") sweatstack[polars,arrow]
Parquet file Path("season.parquet").write_bytes(...longitudinal.data(..., output="bytes")), then duckdb.sql("select ... from 'season.parquet'") nothing extra
from datetime import date

import duckdb
from sweatstack import Client

client = Client()
table = client.activities.longitudinal.data(
    sport="cycling", start=date(2026, 1, 1), metrics=["power"], output="arrow"
)
duckdb.sql("""
    select sport, count(distinct activity_id) as rides, round(avg(power)) as avg_power
    from table group by sport order by rides desc
""").show()

DuckDB reads durations as INTERVAL and timestamps as TIMESTAMP WITH TIME ZONE.

Cache longitudinal queries

While you iterate on an analysis, cache the longitudinal responses on disk:

from datetime import date

import sweatstack
from sweatstack import Client

sweatstack.enable_cache()        # or enable_cache(path="...")
client = Client()
season = client.activities.longitudinal.data(
    sport="cycling", start=date(2026, 1, 1), end=date(2026, 4, 1)
)

The cache key is the query. Use fixed dates while caching: a window that moves with today adds a new entry every day. client.clear_cache() deletes the cache of the signed-in user.