Data output¶
What the SDK returns: Pydantic models, or frames in pandas, Polars or Arrow, and how to query them with DuckDB.
Choose models or a frame¶
Two kinds of methods return collections:
- List methods like
activities.list(),traces.list(),tests.list()anddailies.list()return Pydantic models by default. - Time-series methods like
activities.data(),activities.mean_max()andactivities.longitudinal.data()always return a frame. By default it's a frame of the library you installed: Polars, then pandas, then Arrow.
Pass output= to choose per call:
from sweatstack import Client
client = Client()
latest = client.activities.latest()
activities = client.activities.list() # list[ActivitySummary]
frame = client.activities.list(output="polars") # polars.DataFrame
data = client.activities.data(latest.id, output="pandas") # pandas.DataFrame
table = client.activities.data(latest.id, output="arrow") # pyarrow.Table
raw = client.activities.data(latest.id, output="bytes") # the parquet response
output |
List methods | Time-series methods |
|---|---|---|
"models" |
yes (default) | no |
"pandas" |
yes | yes |
"polars" |
yes | yes |
"arrow" |
yes | yes |
"bytes" |
no | yes |
If you pass an output a method can't produce, the SDK raises ValueError. Account, team, status and Portal methods (users.list(), profile.status(), ...) take no output and always return models.
Set the output once¶
If your code assumes one frame library, set it once instead of per call:
import sweatstack
from sweatstack import Client
sweatstack.set_output("polars") # every client in this process
client = Client(output="pandas") # or one client
The SDK resolves output in this order: the call's output=, then Client(output=), then sweatstack.set_output(), then the method's default. If a configured default is one a method can't produce, the method uses its own default.
Know the frame shape¶
No frame has an index. timestamp, the mean-max metric and the dailies date are ordinary columns. To use pandas time-based operations, set the index yourself:
from sweatstack import Client
client = Client()
data = client.activities.data(client.activities.latest().id, output="pandas")
per_minute = data.set_index("timestamp")["power"].resample("1min").mean()
The timestamp column is timezone-aware UTC. timestamp_local is the athlete's naive local wall-clock time. See Timezones.
Dtypes depend on the library:
- pandas frames use
float64and nanosecond timestamps, so arithmetic doesn't overflow. - Polars and Arrow keep the compact types the API sends (
Int16,Float32, dictionary-encoded strings). Polars upcastsFloat16toFloat32.
From list methods, nested fields become dotted columns in pandas (summary.power.mean) and structs in Polars and Arrow:
from sweatstack import Client
client = Client()
frame = client.activities.list(output="polars")
power = frame.unnest("summary").unnest("power").select("id", "mean", "max")
Page through long lists¶
limit is how many items you get, not the API's page size. The SDK requests pages until it has limit items or there are no more. offset skips that many of the newest matching items.
from sweatstack import Client
client = Client()
season = client.activities.list(sport="cycling", limit=1000)
Query with DuckDB¶
DuckDB queries Arrow tables, Polars frames and parquet files directly. Its Python API reads in-memory frames through pyarrow, so install sweatstack[arrow] for the first two:
| Source | Code | Needs |
|---|---|---|
| Arrow table | table = ...longitudinal.data(..., output="arrow"), then duckdb.sql("select ... from table") |
sweatstack[arrow] |
| Polars frame | frame = ...longitudinal.data(..., output="polars"), then duckdb.sql("select ... from frame") |
sweatstack[polars,arrow] |
| Parquet file | Path("season.parquet").write_bytes(...longitudinal.data(..., output="bytes")), then duckdb.sql("select ... from 'season.parquet'") |
nothing extra |
from datetime import date
import duckdb
from sweatstack import Client
client = Client()
table = client.activities.longitudinal.data(
sport="cycling", start=date(2026, 1, 1), metrics=["power"], output="arrow"
)
duckdb.sql("""
select sport, count(distinct activity_id) as rides, round(avg(power)) as avg_power
from table group by sport order by rides desc
""").show()
DuckDB reads durations as INTERVAL and timestamps as TIMESTAMP WITH TIME ZONE.
Cache longitudinal queries¶
While you iterate on an analysis, cache the longitudinal responses on disk:
from datetime import date
import sweatstack
from sweatstack import Client
sweatstack.enable_cache() # or enable_cache(path="...")
client = Client()
season = client.activities.longitudinal.data(
sport="cycling", start=date(2026, 1, 1), end=date(2026, 4, 1)
)
The cache key is the query. Use fixed dates while caching: a window that moves with today adds a new entry every day. client.clear_cache() deletes the cache of the signed-in user.