Metrics
How to Compare AI Usage Without Misleading Deltas
A defensible AI-usage comparison starts with two explicit slices, checks metric coverage and provenance, and keeps refused deltas visible instead of silently dropping them.
By Novus Stream Solutions Editorial Team. Published 2026-08-20. Last reviewed 2026-08-20. 7 min read.
A delta looks precise because it has a sign. +18% feels more decisive than two totals shown side by side, even when the totals came from different providers, different export versions, or different amounts of missing data. That confidence is exactly why comparison needs stricter rules than a dashboard total.
AI Stats comparison mode treats a comparison as two independently defined slices of imported history. It will subtract a metric only when the two slices make the subtraction defensible. When they do not, the row remains in the report and explains why it is not comparable. The blank is not a broken calculation. It is the calculation refusing to overstate what the exports know.
This guide explains the live workflow on /app/dashboard/compare. For the field-by-field contract, keep Compare views and save the ones you reuse open beside it.
A slice is the question you asked of the imported sessions. Each side has six filter fields: range, custom start, custom end, providers, projects, and a model substring. Side A and side B therefore produce twelve URL parameters, prefixed with a and b. That URL is the complete comparison definition; it contains no prompt, transcript, session title, or imported file.
The safest entry point is the Compare link on the dashboard. It copies the view already on screen into side A. Side B remains independent so you can deliberately choose the contrast. Copying A into both sides would produce a page full of zero differences and teach you nothing.
Useful questions are narrow enough to name:
- This week versus the previous equivalent week, with the same provider filter.
- Codex sessions in one project versus Codex sessions in another project.
- One model substring before and after a workflow change, where the export actually names models.
- The same saved filter view reopened after a later import, when you want a current answer rather than a historical snapshot.
Avoid starting with “everything versus everything.” Consumer chat exports and coding-agent logs do not report the same fields. A mixed-provider comparison can still reveal composition, but many token, cost, tool, and duration rows should refuse a numeric delta.
Date ranges use when sessions started, not when you imported their files. Uploading a six-month export today does not move six months of sessions into the rolling 30-day window. This distinction matters whenever all-time looks complete and a short range looks unexpectedly small.
Before comparing, open each side as an ordinary dashboard view and confirm that its range contains the sessions you intended. “Rolling 30 days” means the local calendar window ending now. “All time” includes the imported history regardless of when the import completed. A custom range is appropriate when your work period does not match a rolling or calendar preset.
If the dates are wrong, a mathematically correct delta answers the wrong question.
Every metric has a coverage fraction: how many sessions in the slice actually reported that field. A tool-call total over 30 of 100 sessions may be a correct total for those 30 sessions. It is not evidence that the other 70 sessions made no tool calls.
Missingness is usually structural. ChatGPT web exports do not suddenly begin reporting token counts because they sit beside Codex sessions that do. Claude web can expose tool blocks in some export shapes while other sessions do not carry them. Model names may be present for one provider and absent for another. Those gaps follow the source format, so they are not a random sample that can safely be scaled up.
That is why comparison is stricter than display. A dashboard may show a partial total with its coverage beside it. A delta between two partial totals would invite you to interpret a change in export composition as a change in behaviour.
For a numeric difference, both sides must pass the same comparability gates:
- Sessions exist on both sides. An empty slice has no baseline.
- Both sides have a value. Unavailable is not zero and cannot be subtracted as though it were.
- Coverage is complete on both sides. Every contributing session in each slice reported the metric.
- Measurement signatures match. Adapter version, normalization version, quality grade, source, and calculation must agree.
The fourth gate protects against a quieter error. Two values may both be present and fully covered but still mean different things. Source-reported duration and a derived active span are not interchangeable simply because both format as minutes. A parser upgrade can also change what an adapter recovers from the same family of exports. Comparison must not erase that provenance.
AI Stats sends both sides through the same comparison policy used by the dashboard’s period-over-period narrative. There is one decision point for comparability, not a looser rule for one screen and a stricter rule for another.
Comparison rows are never filtered just because they cannot print a difference. Each remains in place with its reason:
- No sessions means the filters removed the baseline or result set.
- No value usually points to a capability-matrix limitation in the export.
- Partial coverage means at least one session did not report the metric.
- Different signatures means the two sides were measured differently.
Removing those rows would create a polished but misleading report. You would see only the easy comparisons and lose the evidence that token coverage disappeared, a provider mix changed, or one side was normalized by a newer method.
Composition tables follow the same honesty rule in a different form. Providers, models, and project groups are aligned across both sides. If a key exists on only one side, the other cell reads Not in this view. A new provider appearing only in side B may explain several blocked metric rows more clearly than any percentage could.
Some comparisons need identical provider and project filters. Others deliberately compare providers or projects. Decide which before reading the result.
For a time comparison, keep the non-time filters the same. If A is Codex-only and B includes Claude web, the result mixes a change in time with a change in source capability. For a provider comparison, keep the date and project filters aligned, then expect provider-specific metrics to refuse when the formats differ.
Model filters deserve particular care. A model substring cannot recover a model name an export never contained. Filtering one side to a named model while the other provider exposes no models creates two different populations, not an even contest.
A saved view stores a name plus the six canonical dashboard filter parameters. It does not store the totals that were on screen, a screenshot, or a session list. Reopen it after another import and AI Stats recomputes the slice from current normalized data.
That behavior is useful for recurring questions such as “Codex, this repository, rolling 30 days.” It is not an audit snapshot. If you need to preserve figures for publication, use the product’s explicit sharing or export workflow and include coverage and provenance with the number.
Saved filter values are allow-listed, de-duplicated, and ordered canonically. The name helps you find the question again; the URL remains the inspectable definition of the view.
Before quoting a delta:
- Open each side and confirm the intended sessions are present.
- Confirm the range refers to session start time, not import time.
- Read coverage for the metric on both sides.
- Check whether the providers actually export that metric.
- Keep provider, project, and model filters symmetrical unless changing one is the question.
- Read every refused row and the composition table before focusing on the allowed rows.
- Describe the two slices beside the number so another reader can reconstruct the comparison.
Then keep the conclusion at the altitude of the measurement. More sessions means more recorded sessions, not more value. More prompts means more exported prompt events, not better prompting. API-equivalent cost is an estimate at published API rates where the required tokens and models exist; it is not a subscription bill.
The strongest comparison is often the one that produces fewer numbers. A report that refuses three tempting deltas and permits one well-supported difference has done more analytical work than a table that subtracts everything.
For a guided click path, continue with Compare two dashboard views. For the source-specific gaps behind a refusal, use the capability matrix and /methodology.
Related reading
Metrics
The Capability Matrix: What Each AI Export Can Actually Tell You
A source-by-source reading of the AI Stats capability matrix — Verified versus Beta, Not in export versus Not parsed yet, and why a blank Claude or Gemini tile is evidence rather than a bug.
Metrics
Twelve Totals and Five Quality Labels: What the Dashboard Measures
Ranges, timezone boundaries, previous-period comparison, unavailable versus zero, and the rules that make a cost estimate refuse.
Metrics
AI Active Time vs. Session Time: What Your Statistics Actually Mean
Why an AI session lasting three hours does not necessarily mean the model worked for three hours.
Was this page helpful?
Your answer stays in this browser. It is not sent anywhere, and no account or cookie is involved.