How Claude Code Counts Your Tokens and Money
I use Claude Code every day. I wanted one simple number: what would all of this cost if I paid API prices?
Claude Code doesn't show that number. But it writes everything to disk. So I built Cortex, a native Mac app that reads those files and adds it up.
Turns out adding it up is the hard part. If you just sum every usage field you find, you'll count almost 3x too much.
Here's what's in those files and how to count them right.
TL;DR
- Every session is a JSONL file in
~/.claude/projects/. - Only
assistantlines carry ausageobject. - One API response shows up on several lines. On one of my sessions: 388 assistant lines, only 139 real responses. Dedup by
message.id+requestId. - Subagents have their own files. Miss them and you undercount.
- Cache writes come in two prices. 5-minute writes cost 1.25x input, 1-hour writes cost 2x.
- On a single turn I checked, 95% of the cost was the cache write. Output was almost nothing.
- It's the API-equivalent cost. On a Pro or Max plan your real bill is different.
Where the data lives
Everything is under one folder:
~/.claude/projects/
<project-slug>/
<session-id>.jsonl ← the main transcript
<session-id>/subagents/*.jsonl ← subagent transcripts
The project slug is the folder path with / swapped for -. Each line in a .jsonl file is one JSON object: a user message, an assistant message, a tool result, and so on.
Some of these files get big. Cortex reads them in 1 MiB chunks so a multi-megabyte transcript never sits in memory all at once. It also only re-reads a file when its size or modified time changes.
What a usage line looks like
Only lines with "type": "assistant" carry usage. Here's one real turn from my machine, with the message content stripped out:
"usage": {
"input_tokens": 2,
"output_tokens": 356,
"cache_read_input_tokens": 25448,
"cache_creation_input_tokens": 40451,
"cache_creation": {
"ephemeral_5m_input_tokens": 0,
"ephemeral_1h_input_tokens": 40451
}
}Look at the input: 2 tokens. Everything else came from the cache. That's normal. Claude Code sends your whole conversation every turn, and most of it is cached.
The five kinds of tokens
Each kind has its own price. These are the rates Cortex uses for Opus, in USD per 1 million tokens:
| Kind | Field | Rate | vs input |
|---|---|---|---|
| Input | input_tokens | $5.00 | 1x |
| Output | output_tokens | $25.00 | 5x |
| Cache read | cache_read_input_tokens | $0.50 | 0.1x |
| Cache write, 5 min | ephemeral_5m_input_tokens | $6.25 | 1.25x |
| Cache write, 1 hour | ephemeral_1h_input_tokens | $10.00 | 2x |
The formula is just each count times its rate:
return Double(usage.input) / 1_000_000 * p.input
+ Double(usage.output) / 1_000_000 * p.output
+ Double(usage.cacheRead) / 1_000_000 * p.cacheRead
+ Double(usage.cacheWrite) / 1_000_000 * p.cacheWrite
+ Double(usage.cacheWrite1h) / 1_000_000 * p.cacheWrite1hSo that one turn above:
| Kind | Tokens | Cost |
|---|---|---|
| Input | 2 | $0.00001 |
| Output | 356 | $0.0089 |
| Cache read | 25,448 | $0.0127 |
| Cache write, 1 hour | 40,451 | $0.4045 |
| Total | $0.426 |
95% of that turn was the cache write. The answer itself, 356 output tokens, cost less than a cent.
That surprised me. I assumed output would be the expensive part. It's the price per token, sure, but you write way fewer of them.
Trap 1: one response, many lines
Claude Code streams responses. A single API response can land in the transcript as several lines, one per content block. Each line carries the same usage.
Sum them all and you double or triple count. On my newest session I found 388 assistant lines but only 139 unique responses. That's about 2.8 lines per response.
The fix: key every line by message.id plus requestId, and keep only the newest one. Token counts are cumulative, so the last line has the final numbers.
let key = line.message?.id.map { "\($0)|\(line.requestId ?? "")" }
if let key {
if let idx = pendingIndex[key] {
// Later line of the same streamed response: token
// counts are cumulative, so the newest wins.
pending[idx].usage = u
return true
}
pendingIndex[key] = pending.count
}Trap 2: the same response in two files
Resume a session and Claude Code can copy earlier lines into the new file. Now the same response lives in two transcripts.
So dedup inside one file isn't enough. Cortex hashes every (message id, request id) pair into one set across all transcripts, and each pair bills exactly once.
One detail: Swift's built-in Hasher is seeded per process, so the same key hashes differently every launch. Cortex uses FNV-1a instead, which always gives the same result. That keeps cached results valid between scans.
It also walks the files in sorted order, so the same file always claims a duplicated response. Otherwise a session's cost could jump around between refreshes.
Trap 3: subagents live somewhere else
When Claude Code spins up a subagent, that work goes into <session-id>/subagents/, not the main file.
Skip that folder and you miss real spend. I have 27 of those folders on my machine right now.
Cortex groups them under their parent session. The main transcript gives the session info (title, message count, activity by hour). The subagent files add their tokens and cost to it.
Trap 4: two prices for cache writes
Newer transcripts split cache writes by how long they're kept:
- 5-minute writes: 1.25x the input rate
- 1-hour writes: 2x the input rate
Older transcripts only have the total, cache_creation_input_tokens, with no split. Cortex treats those as 5-minute writes, since that's all that existed back then.
if let cc = mu.cache_creation {
cw5m = cc.ephemeral_5m_input_tokens ?? 0
cw1h = cc.ephemeral_1h_input_tokens ?? 0
} else {
cw5m = mu.cache_creation_input_tokens ?? 0
cw1h = 0
}The turn above was 100% 1-hour writes. Price it at the 5-minute rate and you'd undercount it by about 38%.
Trap 5: model names are messy
The same model shows up under different ids. Cortex cleans every id the same way:
- Lowercase it
- Drop the
claude-prefix - Strip a trailing date like
-20251001
So claude-opus-4-7-20251001 becomes opus-4-7. Then it looks up an exact match first, and falls back to a prefix (opus, sonnet, haiku) if there isn't one.
Two exceptions worth knowing:
- Old Opus is more expensive. Opus 4.1 and earlier were billed at the original rate, $15 in / $75 out, before the price cut. Those ids get their own tier.
<synthetic>isn't a model. It's Claude Code's id for turns it generates locally. Cortex filters it out everywhere.
And if a model has no known price (a local model, an unknown vendor), Cortex counts the tokens but prices them at $0. A wrong guess is worse than a blank.
You can override any rate in ~/.claude/readout-pricing.json.
Trap 6: transcripts disappear
Claude Code cleans up old transcripts after a while. Your history goes with them.
So Cortex keeps its own cache: token totals per day, per model, in a small JSON file. When a day's transcripts are gone, the cache fills it in. When they're still there, the live data wins because it's fresher.
It's not your bill
Everything above is the API-equivalent cost: your tokens times the published API price.
If you're on a Pro or Max subscription, you don't pay this. What it's useful for:
- Seeing which projects eat the most
- Seeing how much of your spend is cache, not output
- Knowing what it would cost if you moved a workflow to the API
That's it. Five kinds of tokens, one formula, and six ways to get it wrong.
Contact
Have a project? Say hello at
