Coming soon! The Kael'Nyrin Scrolls: The Atlas Edict

Trophy and fireworks celebrating DeepSeek victory

Captain Walker

AI Usage: What My Tokens Actually Cost

AI, API, cost, deepseek, input, models, OpenWebUI, output

Estimated reading time at 200 wpm: 4 minutes

I’ve been running most of my AI work through DeepSeek V4 Pro’s API for a while now. On heavy days it’s about three large documents attached, a long conversation, drafted outputs at the end. It felt cheap. I wanted to know whether it actually was, or whether I’d just got used to a low number without checking it against anything.

Whether or not you agree our Fat Disclaimer applies

So I pulled the data and did the sums properly. What I found surprised me enough to write it down and share out here.

The starting numbers

Thirty days of usage on DeepSeek V4 Pro:

MetricValue
Total cost$17.81
API requests2,272
Total tokens83,788,070

On the face of it, that’s a lot of tokens for not much money. The question was why, and whether other platforms would cost roughly the same for the same work.

The actual data

I pulled a week’s worth of daily breakdowns, cache hits and misses included. The following where the heaviest in a month:

DateTotal tokensInput (cache hit)Input (cache miss)Output
2026-07-037,226,2954,641,5362,505,03479,725
2026-07-128,811,3924,841,4723,857,617112,303
2026-07-143,342,2911,657,6001,587,37697,315
2026-07-188,743,2824,737,7923,803,161202,329
2026-07-1913,416,8975,676,4167,587,996152,485
2026-07-208,410,5843,785,6004,503,944121,040

Add it up and the pattern is stark:

  • Input tokens: roughly 98.5% of the total
  • Output tokens: roughly 1.5% of the total
  • Within input, cache hits and cache misses split close to 51/49

I checked the remaining days across the 28 June to 27 July period. They ran at about 40% of these same ratios. The pattern holds across the month, not just a handful of days.

I read three large documents into most sessions, mostly Markdown, some Word, some PDF. The model reads a great deal and generates loads of text – but comparatively little in terms of how data and computing is counted. That single fact turns out to matter more than anything else in this analysis.

Why the split matters so much

Every platform charges far more for output tokens than input tokens. Typically five to six times more per token. Initially I thought half my tokens were the expensive kind. In reality, close to none are.

The 84 million tokens was with the real ~98.5/1.5 split, no caching applied yet. But $37 in the table below is not what I paid. That is the estimate. My bill was about half of that.

Comparison of costs across AI Models
ModelInput / Output rate (per 1M tokens)estimate
DeepSeek V4 Flash$0.14 / $0.28~$12
DeepSeek V4 Pro$0.435 / $0.87~$37
GPT-5.6 Luna$1.00 / $6.00~$90
Gemini 3.6 Flash$1.50 / $7.50~$134
Gemini 3.1 Pro$2.00 / $12.00~$181
Claude Sonnet 5 (intro rate)$2.00 / $10.00~$178
Claude Sonnet 5 (standard)$3.00 / $15.00~$267
GPT-5.6 Terra$2.50 / $15.00~$226
Claude Opus 4.8$5.00 / $25.00~$446
GPT-5.5$5.00 / $30.00~$452

My working pattern happens to sit exactly where input-cheap pricing pays off, and where output-heavy pricing barely bites at all.

Then there’s caching

The revised DeepSeek V4 Pro estimate above, around $37, still doesn’t match my actual bill of $17.81. The remaining gap is caching.

DeepSeek prices a cache hit at $0.003625 per million tokens, against $0.435 for a cache miss on the same content. That’s well over a hundred times cheaper. Because I run long sessions on the approximately three documents at a time, turn after turn, roughly half my input tokens land as cache hits rather than fresh, full-price reads.

Caching rewards exactly the workflow I have: the same source material, revisited repeatedly, across a sustained conversation. Each new message in a long chat resends the accumulated context. Without caching, that would compound cost with every turn. With it, the repeated portion becomes almost free.

What this actually means

Putting it together, three things happened at once, and they all point the same way:

  1. My workload is overwhelmingly input-heavy, and input is the cheap side of every provider’s pricing.
  2. Long sessions on the same documents are precisely what prompt caching is built to reward.
  3. DeepSeek V4 discounts cache hits far more steeply than any of the closed flagships do.

The result is that my actual spend sits below even the corrected no-caching estimate. Switching to a flagship model wouldn’t just cost more per token. The gap would be worse than the headline rates suggest, because I’m not generating enough output to dilute the expensive side of their pricing, and none of them reward repeated context the way DeepSeek does.

None of this is a verdict on quality. DeepSeek V4 Pro trails the frontier models a little on deep reasoning and factual recall, and matches or beats them on coding benchmarks. That’s a separate question from cost. But for the specific shape of my work, document-heavy, long-running, output-light, the pricing model and my usage pattern line up about as well as they could.

The lesson

If you’re paying for any AI API by the token, the question worth asking isn’t “what’s the headline rate?” It’s “what shape is my actual usage?” – input versus output, and whether your pattern lets caching do any work for you. Those two things, more than the rate card, decide what you actually pay.