Estimated reading time at 200 wpm: 8 minutes
Imagine spending time with a free AI service and realising you are wasting your time. The answers open with flattery: “You make a very interesting point.” or “You’re one of the few people to notice that.” You feel special. Then the invented facts arrive, delivered with total confidence. You soon realise it’s nonsense. Fed up of this, you do the obvious thing and subscribe. Your $20 to $30 dollars a month buys calmer, more careful answers, and for a while that feels like progress. Two things nag though. The output still gets lots of things wrong, still hallucinates but not as badly, and needs checking even more carefully. You encounter loads of problems with creating good python scripts. It’s burning up your time and money. Your invoice tells you nothing about what you got for your money.
Whether or not you agree our Fat Disclaimer applies
There are other endings too. Perhaps the subscription months went quiet and $25 evaporated on capacity you never touched. Or in another scenario the credits ran out halfway through a session. Work stops. You are invited to top up. You groan.
If that sounds like you, or someone you know, stick around. There is a metered route that answers both problems: one account, one API key, one bill, a price attached to every call, and a basket of over 400 models to chose from. Several of them are free or almost free. It’s something called OpenRouter. It’s pretty easy to setup from your computer. See: Accessing DeepSeek V4-Pro via API on Windows 11 Using OpenWebUI – The Captain’s Watch.
The account behind this post was opened on 26 August, and its first ten days are on record: roughly 540 generations, just over half on free variants, for $2.28 in total. The one heavy session in there, a conversation that grew past 315,000 tokens, was carried by a reasoning model for $1.96. Pushed through a flagship model at published rates, that session alone would have cost about $85. However, in an efficient system, even a month with two or three sessions of that size stays under $10.
Every figure comes from two OpenRouter activity exports covering 26 August to 5 September 2026. The cost numbers are measured.
1. What ten days cost – broken down
The window runs from the evening of 26 August 2026, when the account was opened, to the morning of 5 September: nine and a half days, rounded up to ten. Two OpenRouter activity exports cover the period, and every figure in this post is drawn from a CSV export of real data. The account logged 538 generations. Fourteen of the four hundred plus models were actually used, served by fourteen different providers, and the total bill came to $2.28, an average of nine tenths of a cent per billed call.
Table 1: billed spend by model, 26 August to 5 September 2026.
| Model | Calls | Spend | Share |
|---|---|---|---|
| DeepSeek V4 Pro (StreamLake) | 55 | $1.96 | 86% |
| GLM 4.7 Flash (four providers) | 133 | $0.13 | 6% |
| Qwen 3.7 Flash (Alibaba) | 42 | $0.08 | 4% |
| GLM 5.3 Flash (Z.AI, GMICloud) | 4 | $0.06 | 3% |
| Seven one-off models | 21 | $0.04 | 2% |
| Free variants (Ling 3.0 Flash, Nemotron 3) | 283 | $0.00 | — |
| Total | 538 | $2.28 |
Column values are rounded to the nearest cent; the total is computed from unrounded figures.
The $0.00 rows do the quiet work. Free variants carried 283 of the 538 calls, a little over half the workload. Ling 3.0 Flash on Novita took the bulk of it, including the small companion calls that appear seconds before or after nearly every paid generation, most plausibly the interface’s own title and follow-up tasks routed to a free model. Nemotron 3 on Nvidia covered the rest.
But none of this means necessarily means changing ‘chats’ or discussion, in a multilayer setup.
Concentration is the other feature of the table. DeepSeek V4 Pro accounts for 86% of everything billed, from just 55 calls, all of them inside a single overnight conversation that the next section examines. GLM 4.7 Flash, spread across four providers, averaged about a tenth of a cent per call over 133 runs. The seven one-off models added four cents between them, including three calls made from OpenRouter’s Chatroom app rather than the usual pipeline.
Two mechanisms made this arithmetic possible. Prompt caching kept the DeepSeek bill low, and the free tier kept more than half the traffic off the meter entirely. Caching comes next.
2. The caching arithmetic
The overnight conversation was one continuous thread: 55 DeepSeek V4 Pro calls between 23:32 on 4 September and 02:08 on 5 September, its prompt growing from 5,554 tokens to 315,723. In all it pushed 7.2 million prompt tokens and about 261,000 output tokens, and billed $1.96.
What keeps a thread like that affordable is prompt caching. Every call resends the whole conversation, so the provider is asked to re-read millions of tokens. But when the opening is recognised as a repeat it bills at a steep discount, and only the new material pays full price. Two calls from the session, nine minutes apart, show the difference:
| Call | Prompt tokens | Served from cache | Billed |
|---|---|---|---|
| 01:52, cache intact | 281,384 | 272,384 (97%) | $0.14 |
| 01:43, cache gone | 265,290 | none | $0.31 |
Read the rows as one lesson. Same model, same provider, prompts within 6% of each other, and the second billed more than twice the first because its opening had lost the cache. OpenRouter logged a $0.35 cache credit on the healthy call; the discount is already inside the billed figure.
Across the session the credits stacked to about $7.90 against $1.96 billed. Uncached, the same 55 calls would have come to roughly $9.85, so caching cut the bill by four fifths. Hit rates began under a third in the early minutes and settled above 97% once the thread was established.
The empty row is the standing risk. The very next call found 97% of its prompt cached again, so the miss reads as a brief eviction at the provider rather than an edit to the conversation, but one unlucky call cost 31 cents. Latency tracked the cache too: a 160,000-token prompt with 77% cached sat sixteen seconds before its first token arrived, while a fully cached 200,000-token prompt answered in three or four. Section 3 prices the same work at flagship list rates, where no such discount applies.
3. The flagship virtual comparison
GPT-6 Astra Pro lists at $10 per million input tokens and $50 per million output tokens, with a context window of 1M tokens, and it went live the same day this thread started. Repricing the overnight session at those rates:
| Component | Tokens | Rate | Cost |
|---|---|---|---|
| Input | 7,177,630 | $10/M | $71.78 |
| Output | 261,050 | $50/M | $13.05 |
| Session total | $84.83 |
Against the $1.96 actually billed for the same 55 calls, that is roughly 43 times. Per turn in the loop it is about $1.54, against 3.6 cents. The output column carries its own warning: the thread was reasoning-heavy, and reasoning bills as output, at five times the input rate.
The 1M context window does not soften the bill. The loop’s call count is set by the tools, not by the window; each of the 55 turns needed its own fresh answer, and each turn carried the thread so far at full price. The rate sheet shows no cached-input rate either, so nothing equivalent to section two’s discount applies. A rough tally across the whole period puts total volume near 33 million tokens; at these rates the same ten days would have billed around $360, against $2.28 actually paid. That is the arithmetic behind the title’s promise of hundreds of dollars of value.
The above figures are rough estimates. Astra is a said to be more token efficient than previous models. Estimating token usage (a measure of compute) is not a straight linear equation in practice because of several other factors. I don’t plan to use Astra just to get an idea of token usage. I don’t want to pay for that experiment.
4. Qualifying considerations
Measured versus impression. Every dollar figure in this post comes from the two activity exports; the claim that flattery and invented facts arrive far less often through API models rests on use rather than tally, and nothing in those files can verify it. Anyone weighing the switch should keep the two kinds of claim apart.
A first fortnight: The window starts at account creation, and a single overnight session carries 86% of everything billed. Ten days make this a profile snapshot; a forecast needs a longer window. The ten-dollars-a-month figure sits in the middle: a month with two or three sessions like the overnight one stays under $10, and a quiet month lands closer to $2.
The free half: Free variants carried 283 of the 538 calls, and free endpoints can be withdrawn, rate-limited or renamed without notice. Nothing in a metered account protects against that; a paid fallback configured in the pipeline is what keeps the pattern workable when one of them disappears.
Routing economics: The same GLM 4.7 Flash model ran through four providers with first-token times from a fifth of a second to nearly fifteen seconds, and five Llama 4 Maverick calls through DigitalOcean returned between twelve and sixty tokens apiece for just under a cent in total. A monthly invoice would show one total.
Conclusion
The question at the top was whether metered access could beat a flat subscription at this scale of use. Ten days of records answer it in numbers: 538 generations for $2.28, just over half the calls on free variants, and the overnight reasoning session at $1.96 against about $85 for the same work at flagship list prices. The habit that follows is small: a monthly glance at the frontier line item, and the comparison re-run whenever the mix of work changes. The number to beat, for now, is $2.28.
The reality is that most people will not need high-powered flagship AI models regularly. What they need is something that is good enough for their needs. That can be had for almost free! The choice is yours.









