> ## Content Index
> Fetch the complete content index at: https://operators-ai.ghost.io/llms.txt
> Use this file to discover other available public pages before exploring further.

# The month the subscriptions stopped covering it
- URL: https://operators-ai.ghost.io/the-month-the-subscriptions-stopped-covering-it/
- Published: 2026-09-16T22:39:29.000Z
- Updated: 2026-09-16T22:39:29.000Z
- Description: In April I published a cost guide with ninety-two dollar figures in it. Most of them are wrong now, and the thing that drains operators was not in it at all. The plan names did not change. The contract underneath them did.
- Author: Justin Zsimovan
- Tags: lessons, tier-2

In April I wrote a Systems-tier guide for this community called "How Much Does This Actually Cost to Run." It had two pricing models in it, subscription and API, a table of per-token prices, five monthly budget scenarios, and a warning about the two traps. I reread it this week. The two-model frame still holds. Almost every number is stale. And the failure that has cost this community the most money since April is not in the guide, because in April it barely existed.

The failure is the cap.

## The problem

Here is what I have watched happen in our own chat over the summer, names removed. One member let his agent loose on a scraping job overnight, got his IP blocked by Cloudflare, and blew through a week of his plan's usage before breakfast. Another signed up for a second thirty-dollar plan from a different vendor because his hundred-dollar plan was locked until the weekend. A third opened his card statement and found a subscription he had stopped using two months earlier. A fourth spent an evening working out whether his X Premium plan and his SuperGrok plan were two different meters or one, and whether he had been paying for both. I have had four different people tell me some version of "this crap is adding up."

I am not exempt. In August I spent as much on overage in one week as my subscription costs in a month. I moved several workflows down to a cheaper model mid-week just to stop the bleeding, and I moved my interactive sessions to a different vendor entirely, then found I was at ninety percent of that plan's weekly limit within days.

None of us were doing anything exotic. We were running agents the way the vendors' own marketing says to run them. The word "unlimited" is on the pricing pages. The word "subject to abuse guardrails" is in the footnote.

## What changed since April

The April guide said there are two ways to pay: a flat subscription, or per-token API billing. That is still true on the invoice. It is no longer true in practice, because every major vendor has quietly turned the subscription into a third thing: a prepaid bucket with a meter in it and an overflow valve that bills you at API rates.

Anthropic's own support pages now describe the plan as a [five-hour session limit plus a weekly limit](https://support.claude.com/en/articles/9797557-usage-limit-best-practices?ref=operators-ai.ghost.io), with a separate weekly meter for Opus. Their flagship reasoning model, Fable, [does not get its own bucket](https://support.claude.com/en/articles/15424964-claude-fable-models-on-your-plan?ref=operators-ai.ghost.io). On Max it can use up to half your weekly allowance, and on Pro it is not in the plan at all; it runs on prepaid credits from the first message. When you run out, [usage credits](https://support.claude.com/en/articles/12429409-?ref=operators-ai.ghost.io) kick in "at standard API rates," with a daily redemption limit of two thousand dollars, which tells you how much some people are spending.

In May Anthropic [doubled Claude Code's five-hour limits](https://www.anthropic.com/news/higher-limits-spacex?ref=operators-ai.ghost.io) after signing a compute deal with SpaceX. Read it carefully: the five-hour window doubled. The weekly bucket did not. They bought a data center to widen the spigot, and the tank is the same size. Four months on, the complaints are all about the tank.

OpenAI went the other way and simply stopped selling the top tier. New signups and upgrades to ChatGPT Pro at two hundred dollars are [temporarily paused](https://help.openai.com/en/articles/9793128-what-is-chatgpt-pro?ref=operators-ai.ghost.io); if you cancel, you cannot buy it back until the pause lifts. There is a hundred-dollar Pro tier now, "5x" versus "20x" Plus. Codex and ChatGPT Work [share one usage pool](https://learn.chatgpt.com/docs/pricing?ref=operators-ai.ghost.io), the vendor's own docs say tasks "that look similar can consume different amounts of your allowance," and the [credits rate card](https://learn.chatgpt.com/docs/pricing?ref=operators-ai.ghost.io) prices their new flagship, Astra, at 1,250 credits per million output tokens against 30 for the cheap model. Support [will not reset your limit](https://help.openai.com/en/articles/9793128-what-is-chatgpt-pro?ref=operators-ai.ghost.io). Wait, or pay.

xAI sells [SuperGrok at thirty, SuperGrok Plus at a hundred](https://x.ai/pricing?ref=operators-ai.ghost.io), and Heavy at three hundred, and Grok's own account [describes the paid tiers](https://x.com/grok/status/2099991712156307916?ref=operators-ai.ghost.io) as "a shared weekly usage pool across features" whose size is "by tier, not fixed." Which is at least honest. The interesting move is that xAI lets you use the subscription inside third-party tools, Hermes and OpenClaw and Warp and Kilo among them, at subscription pricing. That is the opposite of the April situation, where the story was Anthropic cutting harnesses off.

And it is not just us. In the last thirty days: a Reddit thread of [ChatGPT Pro users hitting the weekly limit within one or two days](https://www.reddit.com/r/OpenaiCodex/comments/1wgaxp2/are%5Fother%5Fchatgpt%5Fpro%5Fusers%5Fhitting%5Fthe%5Fweekly/?ref=operators-ai.ghost.io) of the Astra release, one Plus user reporting the cap after "2-4 cumulative hours" of work. A developer on X [paying seven hundred dollars a month](https://x.com/nvandewetering/status/2099853692446945534?ref=operators-ai.ghost.io) across vendors and "rationing prompts" by midweek. People posting their subscription stacks like scorecards: [two Claude Max seats plus Pro plus SuperGrok](https://x.com/one%5Fdyor/status/2099081759442784294?ref=operators-ai.ghost.io), about six hundred a month, "roughly 4 days of continuous operation." A [thread with two hundred likes](https://x.com/TimJayas/status/2097948972702921125?ref=operators-ai.ghost.io) whose entire point is that people want the next Grok to be good so they can cancel the other two.

That last one is the tell. Nobody is choosing a model. They are choosing which bucket to drain next.

## How to identify it

You are in the cap trap if any of these are true.

You upgraded a tier because of a few heavy days, and you cannot say whether the heavy days were you at the keyboard or something running while you slept. The April guide's advice was "track when you hit limits before upgrading." That is still right and almost nobody does it, because the meters are buried three clicks deep and each vendor buries them somewhere different.

You hold more than one paid seat inside the same model family. Two Claude plans. ChatGPT Pro plus a Codex-only habit that bills against the same pool. X Premium and SuperGrok without knowing which features draw on which. A second seat is sometimes the right call. It is never the right call by accident.

You cannot name the single automation that consumed the most of your allowance last week. If you have anything scheduled, a heartbeat, a cron, a bot that wakes up every half hour to check on things, and it is pinned to a premium model, that job is eating your bucket around the clock at a rate that has nothing to do with how much work you did.

That third one is the one that got me. After my August week I had my agent audit why I was still generating overage charges when my overall allowance had room left. The answer was a scheduled orchestrator for one of my businesses, pinned to the premium model, waking up every thirty minutes. Eighty-nine sessions and over a thousand API calls in forty-eight hours, drawing against a bucket that was already empty, so every call went straight to the overflow valve. It was doing useful work. It did not need the expensive model to do it. Nobody had decided it should run there; it inherited the default.

## What the subscriptions are hiding

Last week I ran an exercise I recommend to everyone in this group. I had my agent pull every model call it had logged in the previous thirty days, about forty thousand of them across every provider I use, and price the whole month at published per-token rates as if I had paid for it by the token instead of by subscription.

I pay somewhere in the range of four hundred dollars a month across my subscriptions. The same thirty days, priced by the token, came to a little over six thousand. Roughly fifteen times what I pay.

That number is the real cost of running agents the way I run them, and it explains everything else in this issue. The subscriptions are not a good deal. They are an absurd deal, subsidized by vendors fighting for market share, which is exactly why every one of them has spent the summer bolting meters and valves onto the word "unlimited." The cap is not a bug in the product. The cap is the product finding its price.

Three more things the exercise turned up, which I would not have known any other way:

The bill was not evenly spread. Two models accounted for about eighty percent of the token-priced total, and one of them was a flagship I had switched to as my default eight days earlier. My daily burn went up roughly five times the week I made that switch. Not because I worked five times harder. Because the default changed.

One day was a third of the month. A single day of board-driven subagent work, sessions running ninety to a hundred and twenty calls each with fifteen million tokens of context apiece, priced out at more than a third of the entire thirty-day total on its own. That was the day I noticed I was "burning tokens too fast." I had no idea it was that bad until the exercise put a number on it.

Caching was doing almost all the work. Ninety-six percent of my tokens were cache reads, billed at a fraction of full price. Without cache discounts the same month would have been closer to forty thousand dollars. If you route through anything that does not pass caching through to the vendor, your token-priced cost can go up five to ten times without your usage changing at all.

The exercise took my agent about twenty minutes. It is in the playbook below. Run it before you touch a single plan setting, because you cannot route what you have not measured, and you will not believe the meter until you see the number.

## How to fix it

The frame that fixes this is simple: a subscription is not a flat fee. It is a rate. You are buying a certain flow of tokens per five hours and per week, with an expensive valve at the end. Treat it the way you would treat a utility with a demand charge.

Three moves, in order. Meter before you route, route before you retire.

**Meter.** Once a week, read every bucket you pay for and write the numbers down. Ten minutes. The vendors will not do this for you and their dashboards do not talk to each other.

**Route by bucket, not just by task.** The routing audit I published in August sorted work by what it costs when it is wrong. Add a second axis: which bucket does this draw from, and how fast. Anything scheduled runs on the cheapest model that passes an acceptance test. Premium models are invoked by a human, on purpose.

**Split the work by model.** One of our members passed along a pattern from a developer on his team who went from blowing through three usage resets in a day to using ten or twenty percent of his weekly allowance per day. Plan with the expensive model, build with the cheap one, review with the expensive one. I run the same thing through my task board. The trick is that the expensive model writes the plan and the handoff, and the cheap model never has to think, only build.

**Retire before you upgrade.** No tier upgrade until you have run the meter for three straight weeks with the background jobs already demoted and the plan/build split in place. Most of the time the cap was never the problem. The default was.

Below the line is the full version: the bucket map for all three vendors with where each meter lives, the thirty-day pricing exercise so you can see your own real number, the weekly meter read, the stack audit that finds the seat you forgot, the leak hunt for scheduled jobs, the plan/build/review split with the handoff template, the overflow decision table with the actual credit math, and a handoff you can give your agent to build the whole thing for you.

## The playbook

This is written to be handed to your agent. Each section has the rule, the reason, and the concrete step. If you ingest these issues into your own system, the handoff at the end is the part to keep.

### 1\. Know your buckets

Every vendor now has the same four parts. Learn the four, and the plan names stop mattering.

|                                | Anthropic (Pro / Max)                                                                                            | OpenAI (Plus / Pro)                                                                       | xAI (SuperGrok tiers)                                                                           |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
| Short window                   | Five-hour session limit                                                                                          | Five-hour window (Codex/Work estimates are "per five-hour period")                        | Rate limits per feature                                                                         |
| Long window                    | Weekly limit, separate Opus weekly meter                                                                         | Weekly limit (the one people hit in two days)                                             | Shared weekly pool across features, size by tier (per Grok's own account)                       |
| Premium-model rule             | Fable: up to 50% of weekly on Max; credits-only on Pro; weighs roughly double an Opus session                    | Astra: 250 credits in / 1,250 out per million tokens vs Luna at 5 / 30; Fast mode is 2.5x | Heavy models on higher tiers; Build has its own limits                                          |
| Overflow valve                 | Usage credits at API rates; monthly cap settable; $2,000/day redemption limit                                    | Credits (not API credits); auto-reload optional; balance can go negative mid-task         | Upgrade tier or API key                                                                         |
| Where the meter is             | Settings > Usage on claude.ai (session bar + weekly bars)                                                        | chatgpt.com/codex/settings/usage, or /status in Codex CLI                                 | grok.com account usage; harness-side hermes portal info style commands if you route through one |
| Third-party harness on the sub | Yes via OAuth in some harnesses (Hermes lists "Claude Max + extra usage credits via OAuth"); check current terms | Yes via ChatGPT OAuth device login in Hermes and Codex CLI                                | Yes, first-party: xAI publishes the list of supported tools                                     |

Sources: the Anthropic usage-limit, Fable, and usage-credit support pages linked above; OpenAI's Codex pricing doc and Pro-tier help article; x.ai/pricing; the [Hermes provider docs](https://hermes-agent.nousresearch.com/docs/integrations/providers?ref=operators-ai.ghost.io). Recheck the table quarterly. It was wrong within five months last time.

Two rules that fall out of the table:

Never enable an overflow valve without setting its monthly cap in the same sitting. Anthropic offers "set to unlimited" as a button. Do not press it. OpenAI's balance can go negative if a task finishes after your concurrent jobs drained it, and the next purchase pays the debt first.

Know which of your models are "outside the bucket." Fable on a Pro plan is pay-as-you-go from the first token. If you did not know that, you have been paying API rates while believing you were on a subscription.

### 2\. The thirty-day pricing exercise

Do this once before anything else. It turns "I think I use a lot" into a number, and the number changes how you feel about every other section.

Most harnesses log every model call with token counts. Hermes keeps them in a local database; Codex and Claude Code expose usage in their dashboards and CLI status commands; if you run through an aggregator like OpenRouter or a portal, the billing page has the raw totals. Your agent can find its own logs faster than you can.

The recipe:

1. Pull thirty days of calls grouped by model: input tokens, output tokens, cache-read tokens, cache-write tokens.
2. Price each model at its published per-token API rate. Use the vendor's own pricing page or a portal catalog that lists all of them in one place. Apply the cache-read discount where the vendor publishes one.
3. Sum by model, by day, and total. Compare the total to what you pay in subscriptions.
4. Look for three things: the model that dominates, the day that dominates, and the share of tokens that are cache reads.

What you will find, if my numbers and our members' numbers are typical: the ratio between token-priced cost and subscription cost is somewhere between five and twenty. One or two models are most of it. One or two days are most of it. And the day that dominates is almost never a day you sat and typed; it is a day something ran in a loop.

Write the four findings at the top of your meter log. They are the baseline every later decision is measured against.

### 3\. The weekly meter read

Monday, ten minutes, before you start work. Open each vendor's usage page and record four numbers per plan in a plain text file or a sheet:

```
date | vendor/plan | session % used | weekly % used | weekly reset day | overflow spent this month

```

After three weeks you will know your real pace, which is the number every upgrade decision depends on and the number nobody has. OpenAI's own docs say to check the dashboard "every week or two to understand your pace." They are right. Make it weekly and make it written, because a screenshot in your head is not a trend.

Add one column for "what ran": the biggest thing you or your agents did that week. When the weekly number spikes you want to be able to point at the cause.

### 4\. The stack audit

Once, this week, then quarterly. List every AI subscription you pay for, personal and business cards both.

```
plan | price/mo | model family | what it uniquely gives me | what it duplicates | last used

```

Then apply three rules.

One primary seat per model family. If you hold two seats in the same family, one of them must have a written reason: a genuinely separate bucket for a genuinely separate workload (a team seat for a VA, a second Max seat for an always-on build box). "I ran out on Tuesday" is not a reason; it is a routing problem, and section 5 fixes it.

Find the ghost. Search your email for receipts from anthropic, openai, x.ai, and whatever else you have tried since spring. One of our members found a plan he had stopped using in June. Cancel from the vendor's billing page, not the app store, and note that OpenAI's Pro two-hundred tier cannot currently be repurchased once you drop it. If you are on it and you rely on it, do not cancel it to "test" the hundred-dollar tier.

Untangle the bundle. X Premium and SuperGrok are separate purchases with overlapping access. If you use Grok through a harness, know which subscription is actually authenticating. If it is X Premium, it draws on that plan's limits, not SuperGrok's, and a member of ours learned this by hitting the ceiling on a plan he did not know he was using.

### 5\. The leak hunt: scheduled jobs

This is where the money goes when you are not at the keyboard.

Inventory every job that runs without you pressing a button: cron jobs, heartbeats, board orchestrators, watchers, bots that poll a chat, ambient subagents. For each:

```
job | schedule | model it is pinned to | bucket it draws from | calls per day (estimate) | does the output need a premium model?

```

The multiplication is the whole point. A job pinned to a premium model every thirty minutes is 48 wakeups a day, before it does anything. Mine was drawing over a thousand calls in two days against a bucket that was already empty, which meant every call was billed at the overflow rate. It was also the only job in the fleet that had been pinned by accident: it inherited the model from the profile that created it.

Rule: nothing scheduled runs on a premium or bucketed-separately model. Background work runs on the cheapest model that passes an acceptance test on last week's real inputs. OpenAI describes Luna as built for "routing, classification, extraction, support, background automation" and prices it at a fortieth of Astra. Anthropic's own advice when you are near the limit is to use Haiku or Sonnet. The vendors are telling you where to put background work. Listen.

Re-run the acceptance test when you demote a job. Take five real inputs from last week, run them on the cheap model, compare. If it passes, it stays demoted. If it fails, the job is not background work; it is judgment work wearing a cron schedule, and it should be invoked by a human.

### 6\. Routing by bucket

The August routing audit gave you the first axis: cheap to be wrong versus expensive to be wrong. Add the second.

| Lane        | What it is                                                                     | Where it runs                                                                                                             | Who invokes it  |
| ----------- | ------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------- | --------------- |
| Interactive | You at the keyboard, back and forth                                            | The subscription you are actively paying for; the best model your plan includes without a special bucket                  | You             |
| Background  | Anything scheduled or polling                                                  | Cheapest model that passed acceptance; a cheap API key or a separate low tier so it never touches your interactive bucket | Cron            |
| Judgment    | Decisions with money or customers behind them, copy that ships, hard debugging | Premium model, manually selected, on a bucket you have checked first                                                      | You, on purpose |

The mistake I made in August was letting background work share a bucket with judgment work. When the bucket ran dry, it was the judgment work that got locked out on a Tuesday, because the cron had drained it overnight. Separate them physically if you can: background on its own key or plan, interactive and judgment on the subscription. A cheap second seat used this way is the one case where two seats is the right answer.

### 7\. The plan / build / review split

This is the pattern that moves the needle on interactive work, where sections 5 and 6 do not reach. One of our members relayed it from a developer on his team who consults for someone who had been spending two thousand dollars a week on coding-tool credits. The developer went from burning three usage resets in a day to using ten or twenty percent of his weekly allowance per day. Nothing about it is secret. Almost nobody does it, because you rarely run out early enough in the week to bother.

The shape:

1. **Plan on the expensive model.** Give your flagship (Sol or Astra on OpenAI, Fable or Opus on Anthropic, the top Grok tier) the problem and have it write a full plan. Use an interview-style skill if your harness has one; the more the plan pins down, the less the builder has to reason.
2. **Read the plan yourself and edit it.** This is the step people skip and it is the step that makes the rest work. You are the second expensive reviewer, and you are free. Fix what is wrong, cut what is unnecessary, add the constraint it missed.
3. **Have the expensive model write a handoff.** A short document addressed to the builder: what to build, in what order, what "done" looks like, what not to touch. Plan plus handoff go to the cheap model together.
4. **Build on the cheap model.** Terra on OpenAI, Sonnet or Haiku on Anthropic, a cheap Grok or an open-weights model through a portal. The member's developer reported that even at high reasoning and fast mode, the cheap tier stayed cheap. It is executing a plan, not inventing one.
5. **Review on the expensive model.** After the build, hand the diff or the output back to the flagship for a deep review. This is a single expensive call at the end instead of hundreds of expensive calls in the middle.

Where the savings come from: the expensive model touches the work twice, at the start and at the end. Everything in between, which is most of the tokens, runs on a model priced at a fifth to a fortieth of the flagship. My thirty-day exercise showed a flagship default costing five times what the mid-tier had cost the week before, on the same work. This pattern is how you get that five back without giving up the flagship's judgment.

I run the same shape on a task board: the flagship game-plans and seeds the tasks, workers pick them up on cheaper models, a scheduled check every few minutes looks for blocked cards and hands them back to the planner. The board is the plan and the handoff in one place. You do not need a board to do this. You need the discipline to not let the builder plan and to not let the planner build.

Acceptance test for the split: take one real build task from last week. Run it the old way and time your allowance drop. Run it the new way. If the new way does not cut the drop by at least half, your plan step is not pinning enough down, and the fix is a better plan, not a bigger model.

### 8\. The overflow decision

You hit the cap mid-task. Four options. Decide in advance, not while the session bar is red.

| Option            | When                                               | Cost                                                                                                    | Trap                                                 |
| ----------------- | -------------------------------------------------- | ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- |
| Wait              | Session cap, not weekly; work is not urgent        | Zero                                                                                                    | Weekly cap means days, not hours                     |
| Switch model down | Most of the time                                   | Zero; usually fine for the current step                                                                 | Forgetting to switch back for judgment work          |
| Credits           | The task is judgment work and the deadline is real | API rates. Astra is 1,250 credits per million output tokens; Anthropic credits are "standard API rates" | No monthly cap set; auto-reload on; negative balance |
| Separate API key  | You do this every week                             | Per token, predictable                                                                                  | Now you are paying twice for the same model family   |

Set the credit cap at what you would be annoyed but not harmed to lose in a month. For most operators in this group that is between fifty and a hundred dollars. If you hit it two months running, you have a routing problem, not a budget problem. Go back to sections 5 and 7.

### 9\. The upgrade test

Only upgrade a tier when all three hold:

1. Three consecutive weekly meter reads show you hitting the weekly cap, not just the five-hour session.
2. Every scheduled job has already been demoted and passed its acceptance test, and interactive build work runs on the plan/build/review split.
3. The work that hits the cap is interactive or judgment work you can name.

If the third is not true, you do not know what you are buying more of. Apply the same test in reverse: three weeks under fifty percent of the weekly cap and you are a candidate to drop a tier. Note the OpenAI asymmetry before you drop anything at the top.

### 10\. Handoff for your agent

Paste this to your agent. It builds the meter log, the stack audit, and the job inventory, and it does not change anything on its own.

```
You are helping me control AI subscription spend. Do not change any settings,
cancel anything, or move any job to a different model without my approval.

1. Create a folder called ai-spend with three files: meter-log.md,
   stack-audit.md, scheduled-jobs.md.

2. meter-log.md: a table with columns
   date | vendor/plan | session % used | weekly % used | weekly reset day |
   overflow spent this month | biggest thing that ran.
   Add a reminder for me every Monday to fill it in. Where you can read a
   vendor's usage page or a harness status command yourself, do it and
   pre-fill the row; otherwise leave it for me.

3. stack-audit.md: list every AI subscription you can find evidence of in my
   config, environment, and (if I give you access) email receipts. Columns:
   plan | price/mo | model family | unique value | duplicates | last used.
   Flag any two seats in the same model family and any plan with no use
   in 30 days.

4. scheduled-jobs.md: inventory every cron job, heartbeat, orchestrator,
   watcher, or bot you can find in my setup. Columns:
   job | schedule | pinned model | bucket it draws from | est. calls/day |
   premium needed? For each job pinned to a premium or separately-bucketed
   model, propose the cheapest model to test and write a 5-input acceptance
   test using last week's real inputs. Do not run the demotion. Give me the
   list and wait.

5. Run the thirty-day pricing exercise: pull every model call you logged
   in the last 30 days with input, output, cache-read, and cache-write
   token counts, grouped by model and by day. Price each at the vendor's
   published API rate with cache discounts applied. Write the result to
   ai-spend/pricing-exercise.md with: total token-priced cost vs my
   subscription spend, the top two models, the top two days and what ran
   on them, and the cache-read share. Do not change any routing.

6. Report back with: the token-priced total and the ratio to what I pay,
   the biggest scheduled consumer, any duplicate seats, any ghost plans,
   and the one change you would make first.

```

### Steal this

Price your last thirty days by the token once, so you know what the subscriptions are hiding. Read the meters every Monday. Nothing scheduled runs on a premium model. Plan and review on the flagship, build on the cheap tier, and read the plan yourself in between. Set the credit cap the day you turn credits on. Upgrade only after three weeks of written evidence with the background jobs already demoted. The cap was never the problem. The default was.

Hit reply and tell me what your meter said on Monday. I read every one.