This week: the model wars got a price list
Five vendors shipped in four days, two flagships landed at the identical price, the cheap tier got 100x cheaper than the top, and Musk says his is next. What the benchmarks won't tell you, and what the invoice will.
Monday, Anthropic. Wednesday, Google and Meta. Thursday, OpenAI. Alibaba and Z.ai the week before, and Musk on Wednesday promising xAI's next one "in 10 days." One indie builder counted eight releases in 48 hours. I have been doing this a long time and I do not remember a week where every major lab shipped at once, with prices attached, aimed at the same customer.
That customer is you. The labs are no longer fighting over developers; they are fighting over who runs your quotes, your inbox, and your books. So this issue is the week as a price list, then what the price list does not show.
| Model | Vendor | Shipped | API price, per million tokens (in / out) | Who gets it in the app |
|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Sep 3 | $10 / $50 | Pro $100 and $200, Business, Enterprise. Plus: unconfirmed. Free: no |
| Claude Fable 5.1 | Anthropic | Sep 1 | $10 / $50 (cache reads cut to $0.25) | Max includes it (capped); Pro pays per use; Free: no |
| Gemini 3.8 Flash | Sep 2 | $0.75 / $3.75 until Dec 31, then $1.50 / $7.50 | Google AI Pro and Ultra | |
| Muse Spark 1.3 | Meta | Sep 2 | $1.25 / $4.25, or $0.10 / $0.20 if Meta can train on your data | Developer surfaces only |
| Qwen3.8-Flash | Alibaba | Aug 26 | $0.15 / $0.47 | Hosted API; weights are open |
| GLM-5.3-Flash | Z.ai | Aug 26 | $0.075 / $0.25 on promo through Sep 9, then $0.15 / $0.50 | Hosted API; weights are open |
| Grok 4.6 | xAI | Aug 12 | $2 / $6 | Grok, API; 4.7 promised for ~Sep 12 |
Sources for every row are linked in the gauges below. The spread from cheapest to most expensive capable model this week is roughly 100x on input. Hold that number.
On the bench
My desk runs four vendors, and this week it ran all four on purpose.
The everyday agent moved to Claude Fable 5.1 on Monday afternoon, after a test, and it is where most of my work landed: a technical session prep for a consulting client, a month-end invoice, a Q&A study sheet. A second bot, the one I use for deep research into a specific technical platform, I deliberately pinned on Monday night to OpenAI's GPT-5.6 Sol at high effort instead of max. That was a cost decision. Max effort on a research bot burns tokens like a space heater, and the output at high was good enough that I could not justify the difference. A group chat between two of my agents ran on xAI's Grok 4.6 because it is cheap and fast enough for agents talking to each other. And Thursday, when GPT-6 Astra shipped and I did not get access, I spent the evening writing four long job descriptions for it anyway, one per business, so the day it arrives it goes to work instead of to a demo.
Here is the confession inside that last one. My first Astra draft was a single prompt: audit every project I own, find improvements, propose new features. It is the exact mistake I tell readers not to make. A vague "make everything better" job gives the model nothing to check itself against, and it will burn its budget producing mush. I rewrote it into four prompts, each with its own definition of what "better" means for that business (booked machine-days for the land clearing company, visible profit for the reseller, subscribers for this newsletter). Better prompts. Still prompts for an employee who has not started.
The point of the bench this week is not the four vendors. It is that none of them has my loyalty. Each job on my desk has a model, a reason, and a price, and when a better price or a better behavior shows up, that one job moves. Not the whole desk.
The gauges
The flagships now cost exactly the same, and one of them refuses to guess
What I saw: Claude Fable 5.1 on Monday and GPT-6 Astra on Thursday both landed at $10 per million tokens in and $50 out, both with roughly a million tokens of context. That is not a coincidence; it is a price war where the two leaders agreed on the price and are fighting on everything else. Anthropic's real move is underneath: cached reads dropped from $1.00 to $0.25, which Anthropic says makes long agent sessions up to 45% cheaper. OpenAI's move is a 2.5x price jump over GPT-5.6 Sol, justified by computer use: filling forms, updating CRM records, building decks that follow your template.
The number I keep coming back to is not a benchmark. Zapier's CEO ran both on his company's automation test suite and reported this: asked to answer and log 15 customer inquiries using a reply standard stored in a document, Astra searched, could not find the standard, and stopped. Zero replies sent. Sol could not find it either, took its best shot at all 15, and got partial credit. Vendor test, take it as such. But that is a behavior difference, and it is the one an operator should care about: "Astra will not guess. When it cannot find the instructions, it pauses the work instead of improvising."
Should you pay attention? Yes, to the behavior, not the leaderboard. In August I gave you the routing label: route by what it costs when the model is wrong, not by how hard the task is. For the jobs where a wrong record is expensive (reconciliation, quotes, anything touching a customer's money), a model that stops is worth more than a model that finishes. Zapier found the same model weaker on "outbound comms where the guidance is scattered." Same model, opposite verdicts, depending on the cost of being wrong.
My position: $10/$50 is not a price for "the default model." It is a price for the two or three jobs in your business where a wrong answer costs more than fifty dollars. Everything else belongs a tier down, and that tier is where the real war is.
Google and Meta are fighting on price, and the fine print is the weapon
What I saw: Gemini 3.8 Flash is Google's third Flash model in six weeks (3.6 in July, 3.7 in August, 3.8 on Wednesday), held at $0.75 in / $3.75 out and available to Google AI Pro and Ultra subscribers in the Gemini app, AI Mode in Search, and Gemini in Sheets. Same day, Meta shipped Muse Spark 1.3 at $1.25 / $4.25, with a "Contributor" tier at $0.10 / $0.20. One user did the arithmetic: 7.5x cheaper input than Gemini Flash. Another ran the same build on both and got $0.04 on Muse versus $3.32 on Gemini, with Gemini's output scored one point higher out of ten. Behind both, Alibaba's Qwen3.8-Flash at $0.15 / $0.47 and Z.ai's GLM-5.3-Flash at $0.075 / $0.25 on promo, both with open weights you can run yourself.
Now the fine print, because on the cheap tier the fine print is the product. Meta's Contributor price is cheap because you are granting Meta permission to train on your prompts and completions; the $1.25 tier is the one without that clause. Google's price is introductory: it doubles to $1.50 / $7.50 on January 1, and Google says 3.8 "works harder," meaning more reasoning steps and more tool calls per task, so the same per-token price can produce a bigger bill. Z.ai's promo ends September 9. Qwen and Z.ai are Chinese vendors, which is a data question for some readers and a non-issue for others, but it should be a decision, not a default. And the field's read on the cheap tier's weakness is consistent: fast, cheap, and it "never asks for clarification, no matter how I ask, it does something". Muse Spark failed a from-scratch task in five attempts that Gemini did in one.
Should you pay attention? Yes, more than to the flagships. The routine 90% of a small business's AI volume belongs on this tier, so this tier's fine print is your fine print. Do not route customer data through a price that assumes your data is the payment. Do not budget on an introductory rate without writing down the date it ends.
My position: Google is racing on cadence and Meta is racing on price, and both are good for you as long as you read the two lines above the number. For a quote draft, an inbox triage, a summary, at a few tenths of a cent per task, these are the models. For anything with ambiguous instructions, expect them to guess.
Grok 4.7 is next, if you believe the date
What I saw: on Wednesday, Musk said xAI's Grok 4.7 "comes out in 10 days," which points at roughly September 12, reportedly a 2.1-trillion-parameter model. No price published. The current model, Grok 4.6, shipped August 12 at $2 in / $6 out, which puts xAI in the middle tier: five times cheaper than the flagships, three times pricier than Gemini Flash. For context, 4.6 was originally targeted for around August 7 and shipped on the 12th, and Grok 5's Q1 and Q2 windows both passed without a release. Do not put September 12 on a calendar.
Should you pay attention? Mildly. Grok 4.6 is a real option for internal, agent-to-agent, and low-stakes work (that is what I use it for), and xAI has been aggressive about getting it into other people's tools: Cursor, GitHub Copilot, Amazon Bedrock, Microsoft Foundry all added it within two weeks of launch. The thing to watch when 4.7 lands is not the benchmark; it is whether the $2/$6 price holds. If xAI keeps a frontier-adjacent model at a fifth of the flagship price, the middle tier becomes the most interesting place on the list.
My position: a rumor is not a plan. If your work runs fine on 4.6 today, 4.7 changes nothing until it exists, has a price, and passes your test.
What the benchmarks don't show and the invoice does
What I saw: every launch page this week had the same shape, a wall of charts claiming best-in-class on tests you have never heard of. One HN commenter put it plainly: "The models change constantly and relentlessly and so does the pricing, basically weekly at this point... it feels nearly impossible to have any rigorous approach when choosing a particular model." Meanwhile the field data about the best-benchmarked model of the week was about none of the benchmarks. Three things the charts left out:
The ration. Fable 5.1 is widely called the best model anyone has used and "unusable" in the same post, because it drains subscription limits in a day or two. A team-plan user reported 70% of a weekly limit in one day on the same workflow that used 30% on the previous model; a $200 Max subscriber, 50% in under 24 hours. Anthropic's own plan page says Max can spend at most 50% of its weekly limit on Fable models, and Pro plans pay per use from the first message. OpenAI's help center says Astra arrives in ChatGPT as "GPT-6 Pro" with 200 messages a week on the $200 plan and 50 on the $100 plan, and is "not included with ChatGPT Plus in Chat," while the launch blog says Plus gets it "over the coming days." Two OpenAI pages, two answers. Free tiers got nothing from anyone.
The switch. Fable 5.1 runs safety classifiers on every request, and when one trips, Claude re-runs your request on an Opus model and leaves the picker on Opus for the rest of the conversation. The checks read everything the model reads, "including memory, content from connectors, web search results, and files, so a block can be triggered by content you didn't type." One developer's strategy game tripped the biology filter because an in-game joke used the word "biological." OpenAI says the same about Astra in plainer words: safety checks "can sometimes slow, pause, or stop legitimate work." The model in your settings is a request, not a guarantee.
The outage. Wednesday morning, Claude, ChatGPT, Grok and Gemini all went down or degraded in the same three-hour window. Nobody has published a cause. Four vendors at war with each other, and every one of your fallback options was down at the same time.
Should you pay attention? This is the gauge to pay attention to. A benchmark measures the model on its best day with unlimited budget. Your business runs it on a Wednesday with a weekly cap, a safety filter reading your files, and a status page you were not watching.
My position: usability is the benchmark. How many tasks the ration buys you, whether the model that answered is the one you picked, and what your tool did at 10am Wednesday. Those three numbers decide whether a new model is an upgrade. No launch page prints them, so you have to.
Your move
Ten minutes. Do it today, before you touch a model picker.
- (3 min) Write down every AI tool your business uses, and next to each, the model it runs right now. Not the vendor, the model. If you do not know, that is finding number one.
- (2 min) Next to each model, write its price from the table above, or "subscription" and the weekly cap. Most people discover here that their most expensive model is doing their cheapest work.
- (3 min) For each tool, answer one question: what happened at 10am Wednesday? Queued and retried, fell back to a person, or silently produced nothing? If you cannot answer, open the tool and find out. Wednesday was a free drill and most people did not notice they were in it.
- (2 min) Pick the one workflow where a wrong answer costs real money. That is the only workflow that gets a new model this month, after a side-by-side test on a real task with a known answer. Everything else: write "no change" and close the picker.
Every operator who told me they "upgraded" this week meant they changed a dropdown. The ones who got value from it ran a test first.
The meme

Fourteen models in August. Eight in 48 hours this week. The workflow that already works did not get worse on Thursday.
The playbook: the upgrade gate
Operators: Your move above is the ten-minute version. This is the standing gate, written so you can run it every time a vendor ships, which is now roughly weekly. The August routing audit told you which model belongs on which job. This one tells you how to change a model without breaking the job.
The idea
A new model is a new employee who arrived with a glowing reference and no probation period. The vendor's charts are the reference. Your business is the probation. The gate has five parts, in order: proof of tools, a pinned test set, the ration math, the fallback rule, and the outage rule. Skip one and you find out the hard way.
1. Prove the tools work before you believe they work
Two tiny jobs, in a fresh session, before anything real. This works in any harness: ChatGPT, Claude, Gemini, Hermes, Codex, or whatever desktop app you run.
The tool job. Pick a tool the model must use in your setup (a calendar read, a file search, a database query, a spreadsheet lookup). Give it a task with exactly one correct output that can only come from using the tool: "Read today's calendar and return only the title of the 2pm meeting." If the model returns a plausible title without a tool call showing in the activity log, it guessed. Fail.
The integration job. One hop deeper: a live call into a system of record (CRM, accounting, knowledge base) that returns a value the model could not know. "Return the invoice number of the most recent invoice in the books." If it answers from memory, fail.
Both prompts carry the clause "do not answer from memory and do not merely claim the tool ran." Read the log, not the answer. "I ran that" is a claim; the log is the receipt. New models fail this more often than you would think, usually because the tool surface changed and nobody noticed.
2. The pinned test set: five real tasks, known-good answers
Pick five tasks from your actual business, one per job type you route: a customer reply, a quote or estimate, a document summary, a data pull or reconciliation, and your fifth most common job. For each, save the input and the answer you consider correct. That is your test set. It lives in a folder, it does not change unless the business changes, and every new model runs against it before it runs against a customer.
Score each result three ways. Correct or not. Cost per task, not per token: Anthropic's own charts show a cheaper token can cost more per task if the model uses more of them, and one third-party measurement put Fable 5.1 at 15% more per task than Fable 5 despite the cache cut; Google warns 3.8 Flash uses more tokens than 3.7 at higher effort. And the missing-information case: take one of your five tasks, remove the instruction the model needs, and see whether it stops and asks or improvises. Zapier found Astra stops and Sol improvises; the field found Gemini Flash never asks. Neither is right for every job. Decide per job, and write it down.
Effort levels are part of the test. Fable 5.1 at low or medium effort matches the previous model "at a much lower cost," per Anthropic. Run your five tasks at default effort and one notch below. The notch below is frequently the right answer for routine work, and it is the single cheapest cost cut available to you this week.
3. The ration math
Your subscription is a weekly dollar budget with the dollar sign hidden. Do the arithmetic once.
Find your plan's weekly limit and any per-model ceiling (Claude Max: 50% of weekly on Fable models; ChatGPT Pro $200: 200 GPT-6 Pro messages a week; Pro $100: 50). Estimate the flagship's per-task burn from your test set (a heavy Max user measured about 1.7x the cost per turn versus Opus on his workload). Divide. That is how many flagship tasks you get per week. If that number is smaller than the number of wrong-is-expensive jobs you have, the flagship is not your default for those jobs either; it is your escalation, used when the cheaper model's output fails a check.
Then price the API path for the same jobs. At $10/$50, a 3,000-token-in, 1,000-token-out task costs about eight cents. At Gemini Flash's $0.75/$3.75 it costs about six tenths of a cent. At Qwen's rate, a twentieth of a cent. Most operators find the subscription is the right buy for the flagship and the API is the right buy for the cheap tier, because volume lives on the cheap tier. Write both numbers on the same card.
4. The fallback rule
Decide, per workflow, what happens when the model you named does not answer.
In the consumer apps, switching is on by default and the response is labeled. Rule: any output that feeds a decision gets its model label recorded with it. If your tool cannot show you which model answered, that tool does not get the wrong-is-expensive jobs.
On the API, Anthropic's switching is off by default and OpenAI's task simply stops. Rule: do not enable silent fallback for anything that touches money or a customer. A blocked request should fail loudly, land in a queue, and get a human. For routine internal work, enable fallback to a cheaper model and log it; later you will read the log and learn which of your inputs trip filters. The "biological" story is funny until it is your garden-supply catalog.
One more from the field: a fleet operator found that older client software errors out when a new model needs a newer version, which "can turn a normal model change into a fleet wide silent failure." If you run more than one agent, update the client before the model, and run the tool job on each one.
5. The outage rule
Wednesday was the drill. For each workflow, one of three answers, written down: queue and retry (with a maximum wait and a person notified when it is exceeded), fall back to a human immediately, or fail loud and stop. "Silently produce nothing" is not on the list. If you have a customer-facing bot, the fallback is a human every time; a three-hour hole on a weekday morning is now a documented event.
If a workflow is important enough to have a second vendor as a backup, the second vendor goes through the same gate. Wednesday also showed that a second vendor is not a guarantee; all four were down together. An untested backup is a second outage waiting behind the first.
The model change record
Every model change gets one page, and the page is the deliverable of the gate:
- Date, tool, old model, new model, effort level, price per task at each.
- Tool job and integration job: pass or fail, with the log line.
- Test set: five tasks, correct or not, cost per task, missing-information behavior.
- Ration: tasks per week at this burn rate; which jobs get them.
- Fallback setting, per workflow, and where the model label is recorded.
- Outage behavior, per workflow: queue, human, or fail loud.
- Decision: switch, switch for these jobs only, or no change. Signed.
Fifteen minutes after the tests. The next time a vendor ships (next week, if Musk is right), you open the last record, run the same five tasks, and compare. That is the whole discipline. It does not require a lab and it does not care about the vendor's charts.
Steal this: before the next release, build the five-task test set and save the known-good answers. That is the one-time cost. Once it exists, every "should I switch?" becomes a twenty-minute question with a written answer and a price, instead of a dropdown and a hope.
Hit reply and tell me what's on your bench this week. I read every one.