CO/AI Subscribe
Tuesday · July 21, 2026 · Issue No. 932
I Am Altering the Deal
Daily Briefing

I Am Altering the Deal

Vader kept changing the terms in Cloud City, and Lando could only stand there and take it. Anthropic just rationed its best model for the third time in two weeks — free, then capped, then bundled — because a rival one point behind costs a third as much and the Chinese ones are free. Mira Murati won't alter your deal. She'll sell you the car.

THE NUMBER: 1 point. That’s how far GPT-5.6 Sol sits behind Claude Fable 5 on the Artificial Analysis Index — and it bills at about a third the cost per task. One point. Behind Sol, Alibaba just previewed a 2.4-trillion-parameter Qwen it plans to give away, and Moonshot’s Kimi K3 had to pause signups because too many people wanted the free one. One point is the entire moat Anthropic is defending, and this week it defended that moat by quietly folding its most expensive model into a cheaper plan. Hold that number. Moats one point wide don’t get charged for. They get bundled away.

☁️ Cloud City

There’s a moment in The Empire Strikes Back that every operator eventually lives through from the wrong side of the table. Lando Calrissian cut a deal with the Empire to protect his little mining colony, and then Vader keeps moving the line. First it’s “leave the ship.” Then it’s Han in carbonite. Then it’s Leia and Chewie staying too. Lando protests — “this deal is getting worse all the time” — and there’s nothing he can do, because he was never the one holding the leverage. He just thought he was.

That’s the Anthropic Fable 5 rollout, compressed. Follow the tape. Fable 5 shipped as a free inclusion. Then the free window was set to close July 7. Then July 7 became July 12 after users howled. Then July 12 became July 19. Then, this Monday, the music finally stopped: Fable 5 is now a permanent feature of the Max and Team Premium plans — but capped at 50% of your normal weekly limits. Pro and Team Standard lost free access outright. You get a one-time $100 credit, and after that you’re on the meter at $10 per million input tokens and $50 per million output tokens, the steepest pricing of any Claude model in general release. On the same day, a separate 50% boost to Claude Code weekly rate limits also quietly expired.

Three changes to the deal in two weeks, each announced by tweet, each one worse than the last for the person who’d organized their week around the old terms. A Claude Code lead engineer had promised in early July that Fable 5 would come back as a standard Pro benefit. The Monday announcement doesn’t honor that. This deal is getting worse all the time, and pray they don’t alter it further.

Future Proof Podcast Episode 14

Ep 14 – Google Is Stalling, Elon Bought Cursor for $60B, and AI Still Has No User Manual

Google sheds $200B and loses top talent while Elon buys Cursor to control the developer toll road. Harry and Anthony break down the reality of enterprise AI, on-prem security, and why the best tech doesn’t always win.

🎭 A Price Cut in a Capacity Costume

The company’s explanation is capacity. Demand for Fable 5 was, in their words, hard to predict, and serving it costs a fortune per token, so they had to ration. And look — that’s not a lie. The hardware wall is real. This is a company reportedly leasing up to $10 billion of compute from Meta over two years, on top of a $1.25 billion-a-month deal with SpaceX for Colossus. You don’t sign checks like that if the servers are humming along with room to spare. The three delays weren’t theater. The scarcity is genuine.

But genuine scarcity and a quiet price cut are not mutually exclusive, and this week they’re wearing the same coat. Here’s the tell. When you actually can’t make enough of your most expensive product, you do not fold it into a cheaper plan. You wall it off. You charge a premium for the scarce thing and you make the people who want it pay up. That’s Economics 101, and Anthropic knows it cold. Instead, they took the flagship and made it a bundled sweetener for the $200 Max tier — the tier they most need people to keep paying for. That is not what rationing looks like. That’s what a discount looks like when you can’t afford to call it one.

Remember what Fable 5 actually is. Under the hood it’s the same model as Claude Mythos 5, the version reserved for a small set of vetted partners through Project Glasswing. The only difference is a layer of safety classifiers that kicks queries about cyber exploitation or bio uplift down to Opus 4.8. So this isn’t a story about a scarce, exotic artifact. It’s the same core intelligence everybody’s been using, repriced twice in two weeks — down — while the press release talks about supply.

Why would a company cut the price of its best product in the middle of the biggest demand wave in tech history? Only one reason. Somebody cheaper is standing right behind them.

📉 One Point Wide

That somebody is a crowd now. GPT-5.6 Sol lands within one point of Fable 5 on the Artificial Analysis Index and costs roughly a third as much to run a task through. One point. For the overwhelming majority of real work, one point on a benchmark is the difference between a wine you’d pay $80 for and one you’d pay $25 for — a sommelier can taste it, and you, pouring it at dinner, cannot. And the GPT-5.6 family now ships with five or six reasoning-effort settings, which means “good enough” isn’t even a different model anymore. It’s a cheaper dial on the same one.

Then it gets worse for the pricing power, because the next tier down is free. Alibaba used the World AI Conference in Shanghai this week to preview Qwen 3.8-Max — 2.4 trillion parameters, multimodal, and slated for an open-weight release — while claiming it trails only Fable 5. Take that claim with a boulder of salt; they shipped no benchmarks and no model card, which is its own tell. But you don’t have to believe the ranking to read the strategy. The same week, Moonshot’s Kimi K3 — 2.8 trillion parameters, open weights — had to pause new subscriptions because demand overwhelmed it, and Morningstar started using the words “DeepSeek moment.” Alibaba also open-sourced the software stack for its own AI chips at the same conference, aimed squarely at prying developers off Nvidia’s CUDA. The UK’s AI Safety Institute noted this month that open-weight models now match the frontier’s cyber performance from four months ago, at a fraction of the cost.

Four months. That’s the whole lead the closed frontier has on the free stuff, and it’s shrinking. When the second-best model is a rounding error cheaper and the third-best is free and downloadable, “I use the best model” stops being a business decision. It becomes a badge — the AI equivalent of the guy who leases the S-Class he can’t quite afford because the neighbors can see the badge in the driveway. Fine for him. A terrible way to run a company’s cost structure.

🧰 The Only Reason They Can Still Charge

So why hasn’t everyone already left? Why does the frontier still post revenue that would make a 1999 CFO reach for the smelling salts, if the alternatives are this close and this cheap?

Because of a single, boring fact that is the most important thing in AI right now: coders are the only people who actually redline a Max plan. The whole Fable 5 drama is a coding story — Claude Code limits, the /goal flag, the terminal agents. Developers have a harness. The work lives in an editor, and a passing test tells them, yes or no, whether the machine did the job. Everybody else — the person building the deck, writing the marketing plan, clearing the inbox, wrestling the spreadsheet — has no such thing. For them the tenth-best model was already good enough eighteen months ago. The reason they’re still paying the frontier isn’t quality. It’s that nobody has built the front end that routes their actual work, and switching is a hassle they don’t have the staff to eat.

The evidence is piling up in exactly the places you’d expect. A Salesforce VP wrote this week that the enterprise AI pipeline is a “sieve” — cheaper tokens won’t fix it, because the leaks are architectural: context bloat, ungoverned data, agents wandering the data warehouse burning money. Futurum’s numbers are more brutal: the share of enterprises reporting “significant value” from Salesforce’s flagship Agentforce fell to 6.3%, down from 11.8% last November, even as it shows up in more deals. Only 13.3% of organizations have reached the top rung of AI maturity, and what separates the leaders isn’t model choice — it’s legacy integration and workforce readiness. The one deployment that actually worked this week, Tradeshift ripping out its old BI for an agentic system and getting 30x faster queries and 40% lower cost, only worked because they rebuilt the whole process around it first. Everyone else is bolting a smarter faucet onto a leaky pipe and wondering why the water bill went up.

There’s a smaller, human version of the same truth. A writer at Every published a confession this week about her “Island of Misfit Workflows” — the graveyard of AI tools she built with total conviction and never opened again. The ones that stuck had exactly one thing in common: they paid her back immediately. The rest died. That is the 2.2%-of-households number we led with Saturday, told from the inside of one desk. The models are ready. The harness that would make them stick, for anyone who isn’t a developer, does not exist yet.

That missing harness is the whole moat. We said it last week and it’s worth saying flat: Anthropic stopped charging for intelligence a while ago. What it charges for now is the fact that leaving is a pain. Which is a fine business — until someone makes leaving easy, or better yet, sells you a version you never have to leave in the first place.

🔑 She’ll Sell You the Car

Which brings us to the smartest move in AI this week, and it wasn’t made by a frontier lab. It was made by the person everyone wrote off seventeen months ago.

When Mira Murati left OpenAI in 2024, the shrug was universal — another founder, another “safe AI” startup, another round. Then her company went quiet for a year and a half while the round reportedly collapsed and cofounders drifted back to OpenAI. Easy to file under failure. Wrong. On July 15, Thinking Machines shipped Inkling: 975 billion parameters, fully open weights, Apache 2.0, downloadable from Hugging Face with no waitlist and no API to phone home to. The launch note said plainly that Inkling is not the best model on the market. In an industry that runs on superlatives, that kind of candor is the tell that the rest of what they’re saying is true.

But the model isn’t the move. The move is the razor and the blades. Inkling is the razor — free, open, yours to run on your own iron, permanently. You own it the way you own a car you paid cash for; nobody can send a Friday-afternoon tweet that repriced it. The blades are Tinker, their fine-tuning service: you bring your data and your algorithm, they run the training infrastructure, and — this is the part that matters — the model that comes out the other side is yours. Not rented. Not metered. Not subject to a classifier that reroutes you to a weaker model when it gets nervous. Yours.

Line that up against everyone else’s deal. Anthropic, OpenAI, Google — they rent you intelligence, and they reserve the right to alter the terms whenever the math or the capacity demands it. This week proved they’ll use that right. Murati is, so far, the only major player selling you something you keep. King Gillette gave away the razor to sell the blades forever; Murati is giving away the razor so you’ll finally own the shave. For the enterprise that’s spent two years afraid to hand its proprietary data to a lab that might compete with it tomorrow — the same fear that keeps 89 cents of every enterprise AI dollar locked to the closed frontier — that’s not a product. It’s an escape hatch with a deed attached.

What This Means For You

We’ve seen this movie before. Every platform era starts with a scarce, expensive, centrally-controlled version of the new thing, and every one of them ends when the capability commoditizes and the value moves to whoever owns the workflow and the data on top of it. Mainframe to PC. On-prem to cloud, then cloud margins to whoever owned the app. The frontier labs are not going away — they’ll keep shipping the best model, and coders will keep paying for the top of the range because they actually use it. But the pricing power is leaking, and this week they told on themselves by cutting the price and calling it a shortage.

Your move has three parts. Stop paying frontier prices for badge-tier work — run your heaviest non-coding task on the second-best model for a day and see if anyone can tell; they won’t. Separate the people who actually redline a plan (your developers) from the people who’d never notice a downgrade (nearly everyone else), and stop buying $200 seats for the second group. And take the one workflow you cannot afford to have repriced or rerouted on someone else’s schedule, and move it onto a model you own — Inkling today, tuned on Tinker if you need it — so that the next time a lab decides to alter the deal, you’re not Lando in the throne room. You’re the one who already bought the airline.

You’ve been leasing. Stop renting the point. Own the model.

Three Questions We Think You Should Be Asking Yourself

  1. Which of your AI seats are badges and which are tools? Coders redline plans; the deck-and-inbox crowd stopped needing the frontier a year ago. If you can’t say, per seat, who actually hits the limits, you’re paying the badge price for the whole company.
  2. What would it actually cost you to leave? Not the token price — the switching cost. The front end, the IT lift, the retraining. That number is exactly what the frontier is charging you rent on. If it’s smaller than you assumed, you have leverage at renewal. If it’s huge, that gap is the moat, and it’s the thing to attack.
  3. What in your stack do you rent that you could own? Anthropic altered the Fable terms three times in two weeks. Somewhere in your business is a workflow whose data or continuity you can’t afford to have repriced by a stranger’s tweet. That’s the first thing to move onto weights you hold.

Sources

  • Anthropic Fable 5 tier change — Max/Team Premium permanent at 50% of weekly limits; Pro/Team Standard lose free access, one-time $100 credit, then API at $10/M input and $50/M output; Fable 5 shares the Mythos 5 core (Project Glasswing) behind safety classifiers; third delay since July 7; Claude Code weekly-limit boost expired same day — [Claude (@claudeai), Jul 18, 2026 announcement]; via [TLDR AI, Jul 20, 2026] and [AI Breakfast, Jul 20, 2026].
  • GPT-5.6 Sol within one point of Fable 5 on the Artificial Analysis Index at ~1/3 cost per task; GPT-5.6 family ships 5–6 reasoning-effort settings — day’s Fable 5 coverage, Jul 20, 2026; [Sebastian Raschka, “Controlling Reasoning Effort in LLMs,” Jul 18, 2026].
  • Anthropic compute leases — Meta up to $10B over two years; SpaceX/Colossus ~$1.25B/month — via [AI Breakfast, Jul 20, 2026] and CNN/TheStreet, Jul 17–18, 2026.
  • Alibaba Qwen 3.8-Max (2.4T, multimodal MoE, open-weight promised, “second only to Fable 5,” no benchmarks/model card) and Alibaba open-sourcing its AI-chip software stack to challenge CUDA — WAIC Shanghai, via [TLDR AI] and The Neuron, Jul 20, 2026.
  • Why China’s open-weight AI model Kimi K3 is sparking anxiety in Silicon Valley — South China Morning Post, Jul 20, 2026 (Moonshot, 2.8T; “narrow the gap within weeks”); Kimi K3 subscription pause via TLDR AI / The Neuron, Jul 20, 2026.
  • Open-weight models match frontier cyber performance from four months ago at a fraction of the cost — UK AI Safety Institute, via AI Breakfast, Jul 20, 2026.
  • Salesforce VP on the leaky AI pipeline: why cheaper tokens won’t fix enterprise AI — Fortune, Jul 20, 2026.
  • Agentic AI’s Real Test Is Process Redesign — Futurum, Jul 20, 2026 (Agentforce “significant value” 6.3%, down from 11.8% in Nov 2025; 13.3% at top AI maturity).
  • Evolving from legacy BI to agentic AI at Tradeshift with Amazon Quick — AWS, Jul 20, 2026 (30x faster queries, 40% lower TCO, 98% adoption after process redesign).
  • Why Some AI Workflows Stick—And Others Don’t — Every (Katie Parrott), Jul 20, 2026 (the “Island of Misfit Workflows”; workflows that stick pay back immediately).
  • Mira Murati / Thinking Machines — Inkling (975B, Apache 2.0, open weights, shipped Jul 15) and Tinker (hosted fine-tuning; you keep the resulting model); $2B seed at $12B valuation — “I Didn’t Think Much of Mira Murati Leaving OpenAI. I Was Wrong.,” Anthony Batt, Jul 20, 2026.
  • CO/AI prior issues: Lead, Follow, Or Get Out Of The Way (Jul 19 — the routers aren’t real; 2.2% of households pay for AI); I Bought the Airline (Jul 17 — Inkling and the escape hatch); The Man Behind the Curtain (Jul 13 — 89 cents of every enterprise AI dollar to the closed frontier).
Share: X LinkedIn Email
Daily Briefings

More like this

All briefings →
Lead, Follow Or Get Out of the Way
Briefing

Lead, Follow Or Get Out of the Way

Only 2.2% of American households pay for AI, and the ones that do can't make it pay. The intelligence layer is ready. The management layer isn't. Stop hand-formatting the deck. Go be the dinner.

I Bought the Airline
Briefing

I Bought the Airline

In Inception, Saito doesn't lobby the airline — he buys it. "It seemed neater." This week Mira Murati shipped an open frontier model to your own server and Elon bought a power company in the dark, while New York banned the machines and called it virtue. The doers add. The saboteurs perform.

Ice the Kicker
Briefing

Ice the Kicker

You call timeout to freeze the kicker when he's about to beat you. This week Google and Anthropic asked Washington to build an AI referee, the same week Anthropic handed free Claude to every teacher in America. Watch who's calling for the whistle, not what they're whistling for.

CONSULTING

Outsider
Labs.

A management consulting team focused on AI transformations for executives and business owners.

Work with us →