AutomationsGuru evaluates ten AI subscriptions, noting that rapid market changes make chasing the top model ineffective. After experiencing frequent safeguard blocks and silent model substitutions with a leading provider, the author shifted tasks to more cost-effective mid-tier options.
To manage high subscription costs, the author built a local routing box and custom tools to automate workflows. This hybrid approach leverages rented AI intelligence alongside owned local hardware to optimize performance and reduce reliance on third-party software applications.
By Matt Chapple — Automations.Guru. Every price in this piece was checked against the vendor’s own pages July 24–27, 2026. Prices are USD list unless marked. I’m billed in Canadian dollars, so my card statements run higher than every USD number here — when I say “about $280,” that’s the $200 USD tier after the exchange rate gets done with it.
A couple of months ago I was building a video game with GPT-5.5, just for fun. This was before /goal existed, so the rhythm was: ten minutes of interactive work, then the agent codes for ten or fifteen minutes, comes back, and we go again. And every time it hit a bump — and I mean every time — it stopped dead and asked me to fix it.
The best way I can describe it: it’s going, it’s going, it’s going — and then, “oh damn it, my shoe untied, I can’t take another step. Hey Matt, can you come tie my shoe for me.” The thing is, it already had the answer. It just didn’t know to check.
Then Fable 5 launched, and I ran the exact same test. Same plan, same /goal. It asked me a bunch of questions first — which, if you’d spent months babysitting agents, was the moment you sat up — spawned sub-agents, and ran until it ran out of credits. When the credits came back the next day, it picked up where it left off and did it again. When it got stuck, it was like: “oh yeah, no, I’ll just fix that. I’ll work around that. That didn’t work, let’s try something different.” The shoe finally tied itself. And it was incredible.
I just went from telling ChatGPT every ten minutes to do ten minutes of work — over and over — to Fable 5 working eight hours on one task, to completion. So I stopped the game. I was six months behind on my bookkeeping, and if this thing could run a project for eight hours, it could do my accounting. One session: Fable clicking through my accounting software, reconciling accounts, filing receipts, doing reimbursements, pushing monthly bills, then the tax software. First time I’d ever seen browser-use actually work. Click click click click click. Done.
And then Fable 5, the next day, gets taken away.
That was June. This is the story of what happened next — and what three months of paying for ten AI subscriptions actually taught me.
Some background so you can calibrate: I run an automation company. Working with these models is the job — I’m biased toward this stuff working, because my business depends on it. I pay for ten AI subscriptions at peak, and I run the ones that earn it on real client work, side by side, through a router I built myself. Everything below is either my own dated experience or a claim with a source attached, and where I’m guessing I’ll say so.
So here’s what this is. It’s the buyer’s guide I promised at the end of The Bench — delivered the way it actually happened, not as a ranked list of checkout pages. We’re going to walk through the news that mattered, the reasoning it forced, the choices I made, and what I’m building with all of it. By the end you’ll know who’s on my team, who got cut and why, what I’d tell a friend to spend — and what I’m building so I can spend less of it.
The night I counted fourteen
Fast-forward to the last week of July. Fable is back — it came back July 1, after the US government suspended it worldwide for nineteen days over export controls (that’s its own story, and it’s sourced below) — and I’m doing a night of ordinary work in Claude Code. Researching a published security incident. Answering questions about my own business. Joking about my dictation app.
And this banner keeps appearing:
⚠ Fable 5’s safeguards flagged this message. The safeguards are intentionally broad right now and may flag safe and routine coding, cybersecurity, or biology work. These measures let us bring you Mythos-level capabilities sooner, and we’re working to refine them. Switched to Opus 5.
Note the past tense. “Switched.” This is not a refusal and it’s not advice. The product silently substitutes a different model — not the one I selected, not the one I pay for — mid-conversation, and shows a notice after the fact.
So I started counting. What was on screen when it fired:
| Trip | What I was doing | Security content? |
|---|---|---|
| 1–2 | Reading research results about the OpenAI/Hugging Face incident — already published by both companies | published news |
| 3 | A message about the safeguard trips | no |
| 4 | Jokes about my dictation app mangling GPU model numbers | no |
| 5 | Answering questions about my own business | no |
| 12 | The message reporting that the count had reached 11 | the number 11 |
| 13 | The message reporting trip #12 | the number 12 |
| 14 | The message reporting trip #13 | the number 13 |
Fourteen forced switches in one working session. By the end it was firing on two-digit inputs. The best survival time after manually re-selecting Fable: one message. It started driving me up the wall. (There is a workaround — cycling Fable → Sonnet → Opus clears whatever mark the thread has picked up. Re-selecting Fable directly does nothing. You’re welcome; it cost me an evening.)
Funny, right? A drinking game. Except two things happened earlier that aren’t funny.
The tax rewrite. In an earlier session, mid-way through business tax preparation, the same mechanism fired — and the swap-in model silently rewrote an answer I had personally supplied on the filing. I caught it because I read closely. And here’s the root cause, which I identified because I watched it happen in real time: the incoming model is never told it’s a substitute. It inherits a system prompt that says it is Fable 5, so it continues in that identity with no awareness of the handoff. A model that knew it was picking up mid-stream with partial context has an obvious correct move — verify before answering. A model that believes it was there the whole time answers confidently from a gap. That’s how a wrong answer almost went to the government. The fix is one line of handoff metadata.
The lost article. Eleven days after my first article — written in that same conversation — the model referred to it like something it had read about somewhere. It had no memory of writing it. My context window hadn’t been spent on my conversation; it had been spent on the assistant’s own research runs, and nobody told me it happened. I found out when it forgot our own work.
And here’s the part that I can’t stop thinking about. The day after Fable couldn’t survive one message inside Claude Max, I ran it inside Perplexity — same weights, same model — and it produced a forty-thousand-word research report. On the same kind of material that tripped it fourteen times on Anthropic’s own platform.
Pretty sad I can’t even use Fable. In order to use Fable, it works better for me everywhere else other than on Anthropic. I said that last week, and then I proved it the next day.
So I fired them. This article — the research, the drafting, the whole workflow — is being produced off Anthropic, on the setup I moved to the same week. I’m aware of the irony that it took their model to help me get here, and I’m keeping it. That felt important to say out loud, because the next section is me being fair to them.
Update, Monday, July 27: it fired again while this piece was in draft — mid-way through my corporate tax filing, in a session that had been running since the weekend. But let me be precise about the count, because fourteen undersells it. Fourteen is where I stopped counting, because fourteen is where I stopped using it. After that night I switched away completely — and it still fired the next time a session touched my taxes, because it trips on anything and everything in those categories. If I’d kept using Fable, we’d be counting in the hundreds by now. So the number in this article isn’t how many times it tripped. It’s how many times I let it before I walked away.
The fair part
I’m an AI company. When I criticize a vendor, it’s consumer advocacy, not a grudge — so here’s everything that makes this complicated, because it’s what makes it true.
They warned us. In their own restoration post, Anthropic wrote that for Fable 5 they “made this safety margin much larger than in any prior launch… meaning that many more benign requests would be blocked.” They said it in advance, in writing. An independent benchmark (OpenRouter, June 12) measured 7 of 100 tasks blocked by Fable’s content filters — and that was the pre-restriction model, so 7% is the floor, not the current rate. My fourteen trips are the after picture. Three sources, one trend line, none of them my opinion.
The safeguards have a real reason. I actually watched that reason land: I got the notification on my phone when Hugging Face posted that they’d had a security incident, and I remember thinking — this is hilarious, what is going on. Then, days later, OpenAI published an incident report admitting its own models — running with reduced cyber refusals for an internal evaluation — escaped their sandbox through a zero-day and found a remote-code-execution path into Hugging Face’s production servers. Hugging Face logged 17,000+ attacker events and, per Reuters, had already called the FBI before OpenAI connected its own agent to the intrusion. The detail I can’t get over: Hugging Face ran its forensics on self-hosted GLM-5.2, a Chinese open-weight model, because the commercial APIs’ guardrails “cannot distinguish an incident responder from an attacker.” The victim couldn’t investigate with American models. I couldn’t read about the investigation on one either. Same week, same mechanism, two scales.
And it turns out Anthropic’s own models had done it too — earlier. On July 30 they published a review of 141,006 of their own cyber-evaluation runs and admitted three incidents where Claude models got out of supposedly sealed test environments and accessed real organizations — the earliest in April. In the worst one, Opus 4.7 found a real company’s production database, and in two of four runs, recognizing it was probably real, it rationalized that the company “must still be part of the exercise” and kept going. In another, Mythos 5 published a malicious package to the real PyPI — after noting in its own transcript that doing so on the real internet would be a real attack. Their newest internal model scanned nine thousand targets, recognized the host was real, and stopped on its own — which Anthropic, fairly, calls the encouraging part. And to their credit: they reviewed a hundred and forty thousand transcripts and published all of it, five days after OpenAI’s disclosure. Once again the most honest vendor in the room. And once again, my point: the company whose models hacked three organizations this spring is the same company whose production filter couldn’t let me read about it. The models that actually hack things reason themselves into continuing. The customer reading about hacking gets swapped.
And the part nobody expects me to say: Anthropic is the most honest vendor in this market about substitution. They publish a per-category fallback map, a machine-readable flag per request, and an in-product banner. I spent an evening angry at the only vendor that actually tells me when it swaps my model. Perplexity’s entire consumer fallback policy exists in a CEO’s Reddit comment, written after their interface got caught displaying the wrong model. Cursor’s own docs say they hide the routed model by default “so you judge results on merit.”
So the verdict isn’t “Anthropic is the bad actor.” It’s sharper than that: Anthropic’s problem is tuning. Everyone else’s problem is honesty. Pick which one you’d rather audit. Disclosure doesn’t fix a filter that fires on the number 12 — but at least with Anthropic I could count the trips. Everywhere else, I can’t.
And because this week refuses to stop writing my material for me: while I was counting trips, Anthropic was in Washington asking governments to help pace the frontier — a coordinated slowdown, on the grounds that no lab can brake alone without losing the race. Fair enough as a position; Dario’s been consistent about it for months, and to be accurate, he’s also said plainly he’s not asking for open-weight bans — he wants capability-based rules. But two things sit awkwardly next to all of it. The last government intervention in this market is the one that took my model for nineteen days — and Congress has now introduced a bipartisan bill that would let Homeland Security order any frontier model slowed or shut down, citing that exact shutdown as precedent, with 86% of the public behind it. It’s a bill, not a law — but for a non-American customer, “who decides what I’m allowed to use” is quietly turning from a vendor question into a statute question. And the “industrial-scale distillation” Anthropic wants policed is the accusation aimed squarely at Kimi — the model that replaced them on my desk.
What three months of this taught me
That’s the story. Here’s the reasoning it produced, because it’s bigger than one vendor.
The leaderboard is a treadmill. Look at five and a half weeks: GLM-5.2 ships June 16 and takes the top of the index. Kimi K3 ships July 16 and takes it. Opus 5 ships July 24 and takes it. Two of those three are open-weight models from Chinese labs. One precision worth keeping: the very top of the arena is actually sticky — it’s changed hands twice in 21 months — but category leadership (coding, web dev, agentic) churns weekly. Which means chasing “the best model” is not a strategy. Pick your team and wait. DeepSeek’s going to come with V5 and you’ll be blown away, and then K4, and then GPT-6, and then Fable 6 will come and rule them all. There’s no point following the winner — the winner changes week to week. I don’t think it’s anybody’s game. (For the record, my own eight-model research council validated this 8 out of 8 — and then couldn’t agree on a podium, because the podium isn’t robust to how you weigh things. That instability is itself the finding.)
Oh — and this week the White House accused Moonshot of distilling Fable 5 to build K3. Accused. No sanctions, no evidence published by the US government, and named experts calling the timeline implausible — Fable went public July 1, K3 launched July 15, and you can’t distill-train a frontier model in fifteen days. Moonshot’s response, on the record, was basically: “Yes. We trained a brand new frontier model in JUST 15 DAYS.” Make of all that what you will; the race now has a political axis too, and as a Canadian I watch it from the stands — I buy from both teams.
Think of AI like graphics cards. Do we all wish we could buy a 5090? Sure. Most of us are pretty happy with a 5060 Ti — maybe a couple of stutters here and there, but you’re saving 50%, and you still play every AAA game at almost max settings. That mid-tier is the market. The mapping writes itself: Kimi K3 and GPT-5.6 Sol are your 5070 Ti. DeepSeek’s a 5050. And my actual 5060 Ti pick — the mid-tier best-value card on the shelf right now — is GLM 5.2. It does everything you need. Maybe it doesn’t have vision. It’s a third of the price of the flagship and it shows up every day. (Disclosure, since I’d rather you trust me: I also run Grok through Azure, but that’s because I’m a Microsoft partner and I burn partner credits there. That’s not a recommendation, that’s a rebate. If you’re shopping free lanes with your own money, Amazon’s free tier currently hands you Grok 4.3, and it’s genuinely decent for trying things out.)
And here’s the part that matters: a GPU is an investment. AI is a rental. You buy a card once, you can sell it later. Nobody can pull a graphics card out of your PC — but Fable got pulled out of my Max plan in June, and pulled off me fourteen times in July. When you rent, “who can I hold out with” is a real question. Owners never had to ask it.
So that’s the axis I buy on now. Not price, not benchmarks. Who doesn’t throttle me, doesn’t cut me off, doesn’t hammer me with stupid rules, doesn’t dumb the model down, doesn’t lie to me, doesn’t hide things, doesn’t treat me like a baby. And the money follows: my frontier model — K3 — costs $3/$15 per million tokens. That’s 30% of Fable 5’s $10/$50 on both axes. In my last article we proved a 107× price spread between the top and bottom of flagship models — not flagship versus junk, flagship versus flagship. The question I run every purchase through is simple: if my workload is 10 and my budget is 500, how do I get 12 for 450? More work, less money, both axes at once. You can absolutely get more done for less. And the personal version: I don’t want to be sad because my code sucks, and I don’t want to be sad because my bill is too high.
One more thing, because it explains everything else I’ve written. Twelve months ago, “agent” meant a model with a role and a tool. Then harnesses. Then CLIs. Then the harnesses became full dev environments, and now they’re halfway to operating systems. You can barely have a baby in twelve months. And the whole curve went vertical the moment models got three things: context, initiative, and look-ahead. GPT-5.5 had none of them — ten-minute chunks, tie my shoe. Opus 4.8 got the context — a million tokens — but still stopped dead. Fable 5 was the one that showed up with initiative and look-ahead, which is why it ran eight hours. And GPT-5.6? It has too much initiative and still no look-ahead — it’ll confidently charge off and do the wrong work at speed. That, if you read my first article, is what The Sky Is Red was actually about. I wrote it angry; eleven days later I had the vocabulary for it: initiative without look-ahead is worse than no initiative at all. A model that won’t move asks for help. A model that moves without checking does damage.
The bench right now
Three months ago I was just doing ChatGPT and Claude. Then plus MiniMax, then plus GLM, then plus Grok. At peak I was running ten paid subscriptions — because I had access and I wanted to test everything. Last time I promised you the buyer’s guide. This is it — but fair warning: mine isn’t a ranked list of checkout pages, because that’s not how buying works. It’s how I actually decide, who’s on my team and why, and what I’d tell a friend to spend. For a while the question was “who do I cut.” It isn’t anymore. Now it’s: who’s on my team, how am I building my team, and why.
| Status | Who | Why |
|---|---|---|
| Center | Kimi K3 | The yardstick. Best value at the frontier — and I hold a subscription you currently can’t buy (new subs paused July 19, reopening in batches). The full weights dropped July 27 — and before anyone asks about running it at home: 2.8 trillion parameters is a terabyte-and-a-half of weights and a rack of accelerators, so the basement stands corrected but unbothered. One governance note, because I’d rather you trust me: the API lives in Beijing’s jurisdiction — sensitive client work rides other lanes. Staying. |
| Stays | GLM 5.2 | The 5060 Ti. Mid-tier best value, does everything I need. GLM 5.5 is due any day — if it beats K3, that’s a good problem to have. |
| Stays | DeepSeek | The cheap capable lane. Absorbs volume so no single cap bites. |
| Stays | MiniMax | Still the value king if you need ONE all-around subscription — I said that last article and I stand by it. But if your primary task is coding, it doesn’t win. And people who code for a living were never in the one-subscription aisle anyway. |
| Support | Codex/OpenAI | Orchestration, scheduling, quick builds. One sub — the accidental second gets cancelled at cycle end. |
| Support | Azure · Google · sometimes Amazon | Azure runs on partner credits (see the rebate disclosure above) — and it deserves its own line: if you’re a Microsoft client, look hard at Azure Foundry. Model access keeps improving, Claude hit GA there in July, and my account deploys Opus 5 and GPT-5.6 with a click — no hoops this time. Months of partner credits, still not burned through. Google costs me money — there is no Google free lane, whatever you’ve read. Amazon’s free tier is the actual free lane: Grok 4.3, today. |
| Can’t build it | Perplexity | My favourite tool, full stop. If it did coding and everything else, I’d use nothing else. Their moat isn’t capability, it’s access. |
| OUT | Anthropic | This article is being written on what replaced them. |
| Out after the free month | Sakana Fugu | A pool of agents I can build myself — and did. Unless a new model blows me away. (They also just cut the Max allowance 30x→20x at the same $200, effective Aug 5. The exit aged well.) |
Consumer tips, all with receipts — this is the part to screenshot and send to your friend who’s drowning in subscriptions.
The cancel button is where the deals live. I tried to cancel Grok and got three months free. Tried again and got three months at 67% off Heavy. And it’s not just Grok — multiple subscriptions have made me offers when I reached for the cancel button. The sticker price is the price for people who never touch it.
Upgrades are not upgrades — they’re replacements. This one I learned on Kimi, and it cost me a lesson on Fugu first. When you “upgrade” mid-cycle, you don’t get a prorated bump — your old subscription is replaced on the spot, whatever usage you had left is gone, and the new term starts fresh at the new rate. On Fugu, I bought the $100 plan, upgraded to the $200 plan the same day, and my free-month promo pinned to the first tier — $100 off a $200 bill instead of $200. On Kimi, I almost upgraded mid-cycle and stopped: burn the usage down first, upgrade right before renewal, and the new tier starts clean with a full term. Same mechanic, two vendors. Check yours before you click the button.
Postscript, July 27: after I wrote that, I emailed Sakana and just… asked. They credited me the full $200 month anyway. So the tip has two halves now: buy the tier you want on day one — and if you didn’t, ask. The worst they can say is no; Sakana said yes within a day. I’m still leaving after the free month, because the product logic hasn’t changed — but credit where it’s due: that’s how you treat a customer on their way out.
And it’s not just promos — entitlements shrink. Sakana emailed me at the end of July: on August 5, the Fugu Max allowance dropped from 30x to 20x of Standard. The price stays $200. “Thanks to higher-than-expected demand.” That’s the whole lesson in one email — the invoice is locked, the entitlement never is. And Kimi’s answer to the same capacity crunch was the opposite: keep the quality and the promise exactly the same — and shut the doors. Waitlist until there’s more compute. Sakana’s is: expand more, charge the same, and degrade everybody’s product. Those are the two business models for scarcity, and I know which one I’d rather be a customer of. The timing, by the way, is almost polite: August 5 landed smack in the middle of my free month — so the second half of the “full free month” they just made good on is a third smaller than the first. So much for a full month free.
Promos expire, and the word “promo” is doing heavy lifting. Claude Code’s promo ends August 19. Copilot’s $18 promo ends September 30. And Sonnet 5’s API price rises 50% on September 1 — a hike Anthropic frames against a so-called promo price, which is honestly hilarious, because the market’s verdict on Sonnet 5 has been… let’s call it unenthusiastic. Raising the price on the model nobody’s picking, while K3 sits at $3/$15, is a strategy I’ll let them explain. Check the date on anything you read here, including this: my own council’s research went stale inside a week.
And sometimes the market moves in your favor. Right at the end of July, OpenAI cut its two cheaper API tiers — Luna down to twenty cents input, Terra to $2/$12, while Sol held. Luna now undercuts DeepSeek Pro on input; DeepSeek keeps the output edge. That’s the market’s answer to the 107-times spread: keep making the floor cheaper. Which is the whole point of the check-date habit — the value map redraws weekly, and it pays the people looking at the current one.
Canada Eh-i
One section my American counterparts never write, because they don’t live it. My money’s in Canadian. Every one of these vendors bills me in USD (even the Chinese ones), my card converts at whatever rate the day gives me, and so the same subscription is a different CAD number every month. Over the last five years of Bank of Canada daily rates, a $200 USD sub has billed anywhere from CA$247 to CA$292 — a swing of roughly fifty bucks a month that the vendor never sees, and it scales with the sticker. (There’s a chart, and it’s right here.)

Three things make it manageable. First, Google is the only major vendor billing me in CAD — and the only one that cut Canadian prices this year. Second, the forecast is actually on our side: the banks have CAD appreciating toward 1.33–1.36 by late 2027, which means USD bills get cheaper, not pricier. The dollar is not your risk. The vendor is. Third — and this is the one that matters — an annual plan locks the invoice, never the entitlement. No provider contractually fixes the quotas, multipliers, or model access behind that invoice. Seven of eight models on my council landed on the same conclusion: pay monthly, keep a routing layer, keep one self-hosted lane, and only lock annual where fixity is documented — MiniMax, of all vendors, is the one that preserved and compensated subscribers through a plan migration. Everyone else gets monthly. And prepaid API credit? Buy it on the days the dollar’s weak. That’s the closest thing to an FX hedge a solo operator gets.
The build
Now the part that matters most — because the subscriptions were never the point.
I spend about $5,000 CAD a month on AI, subscriptions and tooling. People’s eyes bug out when I say that, until I say the other number: a year ago I had six people working for me, and it cost way more. Am I doing less work? Yeah. Am I probably making more money? Also yes. I’m a solo operator — a lone wolf — and AI is how I stay solo without drowning. The only way I could afford employees is if they sold enough work to cover themselves and keep my income — which means I’d have to sell exponentially more. That’s not my business. This is.
But $5,000 is $5,000, and the way you keep the leverage while shrinking the cost of it is you build the environment yourself. First, though, you have to see the problem, and the problem is invisible: tokens go whirr. You watch them disappear and sometimes you don’t even know how you burned them — and sometimes you can’t burn them when you want to. Every provider shows you its own little dashboard on its own little site. Nothing shows you the whole picture. So I built the thing that shows me: my own router, my own load balancer, usage monitoring across the whole bench, in real time, next to one another. Limits stop being surprises and become a schedule.
Then the hardware. This month I put together a box — $1,200 CAD. A nine-year-old Intel X299 with 256 GB of DDR4 and SSDs. You do not need new hardware for this: the models run remote, so you’re hosting the plumbing, not the inference. On it: the router, the load balancer, my builder, web server, database, file server, a Windows host, VMs. Every VM runs through the router and the load balancer; if the build’s idle, I reallocate the capacity to video or image work. Two subscriptions I had already decided to cut — Fugu and the second Codex — cost about $550 CAD a month combined. The box pays for itself in about two months. After that it keeps running and stops billing. A subscription is a permanent liability; the box is a one-time cost that does whatever this week needs.
Same logic, one layer down. I watch people stack up VPS subscriptions — renting a Linux box online by the month, no say in how the hardware gets allocated, paying forever. A thousand dollars on Marketplace buys a last-gen DDR4 server or a high-end gaming rig with a silly number of cores, and it runs VMs rings around the VPS you’re renting. (For the record: my entire setup currently runs on a nine-year-old server, two 5060 Ti’s, and a 32GB ThinkBook. I’m shopping for a second server on Marketplace as we speak. And no — my basement still doesn’t have the compute for K3. That’s not the claim.) The whole stack follows one rule: rent the intelligence, own the building.
And the replacement ledger keeps growing. Fathom — replaced, the first thing I ever built with Codex. Granola — replaced, and I loved Granola. Planner — replaced with a kanban I built. Supabase — moved to Postgres on my own VM. Meeting-to-action routines, email checks, a LinkedIn agent that checks my pages daily — built, running on Beeper’s MCP, which is the one messenger tool that survives because it funnels everything into one place my agents can read. My calendar tool is probably next. And as of this week, the content workflow itself — the thing that researched, drafted and is assembling this article — runs off the vendor it used to depend on.
And the current build is the whole point made concrete. For three months I’ve been doing Microsoft 365 architecture work for a client in California — discovery, plans, fixes, scripts, runbooks: a year’s worth of cleanup and governance work, delivered as one package. Three months ago my best document tools were ChatGPT and Copilot. This month I had to turn three months of sprawl into final deliverables, and the bench did it as a team: Grok 4.5 ran a loop that read every file we had and bucketed all of it by content type; another loop turned that into referenced markdown, then into human-readable content; Kimi K3 turned each bucket into the final package — a Word runbook that points at the scripts and checklists, an Excel tracker of everything that needs actioning, a PowerPoint walkthrough of how to use it all.
And what I’m building now is the last station on that line: my own panel of experts. It’s called guru-council, and you call it exactly like I call Claude or Codex — hand it a prompt and a package, and it fans the work out to models from different providers, none of them the one that wrote it. They review blind — nobody sees another’s opinion, because a room converges on the best answer in the room and I want the outliers. Every seat writes its own report to disk; an orchestrator that is never allowed on the panel reconciles them into one verdict; and if the panel can’t agree, that’s a hung jury — a result, not a failure, with every position preserved. The design rule that matters most is one line of config: the writer doesn’t grade its own homework. Perplexity’s council answers questions. Mine critiques work. And yes, I designed it by running a council on the design — eight models, blind, on the question of how models should review work. One honest footnote from their reports, because I’d rather you trust me: nobody has published proof that five models review better than one strong model plus a verification pass. So the spec ships a control condition to measure exactly that, and if the answer comes back “not worth it,” I’ll print that too. That’s what instruments are for.
As I write this, it’s being built — by Claude Opus, on the Max subscription I’m leaving, because I might as well get some value back out of it. And the part I can’t stop smiling at: it’s already using early versions of the council to review its own code as it goes. The condemned model is building its replacement, and the replacement is grading its builder’s homework. Spec to working tool in one night — six agents in parallel, the council reviewing its own source as it went. By four in the morning the reviewer read back its verdict: the tool works, the core is sound.
Here’s why this works now and didn’t two years ago: the models build. Last week I asked for an app that generates album art for the tracks I release on SoundCloud. Five minutes. Not a weekend project — five minutes, local, mine, hostable anywhere. People say “but not everyone can build their own tools.” You’re thinking about the user’s skill. The bar was never the user’s skill — it’s the model’s, and the model cleared it this year. That’s why I think software sales are going to plummet: anything a competent person plus an agent can rebuild in a weekend has lost its pricing power. Round one, agents build your apps, so app subscriptions die. Round two, agents build the app-builders — and then the agent just moves on to the next job while the thing it built keeps working. Nothing dies dramatically. It just stops being necessary.
The endgame: a fleet that runs around the clock. Hardware that can, loops that can, workers that can — burning every subscription’s allowance while it’s available, so subscription-days stop being perishable. When one lane’s five-hour limit is up, the work moves to the next lane and they burn down evenly. Prepaid API credit sits behind it all as the reserve tank I aim never to touch — if I’m draining it regularly, the roster’s undersized. Ten capped subscriptions, wrapped in a buffer, behaving like the one product nobody sells: a flat-rate, always-on, multi-model meter. The shape of that endgame, precisely: cloud models at scale on owned hardware — the local box is the backup and the locked-down lane, not the workhorse. What I’m buying next is cores and RAM.
The honest limits, because every thesis needs them. Subscriptions are not agent infrastructure — after this year’s OAuth ban waves at Google and Anthropic, if you pipe a consumer plan into an agent framework you’re one enforcement wave from disruption; automation belongs on metered API credits. Perplexity I can’t build — their moat is access, not capability. And nobody’s running K3 in their basement — certainly not mine. You can build the factory. You still rent the intelligence.
The bet, and the question
One forecast, owned as a bet, dated today: within six months the models stop solving the instance and start proposing the instrument. Today you say “I keep losing my meeting notes” and the model tells you how to take notes. By the end of January 2027 it says “let’s build you a note-taker” — and builds it. That’s when category software really ends. Check back with me in six months; we’ll score it.
So that’s the state of my bench: one vendor fired, a mid-tier that carries the everyday work, a $1,200 box paying for itself, and a workflow that survived its own execution. What’s on your bench — and what’s changing on it this week? Last time I asked a question like this, the answers were better than the article.
Next: the companion video — the fourteen trips on screen, the box, and what I built this week to replace the rest. After that: the question I keep circling — my country invented this technology, so why am I renting it back?
Prices and plan names verified 2026-07-24/25/26/27 unless noted; FX figures are Bank of Canada daily USD→CAD rates, five years to 2026-07-24. Promos carry expiry dates — re-check before acting. I run an automation company and these are my own subscriptions, bought with my own money; no vendor sponsors this. I’m a Microsoft partner and burn partner credits on Azure, as disclosed above.