Last week Anthropic clarified in their documentation that OAuth tokens from Claude subscriptions can’t be used in third-party tools. Period. Not Claude Code wrappers, not the Agent SDK, nothing. OpenCode removed Claude support the same day, citing “anthropic legal requests.” OpenAI immediately jumped on it and endorsed Codex in third-party harnesses. Classic.
The enforcement had started back in January. Account bans. Server-side blocks. People who built entire workflows around Claude woke up to broken pipelines.
I get why Anthropic did it. Subscription arbitrage is real. People were routing $1,000+ worth of API calls through a $200/month Max plan. That math doesn’t work for them. Fine.
But that’s not what kept me up that night. What kept me up was realizing how dependent I’d become.
How We Got Spoiled
Claude is the model that best understands how my brain works. That’s not a benchmark result. It’s not even a technical claim. It’s the feeling you get when something clicks—you say something messy and half-formed, and it gets what you meant instead of what you said.
ChatGPT has its own character. I never clicked with it the same way. DeepSeek R1 was exciting when it first came out. The other day I tried StepFun’s Step 3.5 Flash and it handled my thinking better than I expected. GLM surprised me too. Each of these models has a personality. Not in a marketing sense. In how they catch things other models miss, or fumble things other models nail.
I think this matters even more when English isn’t your first language. You’re already translating in your head before you type anything. Some models bridge that gap. Others make you feel like you’re talking to a wall. I don’t have research to back this up. It’s an observation from someone who thinks in Turkish half the time.
So you find the model that gets you. And you lean into it. You write prompts that work with its quirks. Your workflows start assuming it. Before long, your muscle memory is wired to one provider.
Then the rules get clarified.
Vendor Lock-In, New Medium
This is the same problem we’ve had in software forever. You build on a platform, the platform changes, your shit breaks. The only difference is that with AI models, the lock-in is cognitive. It’s not a database migration you can plan for over a weekend. It’s how you think and communicate with the tool, baked into your daily habits.
And the vendor can clarify the rules at any time. Anthropic started enforcing in January and spelled it out in a docs page in February. Not even a press release. A docs page.
That scares me more than any model going offline. The rules are whatever they say they are, whenever they feel like saying it.
Structure for the Dumb Models
The answer is not “stop using Claude” or “stop using smart models.” I’m not going to pretend Opus isn’t good at what it does. The answer is stop relying on it for everything.
Agentic workflows already point in the right direction. One orchestrator. Many workers. The orchestrator needs to be sharp—it holds the plan, makes judgment calls, handles ambiguity. But the workers? Extraction, formatting, classification, validation. You don’t need Opus for that. An 11-billion-parameter model handles it fine.
The harder part is what happens between the human and the orchestrator. I ramble in half-sentences and switch between Turkish and English mid-thought. A smart model parses that without breaking a sweat. Throw the same ramble at something smaller and you get back garbage.
So I’ve been thinking about this: use the smart model as a translator. Not between languages—between messy human intent and structured instructions that a dumber model can follow. You brain-dump into Opus. Opus converts that into clear, discrete steps. The steps get handed to cheaper, smaller models that execute them.
I tried building a formal language for this. Called it PDL—a prompting definition language with typed variables, control flow, the works. It added overhead. The model spent tokens parsing syntax instead of doing the work, which defeated the whole point.
Pseudo-Python flows worked better for deterministic outputs. Define a schema, define the steps as methods, let the model “execute” the class. But for anything that needs nuance—judgment, tone, creative decisions—it flattens the thinking.
I haven’t cracked it. The simpler version is probably the right one: talk to the smart model in your natural mess, let it rewrite your intent as clean instructions, then farm those instructions out. But I haven’t battle-tested that at scale yet. Being honest about it.
Models are getting smarter. Smaller models are getting better at understanding natural language without structured prompts. Part of me thinks I should wait for the 11B models to catch up instead of building translation layers.
But the rules changing isn’t a maybe. It’s a when. And the only insurance against that is not needing any single provider to keep your work running.
I don’t know the exact shape of this yet. The dependency is the problem. The smart models made it too easy to ignore that. And a documentation update on a random Wednesday shouldn’t be able to break how I work.