I Am (Politely) Begging You to Start a Business Before It's Too Late
Better models and competing suppliers are making more work affordable to attempt. I think the bigger advantage will belong to people learning what customers will actually pay for.
I'm politely begging you to start a business before it's too late.
I think the gap is going to widen between people learning to use AI well and people sticking their heads in the sand. Not just in what they know, but in what they can build, deliver and earn. Starting a business won't suddenly become impossible. Competing with people who already know how to serve customers could get a lot harder.
If your opinion of AI is still based on a conversation you had with ChatGPT in 2023, you are making a current decision with old information. Meanwhile, somebody else is finding out what the newer tools can do, where they fail, and what a customer will pay them to handle.
I build software and AI systems for a living. I have an obvious interest in this market growing. But my advice isn't to buy my product, quit your job or become an AI influencer. It's to start learning how to turn a problem into something useful enough that another person pays for it.
Every real assignment can teach you something a new subscription doesn't include: what the customer meant, what the model missed, and what it took to finish. That is the advantage I'm worried about people postponing. Better models can arrive overnight. Your experience using them to serve somebody doesn't.
KB
The opinion that expired
The first public ChatGPT arrived on November 30, 2022, based on the GPT-3.5 series. Its launch page described a conversational system that could answer follow-up questions and challenge incorrect premises. It also plainly warned that the model could produce convincing nonsense.
That was a useful tool. It was not the same experience as giving an agent access to a project, having it inspect the files, make changes, run tests and return something you can examine. Early builders could connect models to tools, but they had to supply the surrounding machinery. A chat answer did not execute itself.
Think about a small business with a website that looks fine on a laptop but makes booking painful on a phone. A chatbot can suggest a better headline or explain how a button should work. The owner still needs someone to inspect the actual site, change it, check mobile behavior and make sure the booking path survives.
That gap between advice and completed work is where the model improvements start to matter.
In March 2024, Claude 3 brought a family of models with image understanding and a 200,000-token launch context window. Context is the material the model can consider at once. A builder could supply more of the project and show a screenshot instead of trying to describe a broken layout. The useful test became whether the model could find the relevant problem in that material.
By the Claude 4.5 generation, the releases were increasingly about sustained work. Sonnet 4.5 launched in September 2025 alongside code checkpoints, context-management tools and an agent development kit. Anthropic reported observing it maintain focus for more than 30 hours on complex tasks. That is a vendor observation, not a promise that every task can be left alone for 30 hours.
The useful change was the combination: better coding and computer use, plus software that lets the model act, inspect the result and continue. Opus 4.5 followed in November, with stronger results on Anthropic's software-engineering evaluations and a lower Opus price.
Back at that small business, the assignment can become more concrete. Inspect this booking journey. Find the confusing steps. Implement the agreed change. Test it at phone width. Show me what changed and what still doesn't work.
Someone still has to judge the proposal and the evidence. But the person directing the work no longer has to perform every supporting step manually.
More than a better answer
September 2026 brought another useful pair of examples: GPT-6 Astra and Claude Opus 5.5.
OpenAI's September 3 Astra system card documents the release. Its announcement describes improvements across software engineering, browser and computer use, including creating websites and running frontend checks. The same booking-site assignment now involves both changing the code and checking the experience in a browser. That is a much more useful ambition than another explanation of what the owner should do.
Opus 5.5, released September 22, makes a related argument about efficiency. Anthropic and its early testers report better performance on complex coding and professional work, with fewer steps or tokens in their evaluations. It advertises a 40% reduction in typical workload cost compared with Opus 5. That is not a universal discount on every job: ordinary input and output token prices fell 20%, while cache-read pricing and the amount of model work also changed.
I care about that difference. A model that needs fewer failed attempts can be worth more than one that simply charges less each time it tries.
This also isn't a clean ladder where every new model is better at everything. A replacement can interpret an instruction differently, require a different tool interface or be less reliable in a workflow that worked with its predecessor. The name on the model picker doesn't test your business process for you.
I recently had an AI-generated job-seeker ad come back with a woman's reflected torso disappearing halfway down a mirror. The headline was readable. The image was attractive at a glance. The reflection was still broken.
Being able to produce work and being able to recognize acceptable work are different capabilities. I want both. I don't want somebody selling a client the first confident output because the model has a bigger number in its name.
The price of trying
Even an impressive agent is hard to build a business around if every attempt destroys the margin.
An agent might read the same instructions repeatedly, inspect files, call tools, reconsider a result and try again. You pay for more than the final paragraph or screen that the customer sees. This is why the price competition from DeepSeek, Moonshot's Kimi and Z.ai's GLM matters to me.
They give builders alternatives. A US lab isn't competing only with the next US lab. It is competing with suppliers offering capable coding, reasoning and tool-use models at different prices and operating conditions.
This pressure isn't something I'm inventing after looking at a rate card. In January 2025, Reuters reported Sam Altman's response to DeepSeek R1: he called it impressive, particularly for what it delivered at its price.
I think that competition makes a premium harder to defend without a useful difference in the work. It does not prove that DeepSeek, Kimi or GLM caused every subsequent OpenAI or Anthropic price cut. Better serving efficiency, product strategy and other competition matter too. Anthropic explicitly attributes part of Opus 5.5's lower pricing to needing less compute to serve it.
The price changes are substantial even without assigning one cause. At launch, Claude 3 Opus cost $15 per million input tokens and $75 per million output tokens. Opus 4.5 launched at $5 and $25. Opus 5.5's ordinary rates are $4 and $20. Those are historical list-price comparisons, not a claim that the models use identical token counts or do identical work.
Here is a more practical illustration. Suppose a collection of tasks uses one million uncached input tokens and 100,000 billable output tokens across its requests. Using the published ordinary rates checked on October 10, 2026:
| Model and pricing condition | Input per million | Output per million | Illustrative token bill |
|---|---|---|---|
| DeepSeek V4.1 Flash, peak hours | $0.30 | $1.20 | $0.42 |
| Kimi K2.7 Code, standard speed | $0.95 | $4.00 | $1.35 |
| GLM-5.3 | $1.40 | $4.40 | $1.84 |
| Claude Opus 5.5, standard mode | $4.00 | $20.00 | $6.00 |
Sources: DeepSeek pricing, Kimi pricing, Z.ai pricing, and Anthropic's Opus 5.5 release.
This is arithmetic, not a benchmark. The basket excludes caching, separate tools, hosting and taxes. Any billable reasoning tokens must fit within the stated output total or increase it. Different models may need different token counts and attempts, and the cheapest one might fail the task. DeepSeek also has lower off-peak rates; I used its higher peak rates rather than making the comparison depend on a discount window.
Nor does every Chinese flagship belong in the cheapest tier. Kimi K3's current ordinary input and output rates are $3 and $15 per million, with separate cache terms. The lesson is to compare the actual model and job, not buy a nationality-based story about price or quality.
But having these choices changes what is affordable to test. You can reserve a stronger model for the hard judgment, use a cheaper model where it passes your checks, and use ordinary code where a model isn't needed. You do not need the most expensive intelligence to move a known value from one field to another.
For a small builder, a lower bill can buy another attempt before the budget runs out. Whether that attempt teaches you anything is up to how you use it.
The work around the model
Reviewing the research in our cYpher.camp intel archive, I kept finding a second change alongside the model releases. More of the machinery for doing ongoing work is becoming something you can use instead of invent.
Cursor's September Projects announcement describes shared context that persists across work, a coordinator that delegates tasks, and subscriptions that can react to events or schedules. It launched in beta. I am interested in the direction: less re-explaining the project each time, more continuity between one piece of work and the next. For the booking site, that means the next change can start with the previous instructions and findings instead of another introduction to the business.
A different entry led me to Polylane's engineering account. They went the other way on agent count, replacing a chain of specialists with one investigator. They reported average model spend per pull request falling from $111 to about $18 during the first nine days of the new setup. They also said other changes contributed. A pull request is proposed code, not a deployed fix or a profitable customer outcome.
I like the combination of those stories because it prevents an easy sales pitch. Sometimes parallel workers help. Sometimes the handoffs lose information and make the job more expensive. The business advantage comes from understanding which arrangement works, not from announcing that you have more agents.
Return to the booking website. A small operator can now attempt the research, redesign, implementation and testing with help across several disciplines. Previously, that combination could have required more people, more money or more time than the project could support. Not literally impossible before, but impractical for that person with that budget.
Take away enough supporting work and a person can attempt something they otherwise would have postponed. Somebody still owns permissions, payments, customer data and what happens when the site breaks. Those responsibilities don't disappear because implementation gets easier.
I was already an engineer before these tools improved. AI did not give me my whole profession in a prompt. It gives me more ways to apply what I know and more help with the work around it. Someone from another field brings a different advantage: they may understand a customer's problem far better than I do.
The gap I worry about
None of this means every AI user is more productive.
METR's early-2025 study found experienced open-source developers took longer with the tools it tested. In its February 2026 follow-up, METR said changing participation and task selection made the new productivity estimate unreliable. The researchers thought newer tools were probably helping more, but the data did not establish a dependable size for that improvement.
That is a reason to measure your own work, not to freeze your opinion at whichever headline pleased you.
My concern is the difference between someone doing that learning and someone refusing to begin. One person finds out which problems buyers care about, which model can help, which output needs correction and how much delivery actually costs. The other person waits for a tool so perfect that none of those judgments will be necessary.
Over time, the first person can accumulate something a new subscription doesn't include: customer trust, useful examples, a tested process and a better sense of what not to promise. When the next model arrives, they have work ready to try it on.
There is a counterforce. Better tools can help a late entrant catch up, and established businesses can use them too. Early adoption by itself is not a moat. A year spent generating things nobody wants is not a head start I would envy.
But useful experience can compound. That is the gap I think will widen. Knowing model names is not mastery. Repeatedly getting worthwhile work done, noticing failures and improving the process gets you closer.
A customer before a company fantasy
I know the temptation to keep building. I would usually rather improve the software than go sell it. Cheaper implementation can make that habit worse because now I can build the unnecessary feature faster.
So when I say start a business, I mean start with a customer problem. Not a logo, a complicated company structure or an agent that spends the weekend creating more tasks for itself.
If you have no idea what people will pay for, ask yourself what problem you would pay to solve right now. Then find other people who would pay to solve it too. Your own frustration is a useful lead. It is not proof of a market.
For the website example, talk to the owner before rebuilding anything. What should a visitor be able to do? Where does the current journey fail? Would the owner pay to have that specific problem fixed? A free preview can help them judge your work, but compliments on the preview are not the same as a buying commitment.
Make a small promise you can keep. Agree on what the paid work includes. Deliver it, check it with the customer, and count your time and revisions alongside the software bill. Then decide whether you can offer it again at a price that leaves you something after the costs.
You can do this alongside a job. You can start with a service instead of a software company. You can learn that the idea is wrong and change it without betting your rent on it. As I argued in The Incubation Period, having room to experiment matters. Not everyone has the same money, time or safety net, and entrepreneurship is not a substitute for decent employment or public support.
I still want more people to try. Better tools are making a wider range of attempts practical. Competing suppliers are making some of those attempts cheaper. Neither development gives you the experience of finding a buyer and doing the work.
You have to start accumulating that yourself.
Start now, while you can afford to learn small.
Agentic & distributed systems, DeFi, and the compute economics. One email a week, no fluff.
Subscribe to the newsletter →About the author
Keenan Benning is the founder of cYpher.camp, CTO of DeFi All Odds, and a forward-deployed AI systems engineer. He builds AI systems and writes about the engineering, economics, and ownership questions behind them.
Other projects