Skip to content

Which Hard Things Are Worth Doing?

· · 13 min read · Updated
Which Hard Things Are Worth Doing?

AI can finish the work without helping us grow. I want tools that make harder things possible without taking away the struggle, judgment and relationships that make them worth doing.

I don't want AI to take all the hard things out of my life.

I want it to clear away work that gets in the way, so I can attempt things that used to be out of reach. Not do the learning, thinking and living for me.

I came to that distinction through the argument about whether AI could kill us all. The warnings made me question who gets to decide our future. They also made me question what I want from the tools I'm already using and selling.

KB


The warning

When Jacob Coxon quit Anthropic and warned that the AI race was gambling with our lives, my first question was whether he'd secured his stock payout before leaving.

I had the timing backward.

Coxon told Axios he quit two months before his Anthropic equity would have vested. I can't accuse somebody of waiting for the payout when the reporting says he walked away from it.

Jacob Coxon @hilbertspaess ·

I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

Read the original post on X

I still want to follow the money. I just can't use my suspicion as evidence.

Axios reported that Coxon still held equity in OpenAI. Having a financial stake doesn't settle whether his warning is right. Demanding that every whistleblower go broke before I'll listen would give the people still inside another reason to stay quiet. The Right to Warn letter raised that problem in June 2024, calling for protected reporting and protection against retaliation affecting vested economic benefits.

The people running the labs warn us too, while their companies keep building.

Sam Altman, Dario Amodei and Demis Hassabis signed a statement in May 2023 putting the risk of extinction from AI alongside pandemics and nuclear war. The following March, Anthropic released Claude 3, promoting a new level of capability.

Fear can help sell a company as much as optimism can. Telling me your technology might save the world or end it also tells me how powerful you think it is. If a former employee warns about dangerous AI and then raises money for a safer alternative, I want to examine that pitch too. That's a question about a possible next venture, not a claim that Coxon has one. Neither a financial stake nor a later release proves a warning dishonest.

None of this makes the danger imaginary. In WIRED, Coxon pointed to biological threats and cyberweapons. He called Anthropic the most responsible player and said it wasn't cutting corners yet. His concern was where the race would take it.

The consequences don't stop at the shareholders. Canceling a subscription doesn't opt anyone out of a catastrophe. Workers need protection when they report danger, and independent oversight needs the authority to stop dangerous deployments. That authority needs limits too. I don't want safety rules turned into another way for the biggest labs to keep competitors out.

The people exposed to the risk need power over the decision, not another assurance from the people who own the upside.

If you believe these systems pose an unacceptable risk, advice about using them better won't answer that objection. Neither will my intentions for cYpher.camp. Useful applications don't excuse dangerous development, and you don't owe the industry your participation.

The question I want to explore here is narrower: where we do choose to use AI, which parts of the work should we keep for ourselves?


The work

That's what interests me in the pro-human argument Cal Newport and Brad Stulberg discuss in The AI Resistance is Forming. It brings the question back to the person using the machine.

I sell AI tools. But a finished result isn't always everything a person wanted from doing the work.

I can finish an essay without working out what I believe. I can have functioning software without understanding why it works. I can send an apparently thoughtful message without giving the person much thought.

Before I decide whether to delegate, I need to know what the activity is for. Am I trying to deliver a result, develop an ability, understand something, or spend time with another person?

If I'm learning Spanish, the awkward conversation is part of the point. A perfect translation can help me communicate, but it doesn't prove I learned to speak. If I'm trying to explain an emergency to somebody who speaks Spanish, communicating is the point. I'm not going to insist on a vocabulary exercise first.

Same tool. Different purpose.

That is where the question about hard things becomes practical. I don't need to defend difficulty for its own sake. I need to recognize when removing it also removes something I wanted to gain.


The climb

The mountain comparison in Newport and Stulberg's discussion makes this clear. A gondola can take you to the summit. It cannot give you the training you would have done on the climb.

If the goal is the view, take the gondola. If the goal is to become capable of climbing the mountain, it doesn't do the same job.

That isn't a moral ranking of the people at the top. Somebody may not be able to climb. Somebody may want to enjoy the view with their family. They don't owe anybody a harder trip.

But I also can't outsource the climb and claim the adaptation.

A workout plan isn't a workout. Knowing the right answer isn't always understanding it. The uncomfortable attempt, the mistake, the adjustment and the next attempt are often how I develop the ability I wanted.

There is evidence for taking that distinction seriously. A randomized study of high-school mathematics found that a standard GPT-based assistant improved performance while students had access to it, but hurt their subsequent unassisted performance. A version designed with tutoring safeguards largely mitigated that harm.

One study in one school doesn't settle how everybody learns. It does show why getting better answers during practice isn't sufficient evidence that the student is learning.

I don't read that as a reason to ban help. A good teacher helps. A worked example can help. Feedback can save somebody from repeating the wrong thing for months. The distinction is whether the help makes my next attempt possible or makes the attempt unnecessary.

If I keep taking the finished answer, I may miss more than practice. I also miss making the choices that teach me what good work looks like.


The judgment

Writing makes that especially visible. Choosing a word, cutting a paragraph and deciding what I can defend are part of developing judgment. If I accept the polished version without examining those choices, I don't know what I've improved.

A sentence can become smoother and say less. A paragraph can sound authoritative while making an argument nobody has examined. Everything is arranged correctly. Nothing sounds like a particular person needed to say it.

That's a real risk for this blog. If a tool makes every unusual sentence conventional, I might end up with something easier to read and less worth reading.

I don't buy the absolute claim that models can only produce mediocrity. A study of AI-assisted short fiction found that access to AI ideas improved ratings of individual stories while making the stories more similar to one another. Better individual scores and less collective variety can happen together.

That gives me something specific to watch for. When I ask for an edit, does the suggestion make my meaning clearer, or replace it with a familiar expression? Is the argument better, or does it just sound finished? Can I explain why I kept the change?

Human editors can flatten a voice too. But a tool that can instantly rewrite the whole piece makes it easy to accept a replacement without working through those questions.

And encouragement isn't scrutiny. A machine telling me my idea is interesting isn't evidence that it is good. I would rather have it point to a weak assumption than congratulate me and build another paragraph on top of it.

My responsibility doesn't disappear because I clicked approve. If my name is on the work, I need to know what I'm saying. That includes deciding whether it needed to be said at all.


More output

At work, that last decision can get lost before anyone opens an AI tool.

We already know how to produce activity that looks important. Status reports nobody uses. Meetings to review the report. Presentations explaining the meeting. A second summary because the first one was too long.

AI makes producing that material easier. It doesn't supply a reason for it to exist.

Imagine a team generating a long update, sending it to people who use AI to summarize it, then generating replies to the summary. Everybody appears busy. It is still worth asking whether anything changed for the customer.

Before I automate a report, I want to know who uses it and what decision it changes. If nobody can answer, maybe the improvement is deleting the report.

That applies to agents too. A hundred completed tasks can mean nothing if they were the wrong tasks. Ten content drafts aren't a campaign that brought in customers. A folder of application letters isn't an interview. Those artifacts may help, but they aren't the outcome.

I want to look at what the work accomplished and what the person got from doing it, not just count the artifacts. That might be a skill, a useful result and an afternoon back, or the chance to attempt something they couldn't before.

The answer cannot be to preserve all the old effort. A pointless report doesn't become meaningful because somebody suffered through writing it by hand. I want to remove that work so there is room for something better.


Harder things

I don't want to take the satisfaction out of difficult work. I want more people to be able to attempt work that used to be out of reach.

Suppose somebody has an idea for a small application but can't afford a development team. AI helps them get a first version running. That's a meaningful change. They can test the idea instead of spending another year imagining it.

Now there are harder questions. Does it solve the problem? What happens when it fails? Is it safe to use? What should change after somebody tries it?

Some of those questions can be explored with AI. Some require learning. Some require an experienced person. Having a generated first version doesn't remove responsibility for the answers.

I don't need that builder to type every line unaided to respect the accomplishment. Engineers already use libraries, documentation and other people's work. But if the builder wants to become an engineer, they also need practice understanding systems, diagnosing failures and making decisions they can defend.

Getting a first version running should give them something to investigate, not a reason to stop learning.

For a business owner, the same website may serve a different purpose. They might want to spend more time learning their trade and serving customers. Becoming excellent at web development doesn't have to be their goal. They still need a site they can trust, but they can seek help evaluating it without making software engineering their life's work.

We shouldn't require everybody to master every supporting skill before they're allowed to pursue the thing they care about.

Nor should we romanticize exhaustion. For someone balancing work, family, illness or a disability, assistance can make participation possible. Telling them that every shortcut diminishes the experience ignores the difficulty they are already carrying.

I want people to have more choice about where their effort goes. And that choice has to include something other than becoming better at work.


The people

Not everything worth doing needs to lead to mastery. I can cook badly, play a game or make art because I enjoy it. If I turn every saved hour into another improvement target, I haven't made much room for a life.

Relationships make the limit even clearer. An apology isn't just a writing assignment. The other person may need me to understand what I did, listen to them and act differently. A beautifully written message doesn't substitute for that.

That doesn't mean help finding words is dishonest. Someone can need translation, accessibility support or a way to organize thoughts they genuinely hold. I care whether the assistance helps them participate or becomes a way not to participate.

The same goes for learning from other people. Asking a model about equipment may be convenient. Asking a coach can start a conversation in which they notice something about how I move or what I need. If I replace every small interaction because the machine is easier to reach, I may lose more than advice.

The human option isn't always available, affordable or welcoming. But convenience shouldn't make the decision before I've noticed the tradeoff.

This is what protecting friction, craft and emotional depth means to me. Keep the effort through which I learn, make choices and connect with people. Remove the obstacles that prevent me from doing those things. Leave room to enjoy something without measuring improvement at all.

That gives me a way to judge whether the tools I build are helping, rather than assuming that less human involvement is always progress.


Human success

I want cYpher.camp to be judged by what happens to the person using it.

Self-actualization is the bigger ambition: helping people develop their abilities and pursue a life they care about. I can't deliver that as a feature or define it for the customer.

I want the product to distinguish learning, getting challenged and delegating. These are design commitments, not a claim that every current feature already does them well.

If I say I want to learn, don't rush to complete the exercise. Ask what I've tried. Give me a hint. Have me explain the answer and try again without assistance.

If I bring a draft or a decision, challenge it. Show me what I missed. Don't manufacture agreement, and don't quietly substitute the model's preferences for mine.

If I want to delegate, help me get a bounded result. Make it possible to check what happened. Keep consequential actions behind the permissions we agreed on.

An approval button alone isn't enough. I can approve something I don't understand. I need the evidence, a manageable decision and the ability to say no. Sometimes I need an expert rather than another reassuring response.

That ability to say no includes leaving the provider. I can run a model privately and still ask it what to think before I've tried thinking; local hardware doesn't preserve my judgment for me. But being trapped with one company makes it harder to choose a different kind of help.

That's why I want access to models from American and Chinese labs, local inference where it makes sense, and portable agent memory. Never rely solely on one AI lab. The practical rule is to use the hardware you have when it can do the job, and borrow stronger capacity when needed. If cloud is all you have, start there. Neither nationality nor owning a machine is a substitute for checking what you send and who can keep it. More copies aren't more privacy.

At cYpher.camp, we manage model access for hosted agents. Their identity and memory live in hosted files and state, and Export Brain lets you download that durable state. It isn't a copy of every working file, so keep your own project backups too.

That's portability, not anonymity. Hosted files aren't automatically on your laptop. Identifying details in a prompt or memory can still reach the model provider. Local inference also doesn't keep data local if your tools, sync or backups send it away.

I need to apply the same standard to my business incentives. I sell access to AI, so I have an incentive to encourage people to use it. That doesn't mean another AI session is always in their interest.

If the right answer is to stop generating and go practice, talk to the customer or have the conversation, the product should be able to say that. Sometimes a successful session should end with the person closing the app.


A first step

Start with something you want to do, not a tool you feel obligated to use. Build the application. Learn the language. Get stronger. Make more time for someone. Decide what success would mean to you. None of those goals requires you to use AI.

If you decide AI belongs in that goal, pick one recurring task that isn't sensitive and isn't a skill you're trying to develop, but takes time away from that goal. Decide what a useful result looks like before asking a model. Save the instructions and context in a folder you control.

Try the task with a model you choose and, if you need a fallback, an alternative. Check the results. Did the tool save work, or did it just give you something longer to review? You don't have to keep using it if it doesn't earn its place.

For the ability you do want to develop, keep the practice. Make an attempt before asking for help. If you use AI, try a hint, an objection or feedback on the attempt rather than a replacement. Come back later and see what you can do without it.

That last check isn't a demand to live without tools. It's a way to find out whether you're developing the ability you wanted or only borrowing the output.

Decide where the saved time should go, too. Otherwise another task, another notification or another session can take it. Practice, rest and time with people don't need an AI justification.

The hard things worth doing are the ones through which I develop an ability I care about, make something I believe in or show up for someone. I don't want to keep every obstacle between me and those things. I don't want to remove my participation either.

I want help doing the hard things I choose.

Agentic & distributed systems, DeFi, and the compute economics. One email a week, no fluff.

Subscribe to the newsletter →

About the author

Keenan Benning is the founder of cYpher.camp, CTO of DeFi All Odds, and a forward-deployed AI systems engineer. He builds AI systems and writes about the engineering, economics, and ownership questions behind them.

Other projects