GnothiGnothi
SeriesFieldsCommunity
Sign inGet started free

What Delegation Leaves Behind

Chatting versus delegating

The chapter builds a worked example around a solo founder testing a productized service for small trade businesses — electricians, plumbers, roofers — who miss enquiries because they're on the job all day. Working from a small set of invented, numbered interview notes and review excerpts, it first shows the tempting shortcut: paste everything into a chat and ask "is this a good idea?" The reply is fluent, well-organised, and useless — it mixes claims drawn from the notes with plausible-sounding numbers the assistant simply generated, and none of it survives closing the tab.

The fix is to build somewhere for the work to live. Using Claude's Projects feature, the founder creates a dedicated workspace, understands what a project actually is and uploads her research so it becomes standing material for every conversation inside it rather than something re-pasted each time. She writes a five-part brief — a named deliverable, a strict rule that only the uploaded documents count as evidence, a requirement that every claim carry a quoted source, no outside searching, and a claim-by-claim audit at the end.

The first draft still slips in an unsupported "forty percent" figure and smooths over the fact that most of her sources say nothing about cost at all. Anthropic's own guidance warns plainly that Claude can produce convincing but ungrounded responses and shouldn't be treated as an unverified source of truth — a warning this chapter spells out directly. Rather than arguing with the output, the founder rewrites the project's standing instructions so every future brief must quote its sources, flag unsupported numbers, and end with an explicit list of what remains unknown. The rerun is thinner, less impressive, and far more honest — ending in six real questions no chat could have answered.

What the model, the application and the agent each are

A closing section separates three things people conflate: the trained model that only produces plausible continuations of text; the application layer — here, Claude's project files and saved instructions — that decides what the model actually gets to see; and a tool-using agent, which takes action in other software rather than just producing text. None of the work described here involved an agent — everything that changed was a document inside one project. The distinction matters for verification: to check a claim, you open the uploaded source yourself and compare it to the quoted text, rather than trusting that the interface has done it for you.


An assistant like the one you already have open in a browser tab will answer almost anything you ask it. That is the trap. Ask it whether your business idea is any good and you will get a fluent, confident, well-organised reply in about twenty seconds, and you will feel like you have done research. You have not. You have had a conversation.

This is a course about building and running a real business with AI agents — one that earns money, serves customers, and eventually keeps operating while you are not watching it. That last part is a long way off. It starts here, with a distinction that decides whether everything after it works: the difference between chatting with an assistant and delegating to it.

The difference is not the wording of one clever request. Plenty of people will sell you a list of magic phrases. The difference is what is still there when you close the tab. Chatting leaves nothing. Delegating leaves a container that holds your source material, a result you can check line by line against that material, and a set of instructions that the next run inherits without you retyping them. That is the whole idea of this chapter, and by the end of it you will have watched one built.

So let us have something to work on. Everything in the example that follows is invented by me for teaching. I want to say that plainly and early, because the fastest way to ruin a business is to mistake made-up inputs for evidence about a market. Invented notes prove nothing about demand. They exist so you can watch the method, and then run the method on real material you gathered yourself.

The founder in the example is one person, working alone, with no product yet. She is considering a productized service for small independent trade businesses: electricians, plumbers, roofers. A productized service means a fixed, repeatable job sold at a fixed price, rather than custom consulting where every project is negotiated from scratch. Her hunch is this: the owners of those businesses are on a roof or under a sink for most of the day, they cannot answer the phone, and enquiries go cold because nobody replies for two days.

That is a hunch. It has the shape of a real problem, which is exactly what makes it dangerous. It is easy to believe.

Her research packet is small — about a dozen short items. Some are notes from conversations with tradespeople. Some are the kind of short excerpt you would find in a public review. And there is a separate short document listing how each business handles enquiries right now. Every item is numbered, so any claim can point at exactly which item it came from. That numbering is not decoration. It is the thing that makes checking possible later, and if you take one habit away from this chapter, make it that one: number your sources before you hand them to anything.

Here are five of the items, so you can see what the material actually looks like.

The first is a note from an electrician with two vans: "I check my phone at lunch and at nine at night. If someone rings at ten in the morning, I might get back to them Thursday." The second is from a sole-trader plumber: "I lost a bathroom job because she rang someone else the same afternoon. That was about eleven hundred pounds of work." The third is from a roofer: "My wife picks up when she can. She doesn't know what to quote, so she says I'll ring you back." The fourth is a review excerpt about a heating firm: "Rang three times over two days, never got an answer, went elsewhere." The fifth, from the handling document, records that one electrician sends every missed call to a voicemail greeting he recorded four years ago, which still gives his old landline.

Look at what is and is not in there. Four different people describe the same failure. Exactly one of them attaches money to it. Nobody says what percentage of enquiries they lose, because nobody counts. Hold onto that, because it is going to matter in about ten minutes.

Now the casual version, which is what almost everyone does. Open a fresh chat. Paste in all twelve items. Type: "Is this a good business idea?" Wait.

What comes back is genuinely impressive to read. It will have a heading called Market Opportunity. It will tell her the pain point is real and underserved, that trade businesses lose a significant share of inbound enquiries to slow response times, that the natural offer is a shared inbox and answering service with a same-day response guarantee, and that she should price it as a monthly subscription because trades value predictable costs. It will end by encouraging her.

Every sentence in that reply is plausible. Some of it may even be true. And it is useless for three separate reasons.

It is unchecked. Somewhere in that answer are claims that came from her twelve items, and somewhere else are claims that came from the model's general sense of how the world works, and the reply gives her no way to tell which is which. It is unrepeatable. If she asks the same question tomorrow she will get a differently organised answer with different emphasis, and she will have no idea which version to trust. And it is gone. When she closes that tab, the notes are not stored anywhere she can reuse, the instructions she gave are not saved, and tomorrow she starts from an empty box again.

Nothing was left behind. That is the definition of chatting, and it is not a matter of the question being badly phrased. She could have written three paragraphs of beautiful instructions and still ended up with an answer she cannot verify and cannot reproduce.

Delegating means doing four things instead. Give the work a permanent container. Put the sources inside it. Write a brief that says what a finished result looks like. Then check the result against the sources rather than admiring it.

The container, in the assistant we are using here, is called a project. Claude, from Anthropic, has a Projects feature in its web application at claude dot ai: you open Projects in the sidebar, choose Create Project, and give it a name and a short description. A project is just a workspace that keeps its own files and its own standing instructions, so that every conversation you start inside it begins already knowing your material. Free accounts get a small number of projects; paid tiers allow more. She names hers something dull and accurate, like Trade Enquiry Research.

Inside the project there is a pane called Project Knowledge. That is where the sources go. You click Add Content and upload files — plain text, markdown, PDF and Word documents are accepted. The uploaded files get parsed and indexed, and from then on they are standing context for every conversation in that project rather than something she pastes in again each morning. She uploads two documents: the numbered interview notes and review excerpts, and the shorter document describing how each business currently handles enquiries.

Context is worth pinning down now, because the word gets used loosely. Context is the material the model can actually see while it is producing its answer: your request, the conversation so far, and whatever the application has pulled in from your files. It is not memory in any human sense. It is more like the pile of paper on the desk in front of someone writing a report. If a fact is not on the desk, it does not exist for that report — unless the model reaches for something it half-remembers from training instead, which is precisely the failure we are going to catch.

Now the brief. And here is where delegation earns its name: she is not asking a question, she is commissioning a deliverable. A useful brief has five parts, and I want to unpack each one as she writes it rather than list them and move on.

The first part is the result she wants. Not "analyse this" — a named artifact with a shape. She asks for an opportunity brief, no longer than one page, in five sections: the problem, how these businesses handle it now, what the failure costs them, who she should try to reach, and what is still unknown. Naming the sections matters more than it looks. It means she can tell at a glance whether the work is finished, and it means she cannot be fobbed off with three strong paragraphs and no mention of cost.

The second part is context: what evidence is permitted. She writes that the two uploaded documents are the only evidence for this brief, and that general knowledge about the trades, about small business, or about response times must not be used. This is worth being blunt about, because by default an assistant draws on both what it was trained on and what you have given it, and those two sources arrive in the answer looking identical. Anthropic's own documentation is direct that Claude can produce misleading responses and convincing-sounding quotes that are not grounded in the material, and that it should not be relied on as an unverified single source of truth. So the restriction has to be stated, not assumed.

The third part is the constraints. Every claim must name the numbered item it came from and quote the words. Any figure that does not appear in the sources may not appear in the brief. Anything the sources do not establish is written as the word unknown, not smoothed over and not estimated.

The fourth part is permitted actions. For this run: read the two supplied documents, and nothing else. No searching the web. She is not being precious — she is making the result checkable. If the assistant can go and find outside material, she can no longer tell whether a sentence came from her research or from a page she has never seen.

The fifth part is evidence of completion. She asks for a claim-by-claim list at the end: every factual sentence in the brief, beside the item number and quoted words that support it. This is the part people skip, and it is the part that turns a nice-looking answer into something she can audit in four minutes.

She sends it. Out comes a tidy one-page brief.

Now she inspects it, which means picking claims and trying to break them.

She takes the first claim: that owners routinely go most of a working day without seeing an enquiry. She looks for the line. Item one, the electrician with two vans, checking his phone at lunch and at nine at night. Supported, and the quote matches.

Second claim: that the current workaround is often a family member or an out-of-date voicemail, neither of which can quote a price. Items three and five. Supported.

Third claim, in the section on what the failure costs: "small trade businesses typically lose around forty percent of inbound enquiries to slow response." She goes looking for that number. It is not in item two, which gives one bathroom job at about eleven hundred pounds. It is not in the review excerpt. It is not in the handling document. It is nowhere in the twelve items, because none of the twelve people counted anything.

That number came from the model's training, not from her research, and it walked into her brief wearing the same clothes as the two claims that were fine. That is the failure to understand. It is not that the assistant lied to her. It is that fluent prose is not evidence, and a well-formed sentence carries no information about where it came from.

There is a second problem, quieter and just as expensive. Her brief asked for what the failure costs, source by source. The brief has a confident paragraph on cost — but it is built on that one bathroom job, generalised outward. For eight of the twelve items, the sources say nothing about money at all. The honest output was eight lines reading unknown. What she got instead was a smooth paragraph that reads as though the question has been answered.

Both of these matter because of what she is about to build on them. The offer she designs, the price she sets, the person she rings first: those decisions rest on these sentences. If a made-up forty percent is sitting underneath her pricing, she will not discover the error when she reads the brief. She will discover it months later, when the business does not work and she cannot say why.

So she repairs it. Not by arguing.

This is the instinct to break, because it is such a natural one. The obvious move is to reply "that forty percent figure isn't in my sources, please be more accurate," and the obvious move produces a corrected answer and no lasting improvement. Tomorrow, in a new conversation, the same fault returns. Arguing fixes an output. Changing the instructions fixes the process.

She goes back to the Project Knowledge pane, chooses Set project instructions, and writes rules that will be prepended to every conversation in this project from now on:

  1. Answer only from the documents in Project Knowledge; never use outside or general knowledge, even when it seems obviously true.
  2. Follow every factual claim with the item number and a verbatim quote, in square brackets, from the document it came from.
  3. Never state a number, percentage or currency amount that does not appear word-for-word in the sources.
  4. End every brief with a section headed "Unknown — must be asked in a real conversation," listing each question the sources do not answer.
  5. End every brief with a claim-check list pairing each claim to its supporting quote, and if a claim cannot be paired, delete the claim.

Then she saves them, and re-runs the brief.

The new version is noticeably worse to read, and much better to own. It is about half the length. The problem section now stands on four quoted items and says so. The workaround section is solid, because her sources actually cover that ground well. The cost section is two sentences long: one job of about eleven hundred pounds reported by one plumber, and one review in which a customer went elsewhere after three unanswered calls over two days, with no amount attached. No forty percent anywhere.

And then the section that was missing before, which is the most useful thing on the page. Unknown, must be asked in a real conversation: how many enquiries each business receives in a week; how many of those go unanswered; what an average job is worth to each of them; whether the owners believe they are losing work, or only that answering the phone is annoying; who currently picks up when the owner cannot; and whether anyone has ever paid for help with this.

That thinness is progress. It is not the assistant underperforming — it is the actual state of her evidence, finally visible. She has twelve items of anecdote from people who do not count their enquiries. The first brief hid that behind a paragraph of prose. The second brief shows it, and in showing it, hands her the questions to take to real people. Those questions are the next piece of work, and she could not have written them from her hunch, because her hunch already assumed the answers.

Two things now exist that did not exist an hour ago. There is a research workspace with her numbered sources in it and instructions that will govern the next run, whatever that run is about. And there is one honest brief, every sentence of which she can trace to a quote, plus a list of what she does not know. Neither of those survives a chat window.

The model, the application and the agent

Three words got used almost interchangeably above, and they are three different things. Keeping them apart is what lets you predict what a tool will do before you rely on it, and diagnose it when it disappoints you.

The model is the trained system that produces text. Training is the process that produced it — an expensive, one-off process that adjusted the model's internal numbers, its weights, until it became good at continuing text. That happened before you ever met it, and nothing you type changes it. What you do when you use it is called inference: the trained system takes whatever is in front of it and produces a continuation. This is exactly why the forty percent figure appeared. The model has read enormous quantities of writing about small business, and "typically lose around forty percent" is the kind of sentence that plausibly continues a paragraph about slow response times. Producing plausible continuations is what it is for. It has no separate faculty that checks whether a sentence is in your documents, unless something makes it check.

The assistant application is the software wrapped around the model — in this chapter, Claude's web application, and specifically its Projects feature. The project, the uploaded files, the saved instructions: none of that is the model. It is bookkeeping that the application does on your behalf. When she saved those project instructions, nothing about the model changed. The application simply agreed to put those instructions in front of the model at the start of every conversation in that project. That is a small mechanism, and it is doing most of the work in this chapter, because the difference between a chat and a delegation is mostly a matter of what the application reliably puts on the desk.

Knowing that also tells you how to verify. In the Project Knowledge pane you can click any uploaded document and read its raw text inside the workspace. So the audit is a two-step move: ask in the conversation for the exact sentences it used, then open the document and look. If the quoted words are not there, they were generated rather than found, and you have caught it in the only way that reliably works — by looking at the source yourself. The web chat does not hand you interactive citation widgets to click; that is why she wrote the bracketed-quote rule into her instructions rather than expecting the interface to supply it.

A tool-using agent is the third thing: a system that can take actions in other software — read an inbox, write to a spreadsheet, call another service — rather than only producing text for you to read. Nothing in this chapter was an agent in that sense. She uploaded files by hand and read the result by eye, and the only thing that changed anywhere was a document inside her own project. That distinction is worth holding, because the questions you ask change completely once actions are involved. With text, the question is whether the claim is supported. With actions, the question is what happened in the other system, and whether it should have.

One more distinction, because it will save you grief. Context is what the model can see during a single run. The project's saved files and instructions are application state — durable, but living inside one product. Neither of those is a business record. When she eventually has customers, their details will live in a system of record that exists whether or not she ever opens this assistant again. Confusing the three is how people end up with important company information stranded in a chat history.

So: know when to keep going and when to start clean. If a conversation is long and has drifted, or if you have corrected the same fault twice, start a fresh conversation inside the project. She loses the chatter and keeps everything that matters, because the sources and the instructions are held by the project rather than by the thread. That is the practical payoff of the container, and it is why the first move in delegating any work is to make one.

What she has now is a reusable research workspace and one brief she can defend, sentence by sentence, to anyone who asks where a claim came from. What she does not have is evidence that anybody will pay her. Her own brief says so, in a list of six questions, and every one of them can only be answered by a real electrician, plumber or roofer in a real conversation.