GnothiGnothi
SeriesFieldsCommunity
Sign inGet started free

Your First Reusable Delegation

The problem, honestly framed

The chapter opens on a contractor who loses a bathroom job to silence: he sends a quote, hears nothing, and never finds out why. That failure — quotes going cold because nobody follows up — is the running example for the whole project, and every piece of "research" attached to it (interview notes, a pricing page, a cost figure) is stated plainly to be invented for teaching, not evidence of anything a real contractor would pay for.

Building a packet worth trusting

The real subject is the gap between chatting with an AI assistant and delegating to one in a way that produces checkable, reusable work. That starts with a small, deliberately labeled set of source files — forum threads, two interview notes, a competitor pricing page — chosen so every claim in later output can be traced back to a named document rather than dissolving into an unverifiable summary.

Setting up the workspace

Inside Claude, a Project bundles source files, standing instructions, and conversations. The instructions written here demand that every fact cite its source file and that anything unverifiable be flagged as a gap rather than filled in — rules meant to apply automatically to every future request in that workspace.

The brief, and the fabricated number

A five-part request produces a one-page opportunity brief on who has the problem, how they cope now, and what a missed follow-up costs. One paragraph — a precise dollar figure for annual losses — turns out to have no source in the files at all and no citation in the required claims list, exposing how an instruction can be skipped exactly where following it would reveal a problem.

Fixing the instructions, not just the answer

Rather than simply deleting the bad number, the standing instructions are rewritten to require quoted evidence for any figure and to name what evidence would settle each unresolved question. Re-run, the brief drops the invented cost entirely and instead lists three specific, answerable gaps — a result judged more useful than a confident but ungrounded one.

What was actually doing the work

A closing section distinguishes the trained model from the assistant application built around it, explains context and tokens, and contrasts this closed, file-only setup with a tool-using agent that could browse the open web. It also compares Claude Projects' full-context approach against ChatGPT's Projects feature, which relies on retrieval once files exceed its working window, explaining the choice to keep source packets small.


A contractor finishes a bathroom estimate on a Tuesday. He writes it up that night, sends it, and moves on to the next job. The homeowner reads it, means to reply, and forgets. Nobody follows up. Three weeks later the work goes to someone else, and the contractor never learns why.

That is the problem this series is going to work on. Not artificial intelligence in general. One narrow, boring, expensive failure: independent home-service contractors send quotes and then lose the job because nobody follows up. Plumbers, electricians, painters, small remodel outfits. One or two people, a van, a phone, no office staff.

I have to be honest with you about the example before we touch anything. The business is illustrative. The research packet you are about to see, the interview notes, the pricing page, the numbers in the brief — all of it is invented by me for teaching, and I will label it every time it appears. Nothing in this chapter proves that contractors will pay for anything. It cannot. Real evidence comes from real people, and getting it is work we have not done yet.

What this chapter is actually about is smaller and more useful than a business idea. It is about the difference between chatting with an AI assistant and delegating work to one in a way you can reuse.

You have probably chatted. You open a window, type a question, get a decent answer, and close the window. The answer might be right. You have no way to tell, and tomorrow you start from nothing. That is fine for a recipe. It is not fine for a decision that costs you three months.

Delegating is different in four ways. You supply the sources yourself, so the answer is grounded in material you chose. You write a brief that says what finished work looks like. You get an output you can check, line by line, against those sources. And when you find a mistake, you fix the instructions, not just the answer, so the next run is better than this one.

The whole chapter turns on one idea: a result you cannot check is not a result. It is a suggestion wearing a suit.

By the end you will have two things saved on your computer and in a workspace: an opportunity brief where every claim points at a named source, and a set of working instructions that caught a real mistake and will catch it again.

Start with the packet, because everything downstream depends on it.

My packet is small on purpose. There are six files. Four are collections of forum posts and support-style complaints from contractors talking about quotes going quiet — one file per source of discussion, so I can tell them apart later. Two are short interview notes, a page each, from conversations with a painter and an electrician. And there is one page of pricing and feature copy from two existing follow-up tools that already sell to this trade. Seven files, then, once I count properly, and each one has a name that says exactly what it is and where it came from: the trade forum thread on unanswered estimates, the painter interview from the eleventh, the pricing page capture from the two competitors.

Boring file names are a feature. In a few minutes I am going to ask for a brief where every claim cites the file it came from. If my files are called notes one, notes two and final final draft, those citations tell me nothing and I cannot check them. If the file name says painter interview, the citation is a place I can go and look.

Now, why small? There is a real temptation, once a machine is doing the reading, to shovel in everything you have. Four hundred pages of industry reports, a folder of screenshots, an old spreadsheet you never understood. Resist it, and here is the mechanism.

The assistant reads what you give it and writes something that fits together. It is very good at making things fit together. If your packet is a mess of overlapping, unlabeled, half-relevant material, the thing that fits together best is a smooth summary that sounds like an industry report — because that is the shape of the input. You will not be able to trace any single sentence back to any single fact, so you will not check, so you will believe it.

Seven labeled files, each of which you have read yourself, gives you something better: the ability to disagree with the output. When the brief says contractors mostly follow up by text message, you can go to the two interviews and see whether either of them said that. That is the only reason the packet is worth building.

The other half of honesty is what you leave out. My packet contains no market-size figures, no revenue estimates, no count of how many contractors exist in any country. I do not have those. Watch what happens later because of that gap.

Now the setup. This is where people expect a tour of a settings menu, and I am not going to give you one, because the settings are not the lesson. Two clicks and a text box are the lesson.

I am using Claude, the assistant made by Anthropic, and specifically its feature called Projects. As of the September twenty twenty-six documentation, Projects is available on the paid consumer plans — Claude Pro and Claude Max — as well as the team and enterprise plans. A Project is a workspace with three things attached to it: a set of source files, a set of standing instructions, and however many separate conversations you want to have inside it.

You sign in at Claude's website, click Projects in the left sidebar, and choose Create Project. You give it a name. Mine is called contractor follow-up research, because in four months I want to find it without thinking.

Then the standing instructions. In the project's panel there is a place for Project Instructions. Whatever you write there is applied to every message in every conversation inside that project, automatically, for as long as it exists. This is the part that makes the work reusable rather than a one-off, so it is worth writing carefully rather than dashing off.

Mine says, in substance: you are an evidence-grounded research analyst. Rely only on facts in the attached project files; do not bring in outside knowledge or estimates from your training. Every factual statement must carry an inline citation naming the exact source file it came from. If something cannot be verified from the attached files, label it as a gap rather than filling it in. Be concise and structured.

Read that again as four separate rules, because that is what it is. Where facts may come from. That every fact must be attributed. What to do when there is no fact. And how to write. Nothing clever. But it is now the standing law of this workspace, and I did not have to type it again.

Third, the sources. In the project there is a Project Knowledge area with an option to add content. You upload the files or drag them in and wait for processing. Claude Projects accepts PDFs, plain text and Markdown, Word documents, CSV and tab-separated files, HTML and ordinary data and code files. The per-file limit is thirty megabytes. The number of files in project knowledge is not capped, though an individual chat message supports up to twenty attachments. My seven small text files are nowhere near any of that.

There is a capacity bar showing how much of the project's context is being used, and it is worth understanding what it is telling you. The model has a working window of about two hundred thousand tokens — call it a hundred and fifty thousand words, roughly five hundred pages of ordinary text. Seven or ten small research documents use something like five to twenty percent of it. That matters for a practical reason. At that size, every document is held in the working context on every single message, so the assistant can look at all of them and cross-reference between them each time you ask something. If your knowledge base outgrows the window, Claude switches automatically to retrieval — pulling in the chunks that look relevant to your question rather than reading everything. Retrieval works, but it means a claim can be missed because the passage that contradicts it was not fetched. Small packet, whole packet, every time. Another reason to keep it small.

Now, the actual request. I open a new chat inside the project. I do not re-attach the files; they are already there, and re-attaching them is a habit worth breaking early.

A good brief has five parts, and I will name them as I write mine.

The result I want: a one-page opportunity brief on contractor quote follow-up. The context: I am a solo founder deciding whether to pursue this problem, and the attached files are everything I currently know. The constraints: every claim traced to a named source file, no numbers or facts from outside those files. The permitted actions: read and synthesize the attached files, nothing else. And the evidence of completion — this is the part people skip — a claims-and-sources list at the bottom, one line per factual claim in the brief, each naming the file it rests on, plus a section listing unanswered questions.

I also tell it what the brief should cover: who has this problem, how they handle it today, what a missed follow-up costs them, and how I might reach them.

That evidence-of-completion clause is doing something specific. Without it, I have to check the brief by reading its prose and remembering what was in seven files. With it, the output has to hand me its own audit trail. Checking becomes a five-minute job instead of an hour I will not spend.

Here is what came back, condensed, and remember that all of it is invented material.

Who has the problem: independent contractors and two-to-five person crews, quoting between five and fifteen jobs a week, with no administrative staff. How they handle it now: from the interviews, follow-up depends on whether the contractor happens to remember, usually a text message days later, sometimes nothing at all; the painter said he had a note on his phone that he stopped keeping. What it costs: the brief said contractors lose roughly one in four quoted jobs purely to silence, at an average job value of about four hundred dollars, which it put at something over twenty thousand dollars a year for a typical one-person operation. How to reach them: the trade forums in the packet, supplier counters, and the trade association pages mentioned in one interview.

Read that fourth item again. It is the best paragraph in the brief. It is also wrong, and I would have believed it, because it is exactly the kind of sentence I wanted to read.

I found it by not reading the prose. I went straight to the claims-and-sources list at the bottom and worked up. Two faults.

The first: the average job value of four hundred dollars, the one-in-four loss rate, and the twenty-thousand-dollar annual figure appear nowhere in my seven files. I know this for certain, because I built the packet and deliberately put no numbers in it. The assistant did not steal them from a hidden source. It produced plausible figures because the shape of the request implied a figure belonged there, and it can produce plausible figures endlessly.

The second fault is quieter and, for our purposes, more instructive. My brief required a source line for every factual claim. The cost claim has no source line. It is simply missing from the list. The instruction was not refused; it was skipped, in the one place where following it would have exposed a problem. This happens. An instruction is something the assistant tries to follow, not a wall it cannot walk through. That distinction will come up again and again in this series, and it starts here.

So: what does an unsupported plausible sentence actually cost a founder?

Say I believe the twenty-thousand-dollar figure. Now I have a market. I can charge a contractor a couple of hundred dollars a year against a twenty-thousand-dollar loss, and that is an easy sell, so I stop asking whether it is true and start building. Two months of evenings go into software. Then I talk to actual contractors and find that most of the quotes going quiet were never going to close at any price, that the jobs they lose are the small ones they were relieved to lose, and that the number that matters is not four hundred dollars a job but which two jobs a month were worth chasing. The figure was not just wrong. It pointed my attention at the wrong thing, and I paid for that in weeks.

That is why the check is not optional and why I built the packet to make checking possible.

Now the repair, which is the actual point of this chapter.

The weak move here is to reply, "you made that up, take it out." It works. You get a clean brief. And you have learned nothing that survives to the next task, because next week you will run something else and the same thing will happen in a place you are not watching.

The strong move is to go back to the Project Instructions and add the rules that would have caught this. Mine gained three.

First: no figure, percentage, currency amount or count may appear anywhere in the output unless it is accompanied by a quoted sentence or phrase from a named source file containing that figure. Not a citation — a quote. A citation can be attached to a source that does not actually say the thing. A quote either contains the number or it does not, and I can search the file for it in ten seconds.

Second: where information is unknown, label it as a gap and do not estimate, extrapolate, or offer a typical or industry-standard value. That last clause matters, because "typical for the trade" is the sentence a plausible invention hides inside.

Third: for every gap, state what evidence would settle it and where that evidence would come from.

Then I re-ran the same request in a new chat in the same project.

The second brief was thinner and much better. The who and the how-they-handle-it-now sections survived nearly intact, with quoted lines from the two interviews and the forum threads underneath them. The cost section did not survive at all. In its place were three gaps, each with a note on what would settle it.

The first gap: what a lost quote is actually worth. Nothing in the packet gives a job value or a close rate. Settling it needs contractors reporting their own quoted-versus-won numbers, or a look at somebody's actual quote history.

The second: how many quotes go quiet for reasons follow-up could fix, as against quotes that were never going to close. The forum posts complain about silence but do not distinguish these, and they are different businesses.

The third: whether contractors already pay for something here. The packet has a pricing page from two competing tools, which proves the tools exist and what they ask for, and proves nothing about whether anyone buys.

A brief that names three things it cannot answer is worth more than a confident brief, and the reason is not modesty. The confident brief gave me a plan. The honest one gave me a shopping list of exactly which conversations to have next, and it told me the order, because until I know what a lost quote is worth, nothing else about this business can be priced.

Two things now exist that did not exist an hour ago. There is the opportunity brief, with its quoted sources and its three named gaps, which I can reread in a month and still trust. And there are the working instructions living in the project, which are now better than they were because a mistake happened and I spent the mistake on them rather than on a single answer. Every future run in that workspace inherits them.

That is the loop. Sources you chose, a brief that defines finished, an output with its own audit trail, one real fault found, and the fix written into the instructions instead of the reply. The next chapter takes those three named gaps to actual contractors, which is the only place the answers live.

What was actually doing the work

Before we go further, there are four words this series will use constantly, and you should learn them from what you just watched rather than from definitions.

The model is the trained system that produced the text. Training and use are separate events. Training happened before you arrived: a very large amount of text was used to adjust the model's internal numbers, its weights, until it got good at continuing text sensibly. That is finished. Nothing you did in your project changed a single weight. What you did was inference — the trained system reading your current input and producing output from it.

That distinction explains the fault you found. The model's weights carry a diffuse sense of what business writing sounds like, including that cost sections contain dollar figures. When the input had no figure and the requested shape demanded one, the model produced what fits. It was not lying, and it was not consulting a source. It was continuing text plausibly, which is the one thing it is built to do.

The input the model sees on any given message is its context. Your seven files, your standing project instructions, your request, and the conversation so far all arrive as context, converted into tokens, which are the chunks of text the model actually processes. The two-hundred-thousand-token window is the size of that input space. This is why the small packet mattered: with the whole packet inside the window, the model could compare the interview note against the forum thread on the same message. And it is why context is nothing like memory. Nothing persists in it. Each message rebuilds the whole input from scratch.

Which brings us to the assistant application. Claude, the thing you signed into, is not the model; it is software wrapped around the model. It holds your account, stores your conversations, and — this is the part you used — takes the files sitting in Project Knowledge and the text in Project Instructions and pastes them into the context on every single message, so you do not have to. The application also handles what happens when the packet outgrows the window, switching to retrieval so that the passages most relevant to your question are fetched and inserted instead of everything. Projects is an application feature. The model has no idea a project exists.

A tool-using agent is the next step out, and today you did not use one. An agent is an assistant that can take actions between reading your request and answering it: search the web, run a command, read a file from disk, write to another system, then look at what came back and decide what to do next. What you ran was closed. It read the seven files you supplied and wrote a document. That was the right choice for this job, and deliberately so. A brief whose entire evidence base is a folder you assembled is a brief you can audit. Let an agent browse the open web for the missing market size and you get your twenty-thousand-dollar figure back, now with a link, and checking it becomes an afternoon's work instead of a minute's.

Fourth, and easiest to lose: durable records. The brief on your disk is a durable record. The project instructions are a durable record. A conversation is not. Context vanishes when the message ends. Application storage keeps your chats, but it is not a place a business keeps things it needs to find. Anything the business will rely on later gets written to a file, a spreadsheet, or a proper system of record, and named so a person can find it. The reason this chapter left two named assets behind is that assets are the only part of a session that lasts.

One practical note on choosing where to do this kind of work, kept short because it only matters at the moment of choosing.

ChatGPT, from OpenAI, has an equivalent feature, also called Projects, also on its paid consumer plans, Plus and Pro. The flow is nearly the same shape: create a project in the sidebar, add files as sources, write standing instructions in the project's settings, then open a chat inside it. Its per-file allowance is far larger — up to about five hundred and twelve megabytes, or around two million tokens of extracted text — and it accepts spreadsheets, presentations and images alongside documents.

For the specific job of grounding a brief in about ten small files, the difference that matters is not the file limits. It is how the sources reach the model. ChatGPT Projects indexes your files and retrieves the passages that look semantically relevant to each message. Claude Projects, when the collection fits inside the working window, holds all of it in context on every message instead. For auditing a short packet, holding everything is the property you want, because a claim can then be checked against the file that contradicts it — and a retrieved answer is only ever as good as what got retrieved. If your sources were a thousand pages, retrieval would be the only option and that calculus would change.

That is the whole comparison as it bears on this decision. I am staying with Claude Projects for this series, because the packets we will build are small and auditable and I want them read whole.

Next, we go and talk to contractors about a lost quote is worth.