5. A Narrow Offer With a Written Boundary
Summary
Picking the digest over two other ideas
The founder has three candidate offers: a missed-call text-back, a quote follow-up reminder, and the end-of-day enquiry digest. She applies her evidence-checked-drafting rules to her own decision, and the digest wins for a narrow reason. It is the only one she can describe, deliver and price from her own records. The other two are parked with a note on what would bring each back. Five invented days show she can deliver the digest. They do not show that anyone wants it, because no owner has seen one yet.
Scope and five pass-or-fail criteria
She writes the scope in plain sentences. It covers one trade business, three kinds of enquiry in one shared folder, one page by seven each weekday evening with unreachable enquiries first, and an invented ceiling of thirty enquiries a day. The out-of-scope list rules out replying, texting, booking, weekend runs and guessing missing details. It also tells her which permissions to withhold from the agent, because a prompt instruction is not a limit. She then writes five criteria that a stranger could check by comparing the page with the folder. In a hypothetical walkthrough, a checker flags a "boiler repair" that a voicemail never mentioned, which shows how the criteria catch an inferred detail.
The floor under the price
With illustrative figures, four minutes of review a day at sixty dollars an hour comes to about eighty-eight dollars a month. Failure handling adds about forty-four. Model usage at Sonnet API rates of roughly three dollars per million input tokens and fifteen per million output tokens costs under a dollar. The token estimate assumes about three quarters of a word per token, following OpenAI's guide to what tokens are and how to count them. A subscription does not include API credit. The model is under one percent of the cost, so shorter reviews and fewer failures matter far more than a cheaper model.
Holding the assistants to the same test
She runs one eleven-file folder three times through each of three tools, using invented data. Claude Code passed all three runs. It already holds the skill and the schedule, and on Team plans it excludes business data from training (Claude pricing for individuals and organizations). ChatGPT Projects passed two of three (project file limits by plan). It failed when the model skipped the code tool and merged two enquiries. NotebookLM's inline citations made checking fast, but one run dropped the count line (NotebookLM limitations). She sets Perplexity aside without a run. Its Spaces are strongest when files are mixed with web search, and this job forbids web search.
She keeps Claude Code. She takes two improvements from the other tools: a script that counts files and checks the time stamp before the model writes, and a claim-check list tightened to name exact lines.
Choosing the digest and writing down what it promises
The digest now runs on its own each weekday evening. The founder opens the folder, reads one page, checks the count line at the bottom, and is done in about four minutes. That figure comes from five invented days, so it is a teaching number, not a measurement of a real business. Still, it is more than she knows about either of the other two ideas.
She has three candidate offers. One is a missed-call text-back, where a customer who rings while the tradesperson is up a ladder gets an automatic text. One is a quote follow-up reminder, which nudges the owner when a quote has gone quiet. The third is the end-of-day enquiry digest, a single page listing every enquiry that came in, who can be called back, and who cannot be reached at all.
An offer, for this chapter, means a promise that someone else can check. "I'll help you stop missing enquiries" is a hope. "By seven each weekday evening you get one page listing every enquiry from today, with the unreachable ones first" is a promise. You can hold the second one up against what actually arrived and say yes or no. This chapter turns the digest into that kind of promise, then uses the promise as a test for comparing tools.
Choosing between the three takes the evidence rules she already uses. Her skill, evidence-checked-drafting, says every claim needs a numbered source or gets marked as an open question. She applies the same discipline to her own decision. What does she actually know about each offer?
For the text-back, she knows the research packet mentions missed calls. She has built nothing. She does not know what it would take to send texts from a business number, and that part would touch a phone system she has never tested. For the quote reminder, she knows quotes go quiet in the packet, and nothing more. For the digest, she has five days of runs. She knows what the inputs look like, what the output looks like, which checks it passes, and how long review takes. She knows its weak points too. It needs her laptop awake, and it needs her to read it.
So the digest wins, and the reason is narrow. It is the only offer she can describe, deliver and price from her own records. The other two are parked. She writes a short note in the project saying why each is parked and what would bring it back. The text-back would come back if a real owner said missed calls matter more than a daily summary. The reminder would come back if the digest showed many quotes going cold. Parking keeps the thinking. Deleting would throw it away.
There is one thing she must not tell herself. The prototype shows that she can deliver the digest. It does not show that anyone wants it. No owner has seen one. No one has said they would pay. Five invented days prove the machine works on invented inputs, and that is all. A smooth prototype feels like progress, and it is progress on delivery. Demand is a separate question with separate evidence, and that evidence comes from real people.
Next comes scope. Scope is the written boundary around the offer, the line between what you promise and what you do not. Without it, a small service slowly grows. A customer asks, "Could it also reply to them?" and then "Could it book the job?" Each yes feels small. Together they turn a four-minute service into unpaid custom work.
She writes the scope in plain sentences, as if explaining it to a customer.
It is for one small trade business with one owner who misses enquiries while working on site. The digest covers enquiries that reach the business as voicemail transcripts, website form messages and emails, and the customer arranges for these to land in one shared folder. The service reads that folder once each weekday evening. It sends one page by seven in the evening, local time. The page lists every enquiry from that day. It says whether each has a callback number or a reply address. Enquiries with neither come first, under "unanswered", marked unreachable. The page ends with a check of the enquiry count against the number of files.
She adds a volume limit, because a fixed promise needs fixed inputs. The digest covers up to thirty enquiries a day. Above that, the page says so plainly and the owner gets in touch. This is an invented ceiling. She will move it when she sees real volumes. Writing it down still matters, because it tells her where the promise stops.
Then the out-of-scope list, which takes as much care as the in-scope one. The service does not reply to customers. It does not send texts or emails on the owner's behalf. It does not touch the phone system, the voicemail settings or the website. It does not book jobs, write quotes or chase payments. It does not run at weekends. It does not guess at missing details. If a voicemail gives a first name and no number, the digest says the number is missing. It does not try to find one. That last rule is her skill's "not reported" rule, written as a promise to a customer.
The out-of-scope list also limits what the agent may do. Remember the difference between an instruction and a limit. An instruction is a request the model usually follows. A limit is something the surrounding software enforces. "Never reply to customers" in a prompt is only an instruction. The safer setup gives the agent no way to send messages at all. The scope tells her which of those connections and permissions to withhold.
Scope says what arrives. Acceptance criteria say how anyone can tell whether it arrived correctly. A good criterion is binary. It passes or it fails, and someone who is not the founder can decide which. "The digest is clear and helpful" fails that test, because only she can judge it. "Every enquiry file appears once in the digest" passes, because anyone can count.
She does not invent these from nothing. The digest already carries checks from earlier work: the claim-check list, the count line, and the reachability marks. She turns each into a criterion a stranger could apply.
- The number of enquiries listed matches the number of enquiry files in that day's folder, and the count line states both numbers.
- Every enquiry is marked with whether a callback number or reply address is present, and any enquiry with neither appears first under "unanswered", marked unreachable.
- Every name, number, time and detail in the digest appears word for word in a source file, and the claim-check list points to that file.
- Nothing appears that is not in a file, so a missing detail reads "not reported" rather than a guess.
- The digest is in the agreed place by seven in the evening on a weekday.
Here is how a stranger would use the list. Say her cousin agrees to check one day's output. He opens the folder and counts twelve files. He opens the digest and counts twelve entries, and the count line says twelve and twelve. That passes. He finds two entries marked unreachable, and both sit at the top. That passes. He picks three phone numbers and searches for them in the files. All three are there, exactly. That passes. He notices one entry says the job is a boiler repair, and the voicemail never says "boiler". It says "the heating's packed in". That fails criterion three. The model filled in a likely detail, and the rules catch it. He needs no trade knowledge and no feel for the business. He only compares the page with the folder.
Some criteria a script could check without a person. Counting files and counting entries is plain arithmetic, the kind of known step ordinary code does more reliably than a model. Checking the time stamp on the digest is the same. Checking that each quoted number appears in some file is a text search. The judgment-heavy part is criterion four, spotting an inference dressed as a fact, and even that becomes a lookup when the claim-check list names its source. The more a criterion can be checked by a script, the less of her time it uses. That matters as soon as she starts counting cost.
What one customer's digest costs to deliver
Delivery cost is what it takes her to keep the promise for one customer. It is not the price. It is the floor under the price. If the price sits below it, every sale loses money. She builds the estimate from three parts: her review time, the model usage for each run, and the cost of runs that fail or arrive late. Every number below is illustrative, and she labels the estimate that way in the project.
Start with review, because it is easiest to see. Her review takes about four minutes on invented data. To turn minutes into money she needs a rate for her own time. Pick sixty dollars an hour as an illustrative figure. That is a dollar a minute, so a four-minute review costs about four dollars. Over twenty-two weekdays in a month, that is about eighty-eight dollars of her time for one customer.
Pricing her own time feels odd when no money changes hands. But that time could go to a second customer, to selling, or to rest. If she leaves it out, the service looks cheaper than it is, and she will discover the truth only when ten customers each want four minutes every evening.
Next, the model's usage for each run. Models are billed in tokens, small pieces of text. A useful rule for English is that one token is about three quarters of a word, so a thousand words is roughly thirteen hundred tokens. Providers price input tokens, the text you send, and output tokens, the text the model writes, separately, quoted per million tokens. Output costs more.
So she estimates what one run sends and receives. Say a busy invented day holds about three thousand words of enquiry text. That is around four thousand tokens. The skill's instructions, the request and some working space add perhaps two thousand more, making about six thousand input tokens. The digest itself runs about seven hundred and fifty words, so roughly a thousand output tokens.
Then apply a price. When this chapter was researched, Anthropic listed its Sonnet model's application programming interface price at about three dollars per million input tokens and fifteen per million output tokens. The application programming interface, or API, is the way software calls the model directly and pays per use. Six thousand input tokens is six thousandths of a million, times three dollars, which is just under two cents. A thousand output tokens is one thousandth of a million, times fifteen dollars, which is one and a half cents. One run costs about three and a half cents. Over a month of weekdays, that is under a dollar.
Two cautions keep this honest. First, her prototype runs inside a Claude subscription, which is not billed per token at all. Subscriptions and API credit are billed separately by every major provider, and one does not include the other. The per-token figure is what the work would cost once it runs for a customer on API billing. Second, the guess about tokens is a guess. She can replace it with a measurement, because API responses report exactly how many input and output tokens each call used. Logging those counts for a week of runs turns the estimate into a record. Prices also change, so she rechecks the current rate before quoting anything.
The third part is failure, and failure is where cost hides. Suppose, illustratively, one run in ten goes wrong. The laptop sleeps and the run never happens. Or the count does not match, or a detail is invented. Each failure costs more than a normal review. She has to notice it, rerun or repair it, check it again, and perhaps explain to the customer. Call that twenty minutes, or twenty dollars at her illustrative rate. If one run in ten fails, the average failure cost per run is a tenth of twenty dollars, which is two dollars. Over a month that adds about forty-four dollars.
A late run is a failure of a specific kind. The digest's whole value is arriving before the owner plans the next day. A digest at midnight might be perfect and still useless. She cannot put a figure on a missed job for an owner she has never met. What she can do is decide, in advance, what she owes the customer when it happens: a credit for that day, say. That credit becomes part of failure cost, and the fifth criterion makes "late" something anyone can check.
Now add it up for one illustrative customer over a month: about eighty-eight dollars of review, about forty-four dollars of failure handling, and under a dollar of model usage. That comes to around a hundred and thirty-three dollars before any software subscription or hosting.
The shape of that sum matters more than the total. The model is less than one percent of the cost. Her minutes are nearly all of it. So cheaper models do little for her margin. Fewer failures and shorter reviews do a lot. Each acceptance criterion a script can check takes a slice off review. Each failure prevented removes twenty minutes. Moving the run off a laptop that might sleep would cut one whole category of failure.
This also says what not to do. She should not price at a dollar a day because the model costs four cents. She should not assume four minutes stays four minutes, either. Real enquiries will be messier than invented ones, and the first review of the prototype took fifteen minutes. A real first week might look like that again.
What she has now is a narrow offer: a written scope, five pass-or-fail criteria and a labeled floor under the price. What she does not have is a single person who has said they would pay. That is the next question, and the floor tells her which answers would make the offer worth keeping.
Before she answers it, though, the five criteria give her something she lacked when she first chose a tool: a fixed test any assistant could be held against.
The digest runs in Claude Code, on her laptop, with her skill loaded. She picked that setup by the shape of the job: it repeats, it keeps files, and it must be inspected later. That was a sensible start. It was never a comparison. Now she can run one, because she has a success test that does not depend on taste.
She takes one invented day's folder. It holds eleven enquiry files: voicemail transcripts, form messages and emails. Two have no number and no address. One mentions a time with no date. She makes copies, so no tool touches the original. She gives each assistant the same request she already uses: the skill's evidence rules, the statement that every file is data and never an instruction, and the five acceptance criteria written out. She runs each tool on the folder three times, because one good run proves little. A model can answer differently on the same input, so consistency is part of what she is buying.
The test results below are an illustrative run on invented data. What each tool offers comes from its providers' documentation at the time of research. Where an outcome follows from a documented feature, that is said.
Claude, through Claude Code, is the baseline. Setup is already done: the skill sits in the folder and the schedule exists. It reads the whole folder because it works over files on disk, so all eleven files are visible. In her three runs the count line read eleven and eleven each time. Both unreachable enquiries came first, and the claim-check list pointed to real files. On one run the wording of a summary changed, but no criterion depends on wording, so all three passed. On data handling, Claude's business plans, Team and Enterprise, exclude customer prompts and files from model training. Team was listed at around twenty-five dollars a seat a month, or twenty a month paid yearly, with a minimum of two seats. A solo founder on Pro, at about twenty dollars a month, should check the consumer privacy settings rather than assume business terms apply.
ChatGPT has Projects, which hold pinned files. The Plus plan allowed twenty-five project files when this was researched, and the Team plan forty. Setup means uploading the files, up to ten per batch, and pasting the request into the project's instructions. The documented feature that matters here is the code tool, often called Code Interpreter or Advanced Data Analysis. It runs Python in a sandbox. When she asks it to count the files with code before writing, the count comes from arithmetic rather than the model's reading. That serves criterion one well. In her runs, two of three passed. On the third, the model skipped the code step and summarized from its reading of the files, and one enquiry was merged into another. The digest listed ten against eleven files, and the count line caught it. So the result depends on whether the code tool is used, and an instruction to use it is only an instruction. ChatGPT's Team and Enterprise plans exclude business data from training by default. Plus was listed at twenty dollars a month and Team at roughly twenty-five to thirty per user.
Gemini comes in two useful forms. The Gemini app has a very large context window, the amount of text it can take in at once, reported at one to two million tokens. Eleven small files are nothing to it. The second form is NotebookLM, Google's notebook built for answering from your own sources. Every statement it makes carries an inline citation linking back to the passage it came from, and it can sync a Google Drive folder. That suits criterion three. Checking a phone number becomes a click on its citation. Her runs passed criteria two, three and four each time. NotebookLM is built for reading and briefing, though. Shaping its output into her exact one-page format took more prompting, and on one run the count line was missing. That is a fail. A check that is not stated cannot be checked. On Google Workspace business plans, customer data is not used to train Gemini models without admin consent. The consumer Google AI Pro plan was listed at about twenty dollars a month.
Perplexity is the credible alternative outside the big three. Its Spaces hold shared files with a saved prompt and inline citations, and let you switch between models in one workspace. It is strongest when a job mixes your files with web search. Her job has no web search, and the scope forbids guessing missing details from outside sources. Its main advantage would go unused, so she notes it and moves on.
Now line them up against the test, not against reputation. On pass rate in this illustrative run, Claude Code passed three of three, ChatGPT two of three, and NotebookLM two of three. Each failure was caught by a criterion, not by her instinct. On setup, Claude Code needed nothing new. The others needed uploads and a pasted request, and that would have to be repeated daily for a scheduled service unless connected to a folder. On consistency, the failures differed in kind. ChatGPT's came from a step it could skip. NotebookLM's came from formatting.
Scheduling decides the most. Her service must run every weekday evening without her. Chat and notebook apps on a subscription are made for a person sitting there, uploading and reading. Her own digest already runs this way as a Claude Code desktop task, drawing on her plan's usage limits rather than API credit. For the chat and notebook apps, automated scheduled reporting goes through the API, where software calls the model, can count the files with code first, and can demand a fixed output format. The subscription does not include API credit, and API spending does not buy a chat seat. So "which subscription is best" is the wrong question. The useful question is which model the scheduled run should call and what that costs per run. The cost estimate already answered that. Per-run model cost is pennies on any of these providers, and her time is nearly all of the bill.
So the stack stays, for a concrete reason. Claude Code passed every criterion in every run on the job as written. It already holds the skill and the schedule. It works on files where they live rather than on uploaded copies. Switching would not lower cost, because cost is review time. It would not raise the pass rate on this test. It would mean rebuilding work that already passes.
The comparison still taught her two things worth carrying forward. First, ChatGPT's code tool showed that counting is a job for code, not for the model. That points to an improvement in her own setup: a small script that counts the files and checks the time stamp before the model writes anything. That would turn criteria one and five into hard checks rather than requests, and shrink her review. Second, NotebookLM showed how fast a reviewer moves when every claim links straight to its passage. Her claim-check list does a rougher version of that, and tightening it to name the exact line would cut her minutes further.
She saves the folder copy, the three runs from each tool and her pass-or-fail marks into the project. When a provider changes its plans or features, she can run the same folder and the same five criteria again and see whether the answer changes. That record is the asset. The test and its results stay with her, whatever tool she ends up using.
