OCDevel AI for Marketers

Picking the Assistant You'll Actually Ship With

The founding charter

This one walks through what actually happens when you ask an AI writing tool for help, and why the result so often reads like it was written by nobody in particular. It breaks down four ideas worth carrying forward: that these systems generate text one piece at a time based on what fits, not what's true; that the same prompt run twice can produce very different answers, so no single output proves anything; that "context" — everything visible in the conversation — is the entire world the assistant has to work with, and gaps in it get filled with generic guesses; and that saved, standing instructions (projects, custom instructions, gems, collections) can hold your brand's specifics in context permanently instead of retyping them each time.

From there it lays out a concrete way to choose between tools: pick one real, overdue piece of work, write the brief once, and run it through two assistants — say ChatGPT, Claude, Gemini, or Perplexity — a couple of times each, side by side. It profiles each one's character (Claude for prose with less filler, ChatGPT for structural control and deeper research runs, Gemini for size and search-grounded answers, Perplexity for sourced, checkable research) and gives three things to mark on the output: shaky facts, generic "nobody" sentences, and genuinely useful structure. It closes on the sharpest risk in all of this — confidently stated wrong information, from invented statistics to phantom quotes to imagined product features — and the rule that fixes it: nothing unverifiable leaves the document.

The show's news beat

Two updates worth knowing. OpenAI recently lifted the cap on text conversations for free and entry-tier users and added a switch to slow the model down for harder problems, while Anthropic shipped a much larger-context, cheaper flagship model. Separately, recent industry data paints a tightening picture for visibility in search: Ahrefs found top organic results lose over half their clicks when an AI summary appears above them, and a CMSWire report puts the median brand's AI-answer citation rate around sixteen percent, with an actual link only about six percent of the time.


Here is the thing that probably happened the last time you tried this.

You had something you owed someone. A launch email, a landing page, three ad variants, a blog post that was already late. You opened a chatbot, typed something like "write me a launch email for our new analytics feature," and hit enter. Out came six paragraphs in about four seconds. Clean grammar. Nice rhythm. A subject line with a colon in it. And when you read it back, it sounded like it had been written by absolutely nobody. No company, no product, no person. Just the average of everything ever written about analytics features.

At that point most people go one of two ways. Some look at the speed and conclude this changes everything, and start pasting output straight into their email platform. Others look at the blandness and conclude the whole thing is overhyped, and go back to writing by hand. Both reactions come from the same missing step. Neither person compared anything. They ran one request through one tool, once, and then formed a permanent opinion about a category of software.

So that is where we start. Not with what these systems are, or who is winning, or which one scored highest on some benchmark last month. We start by getting you a working setup and one real comparison, today, using work you already owe somebody. By the end of this you should have a primary assistant, one alternate, and a piece of evidence about the difference between them that came from your own brand rather than from a list on the internet.

Before the workflow, though, you need about four ideas. Not a course on prompting — that is its own subject and it gets its own hour. Just enough to read what comes back and understand why it looks the way it does. I am going to go through them one at a time and then get out of the way.

The first one. When you type a request and the assistant answers, it is not looking anything up in a filing cabinet of correct answers. It is producing text one small piece at a time, and each piece is chosen based on everything sitting in front of it — your request, plus whatever it has already written. It is extremely good at producing text that fits. That is the whole trick, and it explains almost every behaviour that surprises people. It explains why the output is fluent. It explains why the output is generic, because the text that fits best is usually the most average text available. And, as we will get to, it explains why the output can be confidently, smoothly, completely wrong. Fitting and being true are two different jobs, and only one of them is what the machine is built to do.

Second. Ask the same thing twice and you will get two different answers. Not slightly different — sometimes a different structure entirely. This is not a bug and it is not the assistant having an opinion about you. There is deliberate randomness in how each next piece of text gets chosen, so the same input can branch down different paths. The practical consequence is what matters here. It means a single bad output is not proof the tool is bad, and a single brilliant output is not proof the tool is good. You cannot judge anything from one run. When you compare two assistants later in this episode, you are going to run each brief a couple of times, because otherwise you are comparing coin flips.

Third, and this is the one that pays off for years: context. Context is everything the assistant can see right now, in this conversation. Your instructions, the examples you pasted, the document you uploaded, the last eight messages of back-and-forth. That is it. That is the entire world the thing is working from. Anything not in there is not available to it, no matter how obvious it is to you. Your positioning, the fact that your customers are procurement managers and not developers, the phrase your founder refuses to let anyone use — if it is not in the context, it does not exist, and the assistant will fill the hole with the most average guess available. Which is exactly how you get copy that sounds like nobody. Nobody was in the context.

There is a limit on how much fits, and it is generous now. The bigger assistants can hold something in the range of a couple of hundred thousand words of material in a single session, and the largest can take a very long document or a stack of documents at once. So the constraint on most marketing work is no longer size. It is you remembering to put things in.

Fourth. There are two places an instruction can live, and knowing the difference saves you an enormous amount of retyping. A one-off ask lives in the message. It applies to this conversation and then it is gone. A saved instruction lives at the account or project level, and it applies every time, without you pasting it. Every major assistant has some version of this. One calls them custom instructions and projects. One calls them projects with their own persistent instructions and reference files. One calls the saved-persona version gems. One calls the saved-source version collections. Different names, same idea: a standing brief that is always in the context so you do not have to keep saying "we are business-to-business, we do not use exclamation points, our product is not a platform."

Do not build an elaborate one yet. Later in this series we spend real time on grounding an assistant in your brand properly, and what you write today will be replaced. For now just know the shelf exists, and know that when your output is bland, the fix is almost always more context rather than a better tool.

That is the whole set. Now the part you can copy.

Running one brief through two assistants

The workflow is this. Pick one primary assistant and one alternate. Take one real asset you owe someone this week. Run the identical brief through both. Put the two outputs side by side and mark three specific things. Then decide.

Start with the picking, because that is where people stall the longest for the least benefit.

Right now there are four general assistants a working marketer would seriously consider for writing and research. There is ChatGPT from OpenAI. There is Claude from Anthropic. There is Gemini from Google. And there is Perplexity, which is a slightly different animal in that it is built primarily around searching the live web and showing you where each claim came from.

The money is remarkably uniform. The main paid tier on ChatGPT, Claude, Gemini and Perplexity all sit at right around twenty dollars a month, give or take a few cents, with a small annual discount on a couple of them. All four have a free tier that is genuinely usable for trying things and genuinely frustrating for daily work, because the limits bite. Above the twenty-dollar tier there is a heavier tier on each — around a hundred dollars a month, and in some cases two hundred — which mostly buys you more capacity and larger working room rather than a fundamentally different tool. There is also a cheap middle rung on a couple of them, in the five-to-ten dollar range, if you want more than free without committing.

The important thing about that price uniformity is what it tells you: cost is not the deciding factor. Twenty dollars a month is not a procurement decision. Two subscriptions is forty dollars a month, which is less than most marketers spend on stock photography. So the honest recommendation is that you pay for two of them, not one, and we will get to why in a moment.

What actually differs is character, and here is roughly how they differ for our kind of work.

ChatGPT is the strongest all-rounder for structural control. If you want to say "give me this as five sections, with the objection handled in section three, and make the tone a notch drier," it takes direction well and revises without falling apart. It has a live web search tool and a deeper multi-step research mode on the paid tiers, where it goes off and works through a question over several minutes rather than answering immediately. Those deeper runs are metered — a handful a month on the standard paid tier — so you save them for things that deserve them.

Claude is the one most people reach for when the writing itself is the deliverable. It tends to produce prose with less filler and fewer of the tells we will spend a whole episode dismantling — the "in today's fast-paced landscape" opener, the triplet of adjectives, the closing paragraph that restates the whole thing. It holds very large amounts of uploaded material, which makes it good for "here are our last twenty blog posts, write the twenty-first." It leans less on live web browsing, so it is a better drafter than it is a researcher.

Gemini has the largest working room of the group by default and is deeply wired into Google's own search, so its research answers come back grounded and current, and its deeper research mode is generous — you can run a lot of them a day on the standard paid tier rather than a few a month. Its drafting voice is factual and clipped. Some people find that a feature and some find it flat.

Perplexity is not really competing to write your email. It is competing to answer a question with sources attached, inline, so you can click each claim and check it. For competitor research, for "what do people actually complain about with this category," for anything where you need to be able to show your work, it is the one you want open. It also lets you switch which underlying engine answers, which is a nice way to sample the others before you subscribe to them.

So which is your primary? The rule is unglamorous: pick for the work you do most. If most of your week is producing written assets, start with the one known for prose. If most of your week is research, briefs and competitive digging, start with the search-first one or the search-grounded one. And then keep a second one open as your second opinion, because the single most useful move in this whole episode is asking the same question twice in two different places.

Now, the actual test, and please use real work. Not a sample. Not "write a haiku about our product." A real asset with a real deadline, because a fake brief produces a fake result and teaches you nothing. Say it is a nurture email — one of a sequence of emails that walks a lead from mildly interested to ready to talk.

Write the brief once, in a text file, so it is literally identical in both places. Say who it is for, in the specific. Say what you want the reader to do next. Say the format and the rough length. Say something about voice, and better yet paste in two paragraphs of something you have already published that sounds right. Say the constraints — what you cannot claim, what words are banned, what the legal team will kill.

Then paste it into both assistants. Run it twice in each, because of the randomness. That gives you four outputs and about ten minutes.

Now put them side by side and mark three things. Only three.

First, mark where the facts are shaky. Every number, every claim about your product, every "customers report" and every "studies show." Underline each one and ask whether you could produce the source in thirty seconds. Not whether it sounds plausible — whether you could produce it.

Second, mark where the voice is nobody's. Find the sentences that could appear in a competitor's email with the brand name swapped and nothing else changed. Those sentences are not writing, they are packing material. Some assistants produce far more of it than others, and this is the difference you will feel most in daily use.

Third, mark where the structure is genuinely useful. Sometimes the words are worthless and the skeleton is excellent — it opened on the objection you were avoiding, or it put the proof before the promise, or it found an order for the three benefits that is better than yours. That structural gift is often the real value, and it survives you rewriting every sentence.

Now compare. One of them will be better at your work. Not better in general, better at yours. That is your primary. Keep the other subscribed as your alternate, and use it deliberately: when your primary gives you something that feels thin, run the same brief through the other one and see whether the problem is the tool or the brief. Nine times out of ten it is the brief. Knowing that quickly is worth the forty dollars.

The obvious objection here is why not just read a comparison and skip all this. Two reasons. The first is that the standings change monthly, so any ranking you read is describing a world that has already moved. The second, and this is the real one, is that a ranking measures average performance on someone else's tasks. You are not producing average work for a general audience. You are producing work in a particular voice, for a particular buyer, in a particular category, with particular things you are not allowed to say. Nothing about a leaderboard tells you which assistant handles that. Ten minutes with your own brief does. This comparison habit — run the same real work through two tools and read the difference — is the one skill in this episode that will still be useful when every product name in it has changed.

Which brings us to the thing that will bite you this week, probably today.

The confident wrong answer. You will ask for a landing page and get back a sentence saying that seventy-three percent of buyers in your category abandon their evaluation for reasons of unclear pricing. It will read beautifully. It will sit in exactly the right place in the argument. It will have the shape of a real statistic — an odd number, not a round one, the kind that sounds sourced. And it will not exist. There is no study. There is no seventy-three percent. The system produced text that fit, and a plausible number fit.

Notice that this is not the same problem as blandness. Bland copy is an inconvenience; you rewrite it. A fabricated statistic is a brand risk. If it ships on your site, you have published a false claim under your own logo. It gets quoted back to you in a sales call. It gets screenshotted. Depending on your category and what the claim is about, it may be a compliance problem rather than an embarrassment. And the person who catches it will not be you, because you already read it and it sounded right.

The same failure shows up in shapes that are easier to miss than a number. A quote attributed to someone who never said it. A named report that does not exist, from an organisation that does. A feature described in your own product that you do not actually have — that one happens constantly, because the assistant knows what products like yours usually do and helpfully assumes. A legal or regulatory statement delivered with total assurance. A competitor's pricing, stated as fact, invented entirely.

The recognition habit is one sentence, and I want you to make it a rule rather than an intention. Anything you cannot check yourself does not leave the document. Not "gets flagged for later." Not "probably fine." If you cannot verify it right now, you either verify it or you delete it, and if the sentence needs a number to work then you go and find the real number and rebuild the sentence around it. This is also the single best argument for keeping a search-first assistant in your kit, because when it makes a claim it hands you the source alongside it, and you can click through and see whether the source says what the summary says. Often enough it does not.

One more thing about the shakiest ground. The assistant is least reliable exactly where you are most tempted to lean on it: anything recent, anything numerical, anything specific to your own company. Recent because its knowledge has an edge and it does not always know where the edge is. Numerical because a wrong number fits a sentence just as neatly as a right one. Specific to you because you are not in the context unless you put yourself there.

So here is where I am going to leave you, deliberately short of finished.

Your first prompt today should be thin. Say who the reader is, what you want, what format, roughly how long, and paste one example of your own writing. That is enough to run the comparison and enough to get one asset out the door. It is not enough to get consistently good work, and you will feel that gap by the second or third attempt — the output will be broadly right and specifically wrong, close to your voice without being it. That gap is the whole subject of the next stretch of this series, and it is worth arriving there with the frustration fresh.

Because the shape of the climb is this. First we produce good, on-brand work by hand, one asset at a time, with you doing every step. Then we ground the assistant in your actual brand and wire it into the tools you already work in, so it stops guessing. Only after that do we hand multi-step work over to something that runs on its own while you set the goal and approve the output. In that order, and not faster, because every stage rests on the one underneath it. Today's job is the first rung: two subscriptions, one real brief, four outputs, three marks in the margin.

What shipped, and what it means for your visibility

Two things worth your attention, and both have something you can do about them this week.

The first is at the desk. In early August, OpenAI reshuffled what its consumer tiers actually give you. The headline for anyone on the free or entry plan is unlimited text conversations on the new default model, plus a button that switches it into a slower, more careful reasoning mode for hard problems. Caps still apply to file uploads, image generation and stored library space — that is where the free tier still pinches — but the basic act of drafting and iterating is no longer rationed. On the paid side, the model behind chat picked up a slider that lets you dial how much thinking it does before answering, and Anthropic separately shipped a top-end model in late July with room for something on the order of a million words of context and a much longer maximum answer, at roughly half the price its flagship tier used to command.

Why a marketer cares: the free-tier change means the comparison test in this episode now costs you nothing to start, and the reasoning toggle is the difference between a fast bland draft and a considered one. Smallest next action — go and find the reasoning control on whatever you are using, run your brief once with it off and once with it on, and see whether the difference is worth the wait. On most marketing briefs it is, and most people never touch it.

The second is the standing check on whether your brand is visible inside AI answers, and the picture as of this summer is not comfortable. Work published in June by SparkToro with Similarweb found that a bit over sixty-eight percent of United States Google searches in the first four months of this year ended without anyone clicking anything at all — meaning fewer than three hundred out of every thousand searches now reach the open web, down from about four hundred two years ago. Ahrefs reported earlier in the year that when an AI summary sits at the top of the page, the top organic result loses well over half its click-through rate. Not everyone agrees on the aggregate direction — a Semrush analysis of ten million keywords late last year found zero-click rates on affected queries ticking slightly down rather than up — so treat the exact magnitude as contested and the direction as real.

Meanwhile the traffic coming directly from AI tools is still tiny, roughly one percent of publisher traffic by one late-year estimate, though growing fast, with Gemini referrals reportedly up several hundred percent year over year. The interesting part is quality rather than quantity: Similarweb data published this year put the conversion rate on AI-sourced visits at around eleven percent against a bit over five for organic search. Small stream, much better water.

And on citations, which is the thing you can actually influence: research published in July by Webflow with CMSWire found the median company appearing in something like sixteen percent of relevant AI answers, and getting an actual linked attribution only about six percent of the time when mentioned. Overlap between engines is poor — one analysis found only about eleven percent of cited domains were cited by both ChatGPT and Perplexity on the same queries — and the strongest single predictor of getting cited was plain branded search volume.

Smallest next action, and do this one yourself rather than trusting my numbers: write down the five questions a real prospect asks before buying from you, ask all five in two different AI assistants, and record whether you appear, whether a competitor appears, and whether anyone gets a link. That is your baseline. If you want it tracked continuously rather than by hand, the established search suites now monitor AI-answer visibility alongside classic rankings, and there is a growing crop of dedicated tools that watch brand mentions prompt by prompt across several assistants. Check one, but check yourself first — the manual version takes twenty minutes and tells you more than a dashboard you have not learned to read yet.