GnothiGnothi
SeriesFieldsCommunity
Sign inGet started free

Picking Your Assistants

The founding charter

An AI assistant, plainly defined, is a general-purpose text tool that drafts on request, reads what you feed it, sometimes browses live, and occasionally makes images — but it knows nothing about your company until you tell it. Because the underlying tools reshuffle constantly, the case here is for judging them by the marketing jobs they do rather than by whichever demo impressed most recently.

Three mechanics get explained before anything else: reusable saved workspaces that hold instructions and files so you never re-explain your company from scratch; the built-in randomness that means identical requests produce different answers each time (so one bad output proves nothing, and a good one should be saved immediately); and the way long conversations drift as early instructions lose out to everything said since — fixed by starting a fresh chat rather than arguing a wandering one back on course.

From there, a tour of the field, weighted by what each tool is actually for: ChatGPT's twenty-dollar tier for its document editor and deep research mode; Claude for prose and style-guide discipline, with the flat limitation that it makes no images at all; Gemini's enormous context window for whole-book, whole-transcript analysis, offset by stiff default prose and a tendency to miss conclusions buried in huge documents; Perplexity as a citation-anchored search engine rather than a writing tool; and Copilot's advantage of reading a company's own email and files from inside Office, traded against smaller context limits.

The method for choosing: run one real, plainly-worded task from your actual week through two assistants side by side, judge on a checkable claim, whether it sounds like your company, and how much editing shipping it would take, then keep one main tool and one cross-checker. Kept alongside this is a way to catch the "generic-AI tell" — throat-clearing openers, tidy three-part lists, confident sentences with no checkable specific in them — using a simple test: could a competitor publish this exact sentence on their own site unchanged? If yes, cut it.

The show's news beat

Ad-platform moves worth acting on: Google is auto-upgrading legacy dynamic search ads into its newer AI-driven search product, testing journey-aware bidding that optimizes toward mid-funnel milestones, retiring campaign-level language targeting for search and AI-max campaigns in favor of automatic language matching, and rolling out prompt-driven dashboards for ads and analytics.

On the assistant side, OpenAI opened a self-serve ads manager with cost-per-click bidding and no minimum spend, and shipped a new flagship model to paid tiers with a lighter version following for free users. HubSpot added two-way ChatGPT and Claude connectors for managing contacts and deals from inside the chat tool, plus a canvas sidebar for inline drafting. Salesforce rebuilt its marketing suite around autonomous campaign agents and two-way customer messaging.

The visibility picture: AI summaries are cutting organic click-through roughly in half, but pages that get cited by name still pull a third more clicks than those that don't — making citation itself a channel worth tracking with a cheap prompt-monitoring tool or a simple monthly check of one real buying question against two assistants.


There is a moment most marketers have had by now. You paste a request into a chat box, you get back nine hundred words of confident, well-punctuated nothing, and you sit there wondering whether the problem is you, the tool, or the whole idea.

It is usually the setup. And the setup is fixable in an afternoon.

So let us start with what these things actually are, because the word "assistant" gets thrown around until it means nothing. An AI assistant, for our purposes, is a general-purpose text tool you talk to in plain language. You describe a job, it produces a draft. It can read documents you give it, it can often look things up on the live web, and some of them can make images. It is not a marketing product. It has no idea who your customers are. It is a very capable, very fast writer who has never worked at your company and never will unless you tell it everything.

That last part matters more than any feature comparison, and it is why this chapter is built the way it is. The tools reshuffle constantly. Prices move, tiers get renamed, a model that was clearly best in March is clearly third best in June. If you choose your assistant because a demo impressed you, you will be choosing again in ten weeks. So you do not choose by demo. You choose by the marketing jobs you actually do, and you keep the method you used to judge, because the method survives a tool swap and the standings do not.

A few pieces of vocabulary, quickly, because the rest of the chapter leans on them.

A chat window is one conversation. You open it, you type, you get answers, you close it. Nothing carries over unless the tool has a memory feature turned on. A saved setup is different — most of these tools let you create a reusable workspace where you park instructions and reference files once, and every conversation you start inside it inherits them. On some platforms that is called a project, on others a custom assistant. The names change; the idea does not. The idea is that you should never re-explain your company from scratch, and if you find yourself pasting the same three paragraphs of background into every new chat, you are working the wrong way.

Second thing to know: the same request can come back different twice. Ask for five subject lines, delete the answer, ask again with identical wording, and you will get five different subject lines. This is not a bug and it is not the tool being unreliable. These systems generate text one piece at a time, and at each step they are choosing from a range of plausible next words rather than reciting a fixed answer. A little randomness is deliberately built in, because without it the output would be flat and repetitive. What this means practically is that a single bad output tells you almost nothing. Run it twice before you conclude the tool can't do the job. It also means you cannot rely on a lucky result you didn't save — if a draft comes back good, keep it, because you may not get it back.

Third: long conversations drift. You start a chat about a landing page, it goes well, you keep going, and forty exchanges later the thing is writing in a voice you never asked for and has quietly forgotten a constraint you set at the top. The reason is that the assistant reads the conversation back to itself every turn, and there is a ceiling on how much it can hold. That ceiling is often described in tokens, which are just the chunks of text these systems count in — roughly speaking, a token is a short word or a piece of a longer one. When the conversation gets long, the early instructions have to compete with everything since for the assistant's attention, and they lose. The fix is unglamorous: when a chat starts feeling mushy, open a fresh one and restate the job. Do not try to argue a drifting conversation back on course. Start over. It takes thirty seconds and it works every time.

Now, the field. Here is where things stand, with the caveat that "where things stand" has a short shelf life and you should check the current pricing page before you put a card in.

The one most people meet first is ChatGPT, from OpenAI. There is a free tier with lighter models and a rolling cap on messages, which is enough to see what the thing does but not enough to work in. The entry paid plan runs about eight dollars a month and gives you higher limits and more memory, though it holds back the deeper reasoning models. The tier a working marketer would actually start on is the twenty-dollar plan, which opens the flagship models, the deep research mode that goes off and browses for a while before reporting back, image generation, and a side-by-side document editor where the draft sits in one panel and you talk to it in the other. That editor is the thing to notice. For iterative drafting — write it, now tighten paragraph three, now cut a third of it — the side-by-side is a genuinely different experience from scrolling a chat log. Above that, the heavy tiers run about a hundred and about two hundred a month, buying you five times and twenty times the usage room respectively. Most marketers do not need them for a long while.

Claude, from Anthropic, is the other one worth a serious look, and it prices the same at the working tier — twenty a month, or about seventeen if you pay for the year up front. Its free plan will let you feel the model out but the limits bite fast, sometimes inside ten substantial exchanges at a busy hour. What Claude is genuinely good at is prose. Nuanced drafting, long analytical pieces, and — this is the part relevant to us — holding to a style guide without turning stiff and formulaic about it. It also has the saved-workspace feature, where you can load reference material once and work against it, plus a side panel that renders documents as it writes them. Its real limitation is that it cannot make images at all. Not badly, not sometimes. Not at all. If your week has visual work in it, that is a hole you will have to fill elsewhere.

Gemini, from Google, is the long-document specialist. The free tier gets you the faster, lighter models with live search behind them; the flagship models moved behind the paid plans in the spring. There is a cheap entry plan under five dollars, and the one you would actually work on is about twenty a month, bundled with a couple of terabytes of storage. What that buys is a context window running from one to two million tokens, which in practical terms means you can hand it an entire book, or hours of video, or a stack of technical PDFs, and ask questions across the whole thing in one go rather than chopping it into pieces. Its live web lookups are strong, unsurprisingly, since Google's search index is sitting right behind them. Its weakness is voice. Left alone, Gemini's everyday drafting comes out stiff and repetitive, fond of the same formatting patterns over and over. There is also a subtler catch: across those enormous contexts, it can lose a specific point when finding it requires reasoning rather than matching words. It will find the sentence you half-remember. It may miss the conclusion you needed it to draw.

Perplexity is a different animal — a search engine that answers in prose. The free tier gives you unlimited basic searches and a handful of deeper ones per day; twenty a month lifts that to unlimited deep searches within fair-use limits, twenty research runs a day, and file uploads. Its strength is factual reliability, because every answer is anchored to a live crawl with dense inline citations you can click and check. Its weakness is that it is not a writing environment. Ask it to re-architect a long draft and you will be disappointed. It investigates; it does not compose.

Microsoft's Copilot sits inside the software many companies already run. The consumer plan at twenty a month puts it into Word, Excel, PowerPoint and Outlook, plus a monthly allowance of image credits in Microsoft's design tool. The workplace version runs from about twenty-one to thirty per person per month as an add-on to a business licence, and its real trick is that it can read your organisation's own email, calendars and files. That grounding is powerful. The tradeoff is smaller context limits inside the Office apps than you get on the standalone web tools, weaker long-document analysis, and some wobbliness on complicated multi-step reasoning once you step outside Microsoft's world.

That is the field. Notice what the descriptions have in common: each one is stated as a job, not a score. Long documents. Live lookups with citations. Prose that holds a voice. Images. Work inside the files you already have. When the next model lands and the rankings shuffle, those jobs will still be the jobs.

Which brings us to the workflow, and it is one you can run today.

Pick one real task from your actual week. Not a test prompt, not something clever — a thing you have to produce anyway. A product launch email. Three ad variations. A landing page section you have been avoiding. Write it out as one plain request, in ordinary language, saying what it is for and who it is for. Do not engineer it. Do not stack it with instructions. A plain request is the point, because you are not testing your prompting here, you are testing the tools.

Now run that same request in two assistants, side by side, in two browser tabs. Same words, same task, no edits between them.

Then score the two outputs against three questions, and score them honestly.

First: is a single claim in there checkable? Pick one factual-sounding sentence and try to verify it. Not all of them — one. If the assistant has invented a statistic, a customer quantity, a feature you don't ship, you want to know that on your test task rather than on something you were about to publish.

Second: does it sound like our company, or does it sound like nobody? Nobody is the common answer, and it is not a small problem.

Third, and this is the question that actually decides it: how much editing would shipping this really take? Not "is it good." Count the passes. If output A is prettier but needs a full rewrite, and output B is plainer but needs three sentences fixed, B won and it wasn't close.

From that test, keep one as your main desk tool and one as your cross-checker. Two, and here is why not five. Every assistant you add is another set of saved instructions to maintain, another place your brand context lives in a slightly different half-updated form, another interface whose quirks you have to learn. You will end up doing your real work in whichever one your hands go to by habit, and the other four will hold stale versions of your voice guidelines. Two is the minimum that lets you catch a bad answer — you take a claim the main tool made and ask the second one about it cold — and it is the maximum you will actually keep current. If your week has heavy visual work in it and your main tool cannot make images, that changes which two, not how many.

Now the pitfall, and it is the one that will bite you first, because it already has.

Call it the generic-AI tell. It is not that the writing is bad. It is that the writing is unmistakably machine-shaped, and your reader hears it even when they cannot name it.

It has a sound, and once you know the sound you cannot unhear it. There is the throat-clearing opener — the sentence that announces the topic before saying anything about it. In today's fast-moving landscape, businesses are increasingly recognising the importance of. Nothing has happened yet and you have spent a line. There is the tidy three-part list, everywhere, three items every time, balanced and rhythmic and always slightly too neat to be a real observation. And there is the confident sentence with no specific in it — the one that sounds like a claim, reads like expertise, and contains no number, no name, no example, nothing anybody could check or disagree with. Our approach helps teams work smarter and achieve better results. That sentence is a hole with punctuation around it.

Here is the recognition test, and it takes twenty seconds. Read the draft aloud. Then take any sentence in it and ask whether your closest competitor could have published that exact sentence, word for word, on their site, without changing a thing. If the answer is yes, that sentence is doing no work for you. It is not wrong. It is just not yours, and it is not about your product, and it could be deleted without loss — which means it should be.

Read your test outputs aloud right now and you will find several. That is expected. It happens for two reasons, and neither is that you picked the wrong tool. The assistant does not know what your company sounds like, because nobody has ever told it — and your plain request, which was the right thing for a fair test, gave it almost nothing to aim at. It filled both gaps with the average of everything it has ever read. The average of everything is exactly what generic sounds like.

Those are the two things to fix, and they are the next two things we take up: giving the assistant a description of your voice it can actually work from, and building the request properly so it has a target. But you cannot fix either until you have a bench, so go run the test.

Now, the news, because the products underneath all of this moved a great deal this year and some of it changes your week.

The ad platforms first, since that is where money is actually moving. In the middle of April, Google announced that its legacy dynamic search ads are being retired and auto-upgraded into its newer AI-driven search product. If you are still running the old format, that migration is happening to you whether or not you plan for it, so go into your campaign settings and find the migration date field before the date finds you. In early May came a bidding change worth more attention than it got: a journey-aware bidding beta for search campaigns on target cost-per-acquisition, which optimises against mid-funnel milestones — a qualified call, a completed demo — rather than only the front-end conversion click. That is closer to how you actually think about a funnel. To feed it, you need those intermediate steps set up as non-biddable conversion actions under your measurement settings, so do that first or the models have nothing to learn from. Also, from late September, campaign-level language targeting disappears for search and AI-max campaigns in favour of automatic ad-language matching — if you were using language targeting as a rough proxy for geography or audience, that lever is gone. And in August, Google shipped prompt-driven dashboards across both its ads and analytics products, where you type what you want to see in plain language rather than building the report. Worth ten minutes on the dashboards tab.

On the assistants themselves, OpenAI has been busy on the money side. The eight-dollar tier went global in mid-January, and alongside it the company began testing advertising inside the interface on its free and cheaper tiers in the United States. Then in May, ad onboarding opened up broadly through a self-serve ads manager in beta — cost-per-click bidding, no minimum spend, product feed ads for e-commerce, and first-party custom audience targeting. That is a genuinely new inventory surface, and the low barrier means testing it is cheap. If you sell products, the smallest next step is to log into the ads manager dashboard and see whether your existing product feed will upload. Model-wise, a new flagship reached the paid and business tiers in early July, with a lighter counterpart rolling out to free and cheap-tier users from the start of August — relevant chiefly because if you standardised your team on specific reasoning settings, go check they are still where you left them.

In the customer-record tools, HubSpot spent the year making its AI reach outward. In February it shipped two-way connectors for both ChatGPT and Claude, which means you can create and update contacts, deals and tasks from inside the chat tool without opening the customer platform at all. That removes a real manual step. It needs a super-admin to reauthenticate permissions in the app marketplace settings before anyone can link it, so if it does not appear for you, that is why. In May a canvas sidebar arrived in its assistant for inline document and email drafting with version control, and in July the agent products were reorganised and renamed, now running on the platform's AI credits — which is the part to look at, because credit-based pricing changes the arithmetic of anything you were planning to run at volume.

Salesforce, meanwhile, rebranded and rebuilt its marketing suite across its last two big conferences, most recently in early June. The shipped pieces include autonomous campaign-creation agents, two-way email and text channels that handle incoming customer replies directly, real-time page personalisation, and multi-touch attribution — attribution being the question of which touches get credit for a sale. If you are on that stack, the campaign playbooks and the conversational response rules are configured inside the core org, and the response rules are the ones to look at first, because an agent answering a customer reply is a brand-safety question before it is a productivity one.

Which leaves the visibility check, and this is the fixture that will run every time, because search is changing under us faster than anything else here. Treat everything in it as a moving snapshot.

The headline is that clicks are draining out of search. Clickstream measurement suggests that out of every thousand American Google searches, only about three hundred and sixty produce a click out to the open web — so roughly two thirds of searches now end without anyone leaving. Where an AI summary appears at the top, one survey puts the no-click rate at eighty-three percent, and in Google's conversational mode another measurement has it at ninety-three. A study across three hundred thousand keywords found that when an AI summary is present, the click-through rate for the top organic result falls by about fifty-eight percent — from roughly seven and a half percent down to under four. Another, tracking three thousand informational terms across twenty-five million impressions, saw click-through fall from about one and a half percent with no summary to seven tenths of a percent when the summary cited the brand, and about half a percent when it did not. Publishers are feeling it hardest: one portfolio of sixty-four publisher domains measured a forty-two percent organic traffic drop across two years.

But look again at those last two numbers, because the gap between them is the whole opportunity. Being cited was worth roughly a third more organic clicks than not being cited on the identical page, and nearly double the paid click-through. Citation is now a channel.

And getting cited does not work the way ranking works. One analysis of a hundred and twenty-nine thousand domains found that around eighty percent of the pages ChatGPT cites do not rank in Google's top ten at all; another put cited pages at position twenty-one or worse nearly ninety percent of the time. What correlates with getting pulled into an answer is being mentioned around the web — that relationship measured far stronger than raw backlink volume. Freshness matters too: something like ninety-five percent of ChatGPT's citations came from material published or refreshed within the previous ten months.

Citations also concentrate hard. Across a study of six hundred and eighty million of them, the top fifteen domains took about sixty-eight percent of everything. Reddit and Wikipedia are the most pervasive sources across engines, with Reddit taking around forty percent of multi-engine citations on general research questions — but that flips on buying-intent queries, where Reddit and Quora together account for under two percent of product and vendor recommendations, and the engines lean instead on review portals, comparison hubs and vertical directories. Each engine has its habits: one leans heavily on wire services and financial press, another on broadcasters, another on business magazines. Google's own summaries cite video more than any other single domain in the United States, at around a fifth of top-source share.

The surfaces are shifting too. Google's summaries began last year overwhelmingly on informational questions — over ninety percent — and by year's end that had fallen to under sixty as commercial, transactional and navigational queries pulled them in. Product comparison modules, citation pills, and checkout links now live inside those answers.

If you want to watch your own brand, monitoring tools exist and are reasonably cheap to try. The lightest track a fixed list of prompts daily across the major engines from around thirty dollars a month. Mid-market options that also watch competitor prompt share start near a hundred, with the more serious ones running into the several hundreds, and the established search platforms now sell AI visibility as an add-on from about a hundred to two hundred a month.

Your smallest next action, though, costs nothing. Pick the one question a real prospect asks right before they buy from you — the actual words, not a keyword. Type it into two assistants this week. See who gets named. That answer is your baseline, and next month you can tell whether it moved.