Picture a regional automotive consignment business. The kind where a local consultant drives the neighbourhood, shakes hands at the car club, and earns trust one Saturday morning at a time. They have had a flyer built. Brand strategy mapped, audience identified, every element placed for a reason. The consultant’s photo is front and centre because in a provincial market where people buy from people they recognise, his face is the trust signal.
The business owner does something perfectly sensible. They paste the flyer into ChatGPT and ask for feedback. A second opinion. Why wouldn’t you?
Here’s what comes back. “Kill 50% of the text.” “Think like Porsche.” “Luxury brands don’t explain, they assert.” Strip the consultant’s photo down to a tiny name and phone number. Replace the service explanation with two words and a cinematic photo spread.
The advice sounds like a senior creative director at a top-tier agency. It is confident. It is specific. It is polished. And it is catastrophically wrong.
This is not a Porsche print campaign. It is a one-page flyer for a used-car consignment service in a small city. The consultant’s photo is not clutter. It is the entire strategy. The model that said “shrink the broker card” did not know that. It could not know that. It had never heard of the business, did not know the audience, and had no idea why every element was where it was.
It pattern-matched against every brand playbook it had ever seen and told a small business to be Porsche.
The pattern has a name
What happened to that flyer is happening in businesses everywhere, and it has a name. Context collapse. Not the social media kind. A newer, more commercially dangerous variety. It is what occurs when strategic work built with deep business context gets run through a system that has none.
The sequence looks like this. A business owner gets work from their agency or consultant. Copy, design, strategy, whatever it is. The work reflects weeks of accumulated context: brand positioning, audience research, competitive analysis, specific business goals. Then the owner opens ChatGPT, pastes the work in, and types some version of “how can I improve this?”
The model does what models do without context. It reaches for the average of everything it has been trained on. Best practice. Industry norms. Safe, credible-sounding advice that could apply to any business in any market. The output sounds like expertise. It reads like a creative brief from someone who knows what they are doing.
But it is advice from a stranger at a bus stop. A very articulate stranger who has read every marketing textbook ever published, but who does not know your name, your customers, or your Tuesday afternoon problem.
The "second opinion" is now the default behaviour
This is not a quirk of one client. It is the dominant way small business owners interact with AI.
Anthropic’s Economic Index, which analyses over a million conversations on Claude.ai, found that augmentation patterns account for roughly 53% of conversations on the platform. Augmentation is a broad bucket. It covers bringing existing work back for review, iteration and validation, and it also covers learning and explanation, which makes up about half the category on its own. The point still holds. Most of the time, people are not asking AI to create something from scratch. They are asking it to judge, explain or improve something that already exists.
Microsoft’s Work Trend Index, in its May 2024 edition, drew a line between two types of users. Power users treat AI as a thought partner. Conversational, iterative, willing to push back. They were 49% more likely to pause before a task and ask whether AI could help with it at all. That is not restraint. It is deliberation. They think about the tool before they reach for it. Casual users treat AI as a vending machine. Input, output, done. No deliberation happens anywhere in that loop.
Almost nobody has been taught to do it properly. Google and Ipsos surveyed 4,464 American workers in December 2025 and found that only 14% had received any AI training from their employer in the past year. The Marketing AI Institute’s 2026 State of Marketing AI survey, which runs on a self-selected sample of people who already care about AI, found that 32% of respondents get no AI training at all and 53% get so little that it amounts to none. Note that both figures measure AI training generally. Neither measures training on prompting specifically, which is almost certainly rarer still.
So when a business owner sits down to get that second opinion, they are using the tool at the lowest available level of sophistication. One-shot prompts with no context, no constraints, no strategic framing.
They are asking a genius with amnesia for advice.
What the "improvement" actually costs
The gap between contextual and context-free AI output is not subtle. It has been measured.
MIT’s Initiative on the Digital Economy ran a study with roughly 1,900 participants and found that about half the performance improvement from a model upgrade came from how users adapted their prompts, not from the newer model itself. Worth stating the limits plainly: that result comes from a fixed-criteria image replication task using one specific upgrade, and it did not hold the same way in the open-ended version of the task.
The same study tested something closer to home. The newer model shipped with a built-in layer that silently rewrote users’ prompts before they reached the image engine. Users did not know it was happening. That automatic rewriting erased 58% of the benefit the upgrade would otherwise have delivered. The rewritten version still beat the older model, so this is not a story about AI making things worse outright. It is a story about a generic rewriting layer quietly eating most of the value, while the person at the keyboard has no way to see it happen.
The BCG and Harvard study tells the sharper version. When 758 BCG consultants were given AI access, they completed 12.2% more tasks, worked 25.1% faster, and produced work rated more than 40% higher in quality. Then the researchers gave them a task that sat outside the model’s capability boundary, and the picture inverted. Consultants using AI were 19 percentage points less likely to reach the correct answer than consultants working without it.
Here is the part that should stop anyone selling a prompt-engineering webinar. The study ran three arms on that outside-the-boundary task. Consultants with no AI got it right 84.5% of the time. Consultants with GPT-4 got it right 70% of the time. Consultants with GPT-4 plus a prompt engineering overview got it right 60% of the time. The 19 points above is the pooled figure across both AI groups. The arms split it 14.5 and 24.5. Which means the group that received training performed worst of the three, and by a distance. A short overview did not close the gap. It widened it.
Without real strategic context, users cannot tell when AI is helping and when it is quietly dismantling their work. They lack the frame to evaluate the advice, so they accept it. “Kill 50% of the text” sounds decisive. “Think like Porsche” sounds aspirational. Both sound better than the uncertain feeling of staring at a draft and wondering if it is good enough.
"But you use AI too"
This is the objection I can hear forming. If AI built the work in the first place, why can’t AI improve it?
It is a fair question with a precise answer. The AI that built it is not the same intelligence as the AI that is “improving” it. Same technology. Completely different system.
When we build brand strategy, the AI operates inside a project environment loaded with that business’s positioning, past decisions, audience research, competitive context and specific goals. It is not a blank chat window. It is a system that knows why the consultant’s photo is large, why the copy addresses a specific anxiety, and why the layout follows a particular hierarchy. Ask it to improve the flyer and it will suggest tightening a headline or adjusting the visual weight of the call to action, inside the strategic frame.
The generic chat window knows none of this. It sees a flyer. It reaches for the average of all flyers. It produces advice that would be correct for no specific business and is therefore useful to none.
We watched this distinction play out in reverse with another business. The owner had built a solid referral programme concept in ChatGPT. Two reward tiers, database-wide distribution, basic tracking. The instincts were good. But a generic model could not see the gaps that the business context made obvious.
Run through a contextual process that understood the business model, the finance product structure and the clawback risk, that concept transformed. The two reward tiers became an escalating structure for repeat referrers. The “send it to everyone” approach split into differentiated handling, where settled clients get reward vouchers and enquiry-only contacts get a softer brand-awareness touchpoint. A separate acknowledgement loop fires the moment a referral arrives, not when the reward triggers. And a referability audit framework sits underneath everything, asking the question the generic model never thought to raise. Is the service actually worth referring right now?
The gap between “solid concept” and “implemented system” was not incremental. It was categorical.
Same technology at both ends. Context made one useful and the other dangerous.
The convergence tax
The damage from context collapse does not stop at individual businesses making their own work worse. It scales. When every business runs every piece of work through the same models with the same lack of context, the outputs converge on the same average.
This is now measurable. “Artificial Hivemind,” which won a Best Paper award at the NeurIPS 2025 Datasets and Benchmarks track, tested 25 models and found that competing systems built by entirely different companies produce responses with 71% to 82% semantic similarity to each other. At the top of that range sit DeepSeek-V3 and Qwen-Max at 0.82, and DeepSeek-V3 and GPT-4o at 0.81. Different labs, different countries, different training pipelines, near-identical answers. They are not just giving similar advice. They are giving the same advice, phrased slightly differently.
There is a natural experiment on the other side of this. When Italy temporarily banned ChatGPT in 2023, researchers at London Business School tracked the Instagram posts of Milan restaurants through the ban. Lexical similarity between businesses fell 15%. Syntactic similarity fell 12%. Average like counts rose 3.5%. People responded better to marketing that did not sound like everyone else’s. Two caveats belong with that finding. The same study found restaurants posted less often and wrote shorter posts during the ban, so the ban was not a free win. And the paper is an unrefereed working paper, not peer-reviewed research. Take it as a suggestive signal, not a settled fact.
Mark Ritson put it best. “When they invent a zigging machine, the value of a zag goes into the stratosphere.”
The mechanism is mechanical. Reinforcement learning from human feedback, the training process that makes AI outputs feel safe and polished, narrows the range of what a model will say. Think of it as a low-pass filter on creativity. The researchers have a blunter name for it: mode collapse. Kirk and colleagues, publishing at ICLR 2024, found that RLHF significantly reduces output diversity compared with supervised fine-tuning alone.
More recent work sharpens the diagnosis. A paper called Verbalized Sampling traces the flattening not to the algorithm but to typicality bias in the preference data. Human annotators reliably prefer text that feels familiar. The models are not trying to make you average. They were taught that average is what people like. Without specific context pushing them away from the centre, the centre is exactly where they will land.
Invisible to the machines that now decide who gets found
Here is where the cost compounds into something existential.
The discovery layer of the internet is shifting from ranked lists to synthesised answers. When someone asks an AI assistant to recommend a service, the AI does not show ten blue links. It names specific businesses. And the research on what gets a brand named points the same direction every time. Specificity wins.
Princeton’s Generative Engine Optimization study tested 10,000 queries across nine content methods. Adding direct quotations from credible sources lifted visibility by roughly 43%. Citing authoritative sources lifted it by roughly 30%. Adding quantifiable statistics lifted it by roughly 34%. The paper’s headline figure is a 41% improvement from its best-performing methods.
Two of its findings matter more than the winners. Keyword stuffing, the oldest trick in the SEO drawer, went backwards by roughly 8 to 10%. And writing in an authoritative tone, which is what generic AI-assisted marketing copy produces by default, was among the weakest methods tested at around 10%. Sounding authoritative does almost nothing. Citing authority does three times more. If your positioning language reads “comprehensive solutions” and “industry-leading platform,” you have optimised for the tactic that barely registers.
Seer Interactive analysed 541,213 AI responses across 20 brands and found a five-fold gap that sits inside the same brands, not between them. When a brand is named in the body of an AI’s answer, its own site is cited as a source 53.1% of the time. When the brand is not named in the answer, its citation rate falls to 10.6%. Being named is what earns the citation, not the other way around.
Seer’s read on why is that the model decides which brands to name from what it already knows, then goes looking for sources that support what it has already written. The citations are the bibliography, not the brainstorm. Seer is careful about this, and so should we be. They call it strongly supported behavioural evidence rather than proven architecture, because they do not have token-generation logs. It is the best available explanation, not a confirmed one.
Either way, the practical consequence is the same. A business producing generic, context-free marketing copy is building a brand that AI systems have no reason to remember and no evidence to cite. It is not just becoming average. It is becoming invisible to the systems that increasingly decide who gets found.
The gap that won't close on its own
If context collapse were evenly distributed, with everyone getting a little worse together, it might not matter. But it is not. A small group of sophisticated users is pulling away, and the distance is growing.
OpenAI’s data from more than a million business customers shows that workers at the 95th percentile of usage send six times more messages than the median worker in the dataset. For coding, the gap stretches to 17 times. For data analysts using the data analysis tool, 16 times.
Anthropic’s research on learning curves found that users with six months of experience succeed at tasks 73.1% of the time, against 66.7% for newer users. That is a raw gap of 6.4 percentage points, or about 10% in relative terms. Be careful with that number, because it shrinks under scrutiny. Once you control for the type of task being attempted, the gap narrows to roughly 3 percentage points, and to about 4 points with the full set of controls. The effect is real and it is smaller than the headline suggests. Anthropic’s own framing is the part worth holding onto. The report warns that “the benefits from early adoption may be self-reinforcing.”
The distribution is starker than the learning curve. Google and Ipsos found that 5% of American workers are genuinely fluent with AI, reshaping how their work gets done. Another 35% are ad-hoc explorers, dipping in and out. The remaining 60% do not use AI at work at all. That last number is the one most people get wrong. The majority are not dabbling badly. They are not in the room.
The businesses whose leaders develop real contextual AI skill will compound that advantage quarter after quarter. The ones whose leaders keep pasting strategy into blank chat windows will compound something else entirely.
The flyer, one more time
That automotive consignment flyer is still in circulation. The consultant’s photo is still prominent. The copy still explains what the service does, who it is for, and why this particular person is the one to talk to. It works. Not because it looks like Porsche, but because it looks like a business run by someone you could call on a Tuesday afternoon.
The AI that told them to strip all of that out and assert rather than explain was not malicious. It was average. A statistical composite of every brand playbook ever written, applied to a business it had never met.
In an age when the machines are making everyone sound the same, context is the only asset that appreciates. The question for every business owner is not whether they are using AI. It is whether the AI they are using knows anything worth knowing.
Your marketing, looked at properly
Thirty minutes on your current setup — what’s working, what’s quietly leaking budget, and what I’d fix first. You’ll leave with a clearer picture whether we work together or not.
Got something specific bugging you? Flag it when you book and I’ll have it looked at before we talk.
