Why AI Won't Kill Agencies: The Tech Barrier Nobody's Talking About
Why AI Won't Kill Agencies: The Tech Barrier Nobody's Talking About

I keep seeing posts about how AI is going to kill agencies.
How agents are going to automate marketing.
How strategists, creatives, and consultants will all be replaced by a few spicy prompts.
I've spent the last few years properly researching and testing this - to understand the actual technical and economic constraints of how these models work.
My conclusion? The "AI replaces everything" narrative fundamentally misunderstands what these tools can and can't do.
Some Jargon Before We Get Started
Before we dive in, a quick explainer on some of the terminology:
LLM - Large Language Model. This is the technology behind tools like ChatGPT, Claude, and Gemini. It's the AI that reads and generates text.
Tokens - The language LLMs speak. They break everything down into small chunks called tokens - roughly ¾ of a word on average. When we talk about how much an LLM can process, we measure it in tokens.
Context window - The LLM's short-term memory. It's the maximum amount of tokens a model can hold and work with at any one time. Go beyond it and the AI either forgets, compresses, or just says 'nah'.
Right. Now we're ready.
The Brute Force Fantasy
The narrative I keep hearing goes something like this: just plug AI into everything, let it have access to all your data, and you're sorted.
I know people who have Copilot are particularly guilty of thinking this - and I don't blame them. Microsoft make big promises about integrations, and it does integrate, and it is useful.
But it also has severe limitations.
In this blog I'll walk you through how the technology actually works and why full integration is still such a pipe dream.
It falls apart in four ways.
Problem 1 - Cost: "You Can't Afford Me, Honey"
Let's start with the maths.
A typical enterprise business sits on around 400 terabytes of structured data. Customer records, communications, documents, transactions - the accumulated knowledge of the organisation.
And that's just the structured stuff. Include unstructured data - emails, PDFs, Slack threads, the works - and you can multiply by 10. Dark funnel data? Don't even start.
400TB of structured data is still more than enough for me to make my point. This equates to roughly 100 trillion tokens. At current API pricing (November 2025), processing that would cost:

Roughly $250 million at the mid-point. For a single query.
And if you're doing complex strategic work - the kind that actually requires advanced reasoning - you'd be using Claude Opus, which pushes you north of $500 million.
Even a modest 4TB subset - roughly 1% of enterprise data - would cost $2-4 million per query depending on the model.
You cannot pay to read everything. You have to be selective with what you feed it.
Problem 2 - Scale: "They're Called "Data Lakes" For A Reason"
Let's pretend for a moment that money was no object.
You could afford to have AI read everything in your system all at once. Every email, every call note, every deal record. And it gives you good, evidence-based answers to every question you have.
Oh if only it was that easy, my friend!
There's a limit to how much AI can meaningfully process before it starts tripping over itself. Before it starts hallucinating. Before everything it says becomes so generic it's useless.
Here's an analogy to get your head around the scale of this issue.
Around 35K tokens is roughly where models do their best, most reliable reasoning work. Think of that as a single Rubik's cube worth of data.
Now, 400TB of enterprise data? That's 2.86 billion Rubik's cubesubik's cubes. Enough to fill Big Ben 38 times over.
Include all your unstructured data - you're looking at 383 Big Bens. A whole skyline.

And you? You've got a single Rubik's cube worth of analytical capacity before things start degrading. Go to the max limit of 1 million tokens for Gemini 3 (~28 Rubik's cubes) and the model just stops - dead on arrival.
This is a needle in a haystack situation... on steroids with a double espresso chaser.
Yep, your boy gone done the maths on this.
Finding a needle in a haystack is a 1 in 56 million problem (you can ask me how I got to this over a pint, but it's logical - promise).
Searching your enterprise data with an LLM? That's a 1 in 2.86 billion problem.
That's not a needle in a haystack. That's a shattered needle scattered across 50 haystacks.
Even if you could afford the AI processing everything, the AI stops being useful after analysing 0.000000035% of your data (if you're analysing the 400tb).
Why does this happen?
It's architectural - a fundamental constraint of how these models work.
The attention mechanism - how models focus on what matters - gets diluted. More tokens means the model's "attention budget" is spread thinner across more information.
There's a thing in LLMs called the "Lost in the Middle" phenomenon - a U-shaped performance curve. Models are good at recalling information at the beginning of the context (primacy bias) and the end (recency bias), but significantly worse at retrieving information buried in the middle.
If your critical data point is in the middle of a 100K token input, the model is statistically more likely to miss it.
There's an actual AI test called "needle in a haystack" which shows models can find a single fact in massive contexts with high accuracy. Gemini 1.5 Pro scores 99.7% recall up to 1 million tokens for simple retrieval.
But that's misleading.
Complex reasoning - connecting multiple facts separated by thousands of tokens - fails much faster. The RULER benchmark shows that while simple retrieval stays strong, "Variable Tracking" and "Aggregation" tasks degrade significantly as context grows.
Problem 3 - Prioritisation - "It Can Read, But It Can't Think"
So, we've established you can't afford to brute-force your data. And even if you could, the technology trips over itself at scale.
But let's indulge the fantasy again.
You've got unlimited budget. The AI can process your entire database without hallucinating or turning everything into beige mush.
There's still a problem: it doesn't know what matters.
Imagine an SDR on a Monday morning asking: "Who should I prioritise this week?"
Simple question. But to answer it properly, the AI would need to process your CRM data, engagement history, ICP criteria, deal stage progression, historical win/loss patterns, recent email opens, website visits, intent signals - the list goes on.
Let's say it does that. It pulls everything that could possibly be related to the question.
Now it's sitting on a mountain of data. But it doesn't have a Scooby which bits are actually meaningful.
It doesn't know that your Head of Sales ignores any account under £50K ACV. It doesn't know that the sales team added pipeline to a few prospects just to boost their numbers. It doesn't know that most of your CRM is out of date and that the Ops team are too scared to delete old data. It doesn't know that "engaged" in your world means something different than the textbook definition.
So what happens is it gives you a generic list that looks plausible but misses the point entirely.
Cost says you can't afford it. Scale says it'll break before it finishes. And even if neither of those were true, it still can't tell the difference between signal and noise.
Problem 4 - DIY: "Fine! I'll do it myself"
I hear this a lot. "Fine-tune a model on your data and it knows what to do."
Not quite.
Let's start with the nuclear option: training a model from scratch.
GPT-5 reportedly cost over $1 billion to train. Gemini 3 is estimated at $500-800 million. Claude Opus 4.5 somewhere in the $400-600 million range.
Just to reiterate - that's just on compute...
Unless you're a sovereign nation or Big Tech, that's off the table.
So what about fine-tuning? That's more accessible - typically $100k-500k per model version. But there's a doozy of a catch: fine-tuning doesn't teach knowledge. It teaches behaviour.
The research confirms it - fine-tuning is good at making models sound like your brand. It's rubbish at making them remember your actual data (Ovadia et al., 2024).
Fine-tuning is like sending an employee to a communication seminar. It doesn't give them access to your file server.
And then there's the retraining problem.
Your data isn't static. Products change. Customers evolve. Competitors shift. That fine-tuned model? It's frozen in time from the moment training finishes.
To keep it current, you'd need to retrain constantly - potentially daily if you want it to reflect recent emails, deals, or conversations. That's millions per year in compute, plus the versioning nightmare of managing multiple model states.
So you can't afford it. It breaks at scale. It doesn't know what matters. And fine-tuning just teaches it to sound like you - not think like you.
Four ways in. Four dead ends.
I think I've made my point. There is no brute force shortcut here.
Why This Matters for the "Agencies Are Dead" Argument
This is where the technical constraints connect to the bigger picture.
AI is incredibly powerful. Us marketers are already doing great things with it - testing new products as they come out, finding where it fits, pushing what's possible.
But don't let that distract you from a fundamental truth: the good outputs coming out of AI aren't exclusively because of the AI. They're because of the people using it.
(And let's be honest - there's a lot of bad stuff coming out of AI too. But that doesn't mean you can't do good with it.)
People who know what to ask. People with the experience and the judgement. People who've built up a fountain of knowledge over years that they're now channelling through a new tool.
The AI can't curate context for itself. It can't decide what's relevant without understanding:
- Your specific objective
- Your ICP nuance
- Your internal politics
- Your brand tone
- What actually moves pipeline
- What the Head of Sales will immediately veto
- What's legacy/inaccurate data
That requires judgement. It requires understanding what matters and what doesn't for a particular situation. It requires strategy.
The stuff that makes marketing actually work.
If marketing could be solved by dumping everything into an AI and letting it figure things out, someone would have built that system by now. We'd have one agency. One approach. One standardised solution.
We don't.
Because context isn't just data. It's knowing which data matters, in what combination, for what purpose.
And that's human judgement. My way of doing things is different to the next agency's. And theirs is different to the one down the road. That's why there are thousands of us - each with our own approach, our own perspective, our own way of solving problems.
That's kind of the beauty of marketing - it's always evolving. It needs innovation. Fresh ideas. Nuance.
The stuff that doesn't fit in a context window.
How The Tech Actually Works
If you've made it this far - well done, and thank you. Genuinely.
I'm now going to get into the weeds a bit on how the tech behind the context window actually works, and some practical implications for how to use it better. Feel free to skip to the end if you just want the takeaway.
The major AI platforms are all built in slightly different ways. Each has made different architectural choices to handle context limitations. None of them solve the fundamental problem - but understanding the differences helps you use them better.
ChatGPT (GPT-5):
Uses an adaptive window approach - a mix of sliding window with internal compression and priority zones for "important" messages.
When you hit the limit, it doesn't just drop everything. It reweights and compresses - low-priority turns get merged or de-emphasised, while high-priority instructions and facts stick around longer. But the exact wording? That can still get lost.
There's a separate "Memory" feature for persistent facts ("My name is Jake"), but this doesn't preserve conversation flow or context.
Google Gemini 3:
Takes the "massive vault" approach. With a 1M+ token context window, it simply holds more raw information. No summarisation, no compression - it relies on sheer capacity to delay the problem.
When you finally hit that limit (which, to be fair, is rare in normal usage), the system stops or requires a new session rather than gracefully degrading.
Microsoft Copilot:
Best at integrating within your environment - it pulls directly from your files, emails, and workspace. But it's the worst at actually remembering what you talked about.
Give it a few chats and it's already lost the thread. Great for environment access, terrible for continuity.
Claude 4.5:
Takes a different approach. When conversations approach the context limit (around 95% capacity), it automatically compacts older messages into a summary. You'll see a notification: "Compacting our conversation..."
This trades perfect recall of early details for the ability to continue longer conversations. The meaning is preserved even when verbatim text isn't.
The comparison:

I'm not picking favourites here - they're all genuinely good for different things because of the way they've been built.
Claude for strategy and complex reasoning - it mirrors how humans actually think, remembering the key points rather than every word.
Gemini for large-scale research, app prototypes, and analytics - that massive context window is perfect when you need to process volume.
Perplexity for verifying sources and fact-checking - purpose-built for research with citations.
GPT for impulses and administration - it feels like it "gets me" the most which is great for quick tasks and day-to-day workflow.
Copilot for ecosystem integration - instant insights from what's open on your desktop, and occasionally pulls a rabbit out of the hat by surfacing relevant stuff from your wider network. But it's searching by semantic meaning, which is unreliable. And if you want to do something useful with what it finds? You're moving to another model.
The one thing they all share is that none of them can automatically determine what context actually matters in the first place.
The curation problem - deciding what information is relevant to your specific objective - remains unsolved by every platform.
Practical Implications
If you're using AI in your work, here's what actually helps:
Start fresh conversations when switching topics. Accumulated context dilutes the model's focus. A clean slate for a new problem often produces better results.
Curate what you include. Five relevant documents beats thirty comprehensive ones. More isn't better - more targeted is better.
Put critical information at the start or end. Remember the "lost in the middle" problem. Don't bury your most important context in the middle of a long prompt.
Be intentional, not exhaustive. Think about what the model actually needs to know for this specific task, not everything that might possibly be relevant.
Accept that no single tool does everything. The harsh reality is there isn't one model that handles strategy, research, fact-checking, and administration equally well. Use the right tool for the job - and know when to move between them.
To Finish Up
We're a long way from AI superintelligence. The models are genuinely impressive - I use them daily and they've transformed how I work. Honestly, I couldn't do my job without them now. It's so baked into my workflow.
But the brute force fantasy - plug it in, let it read everything, job done - doesn't hold up.
It fails on cost. It fails on scale. It fails on prioritisation. And training your own model doesn't fix any of it.
The tools are still brilliant. But a fighter jet without a pilot is just a very expensive metal box. PowerPoint without someone to build the slides is just a blank screen.
The value isn't in the tool. Everyone has access to the same models, the same APIs, the same capabilities.
The value is in knowing what context to give it. And that requires judgement, strategy, and an understanding of what actually matters for your business.
That's not getting automated any time soon.
Oh, one more thing.
There is something that's a really big step in the right direction. It's called RAG - Retrieval Augmented Generation.
It's not going to replace people. You still need to do the strategy. But it can help you store data in a way the AI actually understands, and scale your efforts.
Getting the right context, to the right model, at the right time.
But that's one for a separate blog.
Sorry to leave you all on tenterhooks.
xoxo






