AI's functional capability is created by marketing
AI is only as good as its people (and the people who market to them)
AI is only as good as the people who use it. Labs can spend billions developing new models but they’re functionally useless if no one uses them (see all the old, obsolete models they’ve trained).
Labs also can’t can’t know the complete functional capability of models. People use them in ways they didn’t expect and won’t understand. This doesn’t mean most people will do this though. Most people will use models in ways they know that work.
When a new model comes out, people are happy to just be told “it’s better.” The average person doesn’t experiment or evaluate new models for new capabilities. Basically no one changes from the default models. At max, they are trying 2-3 models at a time, but there are always more models, effort levels, context engineering, and workflows they could be trying.
Instead, people rely on:
Vibes
What other people are saying (AKA marketing)
Actual model capability is formed from what the labs are doing, but adoption significantly lags the frontier. This means functional capability is largely a product of marketing.
Fable: A case study
You could not ask for a better piece of marketing than the government banning your product because it’s too powerful. Mythos felt like the most hyped model ever, but did Fable functionally live up to its expectations?
Fable as planner
The functional capability of Fable was more shaped by its constraints than its potential. Specifically, usage was guided by the fact that you could only use 50% of your Claude plan on Fable.
Because usage was constrained and the API was expensive, people focused more on making the most of that usage limit than experimenting with the limits of capabilities. It created a scarcity mindset.
It seemed (at least in my feed) that this meant using Fable as a planner and have other models execute. Dozens of popular tweets like this one captured my attention:
Fable was probably the most capable model at planning, but this doesn’t mean there weren’t other capabilities to explore. Planning makes sense given the constraints. It was even further encouraged by Anthropic through the pattern of using Fable as an advisor.
In an alternative universe where Anthropic didn’t put the 50% constraint on Fable, we wouldn’t have seen such as focus on Fable as a planner. Even if API costs were expense, people would have treated it like any other model. New models of the past were also better planners, but never didn’t have the constraint limiting their functional potential. Fable as planner became an echo chamber that drowned out other potential capabilities.
Pet problems
What does the opposite of letting marketing guide capabilities look like? An example is pet problems. Things that past models couldn’t solve that current ones can, accomplished at the individual level rather than benchmarks.
Why these? Because reality has a surprising amount of detail. Every pet problem a new model solves is a new capability.
I never saw anything about what Simon Willison calls Fable’s “relentless productivity.” This is how it goes to great lengths to accomplish tasks: building its own debugging tools, using every bit of software available to it, trying non-obvious angles when obvious ones fail. For me, this capacity seems like a newer, bigger deal, but it’s only really discovered through someone’s individual usage and reflection.
The problem is that most people don’t have pet problems to solve. They don’t keep track of what current models can’t solve and don’t even think about problems they might want models to solve. They are mostly content with their existing workflows and this limits the functional capabilities of AI.
What kind of marketing is most important to AI capabilities?
Marketing means a lot of things, but there are three areas that are most specifically relevant to the creating of AI capabilities:
Don’t underestimate the impact of creators
If you in the camp like me who thinks AI is good and want functional capabilities to expand, don’t underestimate the impact creators can have.
Creating content, reviewing models, and sharing examples pushes the frontier a little further. Diversity of use cases matters too. Software development gets a lot of focus, while an area like writing gets very little.
Although there is a lot of bad creators, ones like Every and Theo are doing a great service with their model reviews and opinions. They deserve the attention they are getting and more.
Thousands of people are evaluating the model and building their mental models of what they do because of people and companies like these. Their reviews inspire people to apply new use cases to their work and try models in new ways, expanding functional capabilities.
Model announcement as capability creation
The model release announcement is one of the best ways labs can actually expand the functional capabilities of models. This is because they get a lot of attention and people expect them to talk about capabilities.
People mostly want to know this model is better than the last one. This is done in two ways.
Show that this is the best model that you should be using (graphics, benchmarks)
Telling you how you should be using it
For example, in the Opus 5 announcement you get a paragraph like this:
An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly.
The key phrase is “previous models could not complete this task.”
A trap a lot of model announcements make is spending too much time on benchmarks and graphs rather than examples. There is a balance between convincing people to use your model and expanding functional capabilities. Too often (like the Opus 5 announcement), it falls far on the convincing side, focusing too much on benchmarks and graphs. In the 2,821 word post, only 244 words from Anthropic and 410 words from customers were dedicated to new capabilities.
Benchmarks are good marketing, but less good at creating capabilities
Benchmarks are maybe the most important piece of model marketing. They provide a more robust and believable proof that models are getting better which is basically what people want to know. Benchmarks have even created their own brands in support of this, you think the name “humanity’s last exam” is anything other than marketing?
The other major benefit of benchmarks, at least for their creators, is that labs optimize towards them. For example, Zapier has AutomationBench which Opus 5 “topped ... without spending more tokens than prior Claude models.” You can infer that Anthropic used AutomationBench in training, which provides the most benefit to Zapier.
This is also the downside of benchmarks, models can be overfit to them by optimizing directly on them. For example, the recent Grok 4.5 model “unintentionally” included a snapshot of the Cursor codebase in training, which was likely a factor in its good performance there.
And although benchmarks show that models are getting better, it’s unclear how that translates to functional capabilities. I’ve never looked at the details of benchmarks and I doubt the average AI user has either. An improvement to a benchmark does not mean an equivalent improvement to functional capabilities.
As useful as marketing is, it’s still marketing
Marketing is meant for the masses, but how you use a model should be unique to you. Model capabilities are spiky, it might be good at the benchmarks but bad at the actual capabilities you care about. The quirks of how Theo uses models are not the same as you.
This means you should have your own way of evaluating models like pet problems that models can’t solve yet or a personal benchmark, like PelicanBench or the ability to write a banger tweet:
This is especially important if you aren’t a software developer. A lot of marketing, benchmarks, optimization is done for software developers. They do a lot of experimenting and have clear metrics to show if they succeeded. It’s not so easy to evaluate capabilities in other roles.
And if you do this, share it! Reflecting on how you use LLMs and sharing that with others is probably the best way the average person can increase functional capabilities of LLMs. It has the potential to help others more than you realize.








