Back in February I wrote, almost in passing, that I was building AI agents around my product marketing Centre of Excellence. I had a clear vision, but I had not yet worked out exactly how to build it. Six months later, I put everything in practice and built a system of reusable Claude skills, reference material and workflows to fit with the deliverables at hand. Some of it is working quietly in the background, some of it has failed, and the polished, “everything-worked” version is probably on an iteration level v3 or 4 and is about to be updated.
If I had to compress the last six months into a single lesson, it would be this: AI did the easy twenty percent and left me the hard eighty. The easy twenty is the drafting, the summarising, the first pass nobody was ever precious about anyway, while the hard eighty is the work of deciding what the company actually means in the market, which buyers matter, what we will not say, and whether the internal story is even agreed enough to be worth automating at all. The skills carry my decisions into the work. They do not make those decisions for me. Every time I forgot that distinction, the system bit me; every time I respected it, the system saved me hours. That, more or less, is the lesson. Everything after this is detail, though the detail is where it gets interesting, because some of what I built worked quietly in the background and some of it fell apart in my hands.
Why the blueprint had to come before the agents
There is a question I keep getting asked, and it usually arrives as a challenge rather than a real question: everyone uses AI now, so why would anyone bother with a blueprint?
I answered that back in February, and it is worth repeating now because the last six months have proved it the hard way. When the foundations are unclear, when ownership is blurry, when the story shifts depending on who happens to be presenting it or who has the biggest voice in the room, AI does not fix any of that. It accelerates it, so you end up with the same confusion spreading faster, into more places, with more misplaced confidence stapled to it. A skill built on a shaky foundation is simply an efficient way to be wrong at scale.
That is exactly why the blueprint had to come first, and not out of any love for documentation. The blueprint is the thing that decides the story once so that everyone else can run with it, which is narrative governance rather than narrative creation: you make the call, you defend it with evidence, and you resist the temptation to reopen it every other week. The Centre of Excellence captured what good looked like so it would hold when the surrounding work started moving. The skills came second, entirely on purpose. You can automate a system but you cannot automate a guess.
Building Claude Skills
By “skills”, I mean reusable Claude instructions, reference material and workflows and not a fleet of autonomous agents. Building them does not come down to writing clever prompts. I find that a clever prompt is a party trick: it works once, for the person who wrote it, on a good day, and it does not survive first contact with a scaling business. What I built is the blueprint turned into runnable internal policy, where each of my more contrarian beliefs about this function now has an artifact that carries it into the work. Because product marketing is the value clarification department rather than the launch department, the system starts from the problem instead of the feature.Because I treat positioning and qualification as two of the first hypotheses to test in a lost deal, I tried to wire competitive intelligence in as a live input. And because positioning that sales do not repeat in their calls is only content, the ICP and the proof live in one governed place that everything else reads from. Because an unedited AI draft rarely respects the reader’s time, every skill produces a starting point and then hands it back to me to finish.
Holding the belief is the senior part of the job and building the skill is the execution part, and I wanted to show both, because a point of view you cannot operationalise is only an opinion, while execution without a point of view is only activity.
Under the hood
I want to be precise here, because “I built a system of AI agents” is the kind of sentence that invites people to imagine something far more finished than the truth. The honest status is three tiers rather than one.
Some of the skills genuinely run day to day. The ICP and positioning skills are in more or less constant use across marketing and have now been made available to sales as a shared source of truth, the copywriting skill and the messaging builder come out whenever an asset or a value chain is needed, and the executive brief runs on its own separate track for leadership reporting, pulling performance data rather than positioning. Those are the parts of the system I would comfortably call “live”. Some skills, by contrast, have run exactly once: the launch tiering, launch brief, feature comms, and the VFD framework audit (based on the beloved Emma Stratton) were all built and used for the last product launch, simply because that was the first time I had a real launch to give them structure around. One launch is not a proven cadence, though; it is a single trial with a lot of lessons attached, and I will come to those shortly. And some skills only run ad hoc, when something triggers them, so the ICP research refresh and the competitor analysis fire when someone tells me a piece of information is wrong or out of date, and when they do, they feed the corrections back into the Confluence records that hold the source of truth. They do not yet run on a consistent review schedule. For now, they are triggered when someone flags that information may be outdated, which is useful but still short of proper governance.
| Skill | Status | What it does |
| positioning, icp | Running daily | Source of truth. Claims, proof, buyers, behavioural segments. Everything reads from here first. |
| copywriting, messaging | Running | Drafts market-facing assets and value chains off the source of truth. |
| executive brief | Running, separate track | Leadership reporting from performance data. Also used as part of the win-loss analysis. |
| launch tiering, launch brief, feature comms, vfd framework | Used once in the last launch | Sized, planned, aligned, and audited a single launch. One trial, needs updating and adapting for better results. |
| icp-research-refresh, competitor-analysis | Ad hoc, when flagged | Firing when someone reports the information is wrong and feeds corrections back into Confluence. |
It is worth pausing on the pattern in that table, because it is probably the most useful thing in this whole post. The skills that run reliably are the ones sitting on a stable, agreed foundation, and the skills that ran once and then struggled were the ones sitting on top of a decision the business had not actually made yet. Skills work exactly as well as the agreement underneath them. Hold that thought.
What did I change?
The real result is behavioural change. The skills have helped colleagues work faster and changed how they use product marketing support.
Before all this, the marketing team came to me for raw information: what is the positioning, what is the ICP, how do we describe this, what do we say to this particular partner. I was like a vending machine, and every piece of event positioning and every partner email routed through me for the simple reason that I was the only place the answer actually lived. Now they do the first pass themselves and come to me for confirmation and review instead. The event positioning gets drafted against the right skills in a dedicated events positioning artifact i created, and the partner email gets written straight off the source of truth, so by the time it lands on my desk the question becomes “is this right?” I have stopped overthinking every small piece of copy, the team has stopped waiting on me before they can start, and that shift, from being the author of everything to being the reviewer of everything, is the one story engine doing precisely what it was designed to do. Truth now lives in fewer places, so people stop guessing which version they are supposed to follow.
The moment that made me laugh came during the same launch: a piece I had written entirely by hand was labelled “AI slop”, while an AI-assisted section passed without comment. This proved that instinct is a poor way to judge AI use, and that quality, accuracy and usefulness are better tests. Either way the lesson holds: the tool was never the useful test. The judgement was. That is why I built a voice profile and banned-phrase filter into the system. The useful questions are whether the work is good, whether it is true, and whether it serves the reader.
What did I break?
So… yeah, about that..
The launch is where the skills produced the most garbage, for reasons that are both worth spelling out. The first challenge was that the product continued to evolve during the launch. Some feature details changed after I had already built the skill around an earlier version, so I was effectively automating a moving target. Each change meant revisiting both the feature communications and the underlying instructions. The second reason runs deeper: the terminology and supporting information were still being aligned across the business. The same capability was being described in different ways across the organisation, which meant there was not yet a single, stable story to encode. After several revisions, I paused the skill, completed the critical work manually, and worked with the wider team to close the remaining gaps.
So the lesson is to not put a freshly built skill on the critical path of a live launch with a deadline breathing down your neck and no time to test it, because a skill needs a controlled environment and enough slack to fix and tweak before you can trust it with real work under pressure. I broke that rule because the launch timeline was fixed. The result was additional manual work and a skill I had to pause mid-flight.
It is worth looking at why it broke, because the problem was not simply how the skills had been built. It broke because the underlying story was still being agreed, and automation cannot replace alignment that has not yet happened. The launch chain exposed the exact risk the blueprint was designed to reduce: automating before the underlying decisions are stable.
That is also why I have to correct something I may have implied both earlier in this post and back in February. These skills are a strong default rather than a control, in that they make the right path the easy path and they make skipping the thinking visible, but they do not and cannot force anyone to follow them, myself very much included. Under enough pressure, I overrode my own system. That made the limitation clear: these skills can support governance and make deviations visible, but they cannot enforce it on their own.
What I cannot prove yet
If I am going to hold myself to my own standard, I have to be also honest about the biggest gap.
I cannot honestly tell you that sales are repeating the positioning in every call, or that the ICP is keeping deals on track out in the field. I would love to, since that is the reason I built the ICP skill in the first place, but I do not have the feedback loop that would prove any of it. I do not yet have reliable visibility into who uses the ICP, when they use it, or how it affects live opportunities. The sales-to-marketing feedback loop is still too informal to prove adoption. The tooling and process needed to measure that feedback loop are not yet in place. By my own rule that means I cannot claim it, because if sales adoption is not something I can point to with evidence, then any statement I make about it is content rather than proof, and content does not count. So I am not going to dress a hope up as an outcome. The honest position is that the marketing-facing half has visibly changed how the team works, while the sales-facing half remains unmeasured and needs a more structured adoption signal. Building that feedback loop, some structured signal of whether the ICP and the proof tiers are actually showing up in real deals, is the next thing on the list.
The guardrails
This system is built so that the direction comes from a person of knowledge. AI here is the starter and the accelerator; it surfaces patterns, processes far more than I ever could alone, and hands me a strong first draft, and a machine-made first draft is perfectly fine. Shipping it unedited is not, so every output is reviewed and rewritten before it reaches an audience. Underneath it all, the voice profile and banned-phrase list remove much of the predictable AI cadence and puffery before an output reaches anyone who might put their name to it. The AI-slop incident reinforced the broader lesson: every draft needs to be judged and edited on its own merits regardless of how it was produced.
There is also a “no” built into the system, and it is the guardrail people skip most often. Not every feature deserves to become a headline, and some capabilities matter as proof or context rather than as the main message, which is exactly why the launch tiering skill exists: so that “this deserves a full campaign” becomes a decision with evidence behind it rather than a reflex from whoever built the feature and would like to see it celebrated. When the tiering process and supporting evidence point to “Tier 3, no brief”, that is the function doing its job. The friction job. And that might be a whole discussion for another day.
Where this leaves me
The Centre of Excellence started as eighteen pages in Word, written over three days, and it felt like finally writing down what should have been obvious all along. The blueprint made the decisions, and the skills were meant to carry those decisions into every task, every day, so that consistency would not depend on whether anyone felt like following a process on a given Tuesday.
Six months in, the honest scorecard is mixed, and I am fine saying so. The steady work is genuinely better and faster, and the team has changed how it operates around the system. That behavioural shift is the result I care about most. The launch chain proved that I can build the machinery, but also that it cannot perform reliably while the story and product details are still evolving. The sales adoption I built all of this for is still unmeasured, and that one is on me to fix next.
That is the real state of it: not finished and certainly not magic, but working where the foundation is solid and breaking where it is not. That may be uncomfortable, but it is also the most useful lesson the system has produced.
I also want to thank the people whose work helped me get started and challenged my thinking along the way: Andrea Saez, Yi Lin Pein, Rory Woodbridge, Mary Sheehan, April Dunford and Emma Stratton, among many others. The more experienced product marketing minds I can learn from, the better my own work becomes. You may never read this, but thank you.