Back in February I wrote, almost in passing, that I was building AI agents around my product marketing Centre of Excellence. It was the kind of thing you say when you have promised yourself something you have a vision for but not fully worked out yet, and then over the months you make it happen. So I pulled together an honest account of how that went, and it needs to be honest, because the polished, everything-works version of this story would be an absolute untruthful thing to do.
If I had to compress the last six months into a single lesson, it would be this: AI did the easy twenty percent and left me the hard eighty. The easy twenty is the drafting, the summarising, the first pass nobody was ever precious about anyway, while the hard eighty is the work of deciding what the company actually means in the market, which buyers matter, what we will not say, and whether the internal story is even agreed enough to be worth automating at all. The skills I built carry my decisions into the work, but they do not make those decisions for me, and every time I forgot that distinction the system bit me, just as surely as it saved me hours every time I respected it. That, more or less, is the lesson. Everything after this is detail, though the detail is where it gets interesting, because some of what I built worked quietly in the background and some of it fell apart in my hands.
Why the blueprint had to come before the agents
There is a question I keep getting asked, and it usually arrives as a challenge rather than a real question: everyone uses AI now, so why would anyone bother with a blueprint?
I answered a version of it back in February, and it is worth repeating now because the last six months have proved it the hard way. When the foundations are unclear, when ownership is blurry, when the story shifts depending on who happens to be presenting it, AI does not fix any of that. It accelerates it, so you end up with the same confusion spreading faster, into more places, with more misplaced confidence stapled to it. A skill built on a shaky foundation is simply an efficient way to be wrong at scale.
That is exactly why the blueprint had to come first, and not out of any love for documentation. The blueprint is the thing that decides the story once so that everyone else can run with it, which is narrative governance rather than narrative creation: you make the call, you defend it with evidence, and you resist the temptation to reopen it every other week. The Centre of Excellence was me writing down what good looks like so that it would hold when everything around it started moving, and the agents came second entirely on purpose. You can automate a system. You cannot automate a guess.
Building Claude Skills
Building skills, in the sense I mean it, does not come down to writing clever prompts. I find that a clever prompt is a party trick: it works once, for the person who wrote it, on a good day, and it does not survive first contact with a scaling business. What I built is the blueprint turned into runnable internal policy, where each of my more contrarian beliefs about this function now has an artifact that carries it into the work. Because product marketing is the value clarification department rather than the launch department, the system starts from the problem instead of the feature. Because every lost deal in my eyes is a positioning failure before it becomes a product failure, i tried to build competitive intelligence wired in as a live input. And because positioning that sales do not repeat in their calls is only content, the ICP and the proof live in one governed place that everything else reads from. And because the unedited AI draft is contempt for the reader, every skill produces a starting draft and then hands it straight back to me to finish.
Holding the belief is the senior part of the job and building the skill is the execution part, and I wanted to show both, because a point of view you cannot operationalise is only an opinion, while execution without a point of view is only activity.
Under the hood
I want to be precise here, because “I built a system of AI agents” is the kind of sentence that invites people to imagine something far more finished than the truth. The honest status is three tiers rather than one.
Some of the skills genuinely run day to day. The ICP and positioning skills are in more or less constant use for marketing – and now sales, too – as the source of truth, the copywriting skill and the messaging builder come out whenever an asset or a value chain is needed, and the executive brief runs on its own separate track for leadership reporting, pulling performance data rather than positioning. Those are the parts of the system I would comfortably call “live”. Some skills, by contrast, have run exactly once: the launch tiering, launch brief, feature comms, and the VFD framework audit (based on the beloved Emma Stratton) were all built and used for the last product launch, simply because that was the first time I had a real launch to give them structure around. One launch is not a proven cadence, though; it is a single trial with a lot of lessons attached, and I will come to those shortly. And some skills only run ad hoc, when something triggers them, so the ICP research refresh and the competitor analysis fire when someone tells me a piece of information is wrong or out of date, and when they do, they feed the corrections back into the Confluence records that hold the source of truth. They are not on a clean schedule yet so much as on a “someone complained” schedule, which is better than nothing and some way short of governance.
| Skill | Status | What it does |
| positioning, icp | Running daily | Source of truth. Claims, proof, buyers, behavioural segments. Everything reads from here first. |
| copywriting, messaging | Running | Drafts market-facing assets and value chains off the source of truth. |
| executive brief | Running, separate track | Leadership reporting from performance data. Also used as part of the win-loss analysis. |
| launch tiering, launch brief, feature comms, vfd framework | Used once in the last launch | Sized, planned, aligned, and audited a single launch. One trial, needs updating and adapting for better results. |
| icp-research-refresh, competitor-analysis | Ad hoc, when flagged | Firing when someone reports the information is wrong and feeds corrections back into Confluence. |
It is worth pausing on the pattern in that table, because it is probably the most useful thing in this whole post. The skills that run reliably are the ones sitting on a stable, agreed foundation, and the skills that ran once and then struggled were the ones sitting on top of a decision the business had not actually made yet. Skills work exactly as well as the agreement underneath them. Hold that thought.
What did I change?
The real result is a change in behaviour, which by my own definition is the higher bar to clear, since good product marketing changes how people work. And I have definitely helped my colleagues to go faster in their productivity with the above skills.
Before all this, the marketing team came to me for raw information: what is the positioning, what is the ICP, how do we describe this, what do we say to this particular partner. I was like a vending machine, and every piece of event positioning and every partner email routed through me for the simple reason that I was the only place the answer actually lived. Now they do the first pass themselves and come to me for confirmation and review instead. The event positioning gets drafted against the right skills in a dedicated events positioning artifact i created, and the partner email gets written straight off the source of truth, so by the time it lands on my desk the question becomes “is this right?” I have stopped overthinking every small piece of copy, the team has stopped waiting on me before they can start, and that shift, from being the author of everything to being the reviewer of everything, is the one story engine doing precisely what it was designed to do. Truth now lives in fewer places, so people stop guessing which version they are supposed to follow.
The moment that made me laugh, and it is also rather the point, came during that same launch, when I got pulled up for “AI slop” on a piece I had written entirely by hand while the section that had genuinely been generated with AI sailed straight through untouched. If my own hand-written prose reads as machine output, I do not take that as a knock on my writing so much as a sign that the model has been trained hard on my voice, which is exactly what the voice profile in my system is built to do. Either way the lesson holds, because the tool was never the tell. The judgment is. It is why I built a voice profile and a banned-phrase filter into the system in the first place, so that the question is “is this any good, and is it true?”, which are the only two questions that matter, and neither of which is answered by who or what happened to type the first draft.
What did I break?
So… yeah, about that..
The launch is where the skills produced the most garbage, for reasons that are both worth spelling out. The first is that the product itself simply would not hold still, because even during code freeze there were adaptations of features, usually telling me only after I had already built the skill around the previous version, so I was effectively automating a moving target; every time I thought I had the feature comms locked down, the feature underneath it had already moved. The second reason runs deeper: the business had not yet settled internally on what to call things, and there was genuine turmoil around the information itself. The same capability was being described in more than one way across the organisation, so while all of that got worked through there was no single story stable enough to encode. I lost count of how many times I updated that skill before I eventually stopped altogether, did the important parts by hand, and let the wider team pitch in to close the gap.
So the lesson is to not put a freshly built skill on the critical path of a live launch with a deadline breathing down your neck and no time to test it, because a skill needs a controlled environment and enough slack to fix and tweak before you can trust it with real work under pressure. I broke that rule because the launch was happening whether I was ready or not, and I paid for it in manual hours, a skill I abandoned mid-flight, and everyone else being annoyed at me.
It is worth looking at why it broke, though, because it did not break on account of the skills being badly built. It broke because there was no agreed story underneath them, and no automation on earth can manufacture an agreement that the humans involved have not actually reached. The launch chain failed for the precise reason the blueprint exists in the first place, which was confirmed at a later date.
That is also why I have to correct something I may have implied both earlier in this post and back in February. These skills are a strong default rather than a control, in that they make the right path the easy path and they make skipping the thinking visible, but they do not and cannot force anyone to follow them, myself very much included. Under enough pressure I overrode my own system without much hesitation, so anyone telling you that their AI setup “enforces” governance has either never run it during a real launch or is trying to sell you something.
What I cannot prove yet
If I am going to hold myself to my own standard, I have to be also honest about the biggest gap.
I cannot honestly tell you that sales are repeating the positioning in every call, or that the ICP is keeping deals on track out in the field. I would love to, since that is the reason I built the ICP skill in the first place, but I do not have the feedback loop that would prove any of it. Nobody tells me who opens the ICP file, or when, or why, sales feedback to marketing is thin at the best of times, and I have simply not yet built the mechanism that would close that loop. I have asked for software that can help me close the loop but that was 12 months ago with no decision having been made. By my own rule that means I cannot claim it, because if sales adoption is not something I can point to with evidence, then any statement I make about it is content rather than proof, and content does not count. So I am not going to dress a hope up as an outcome. The honest position is that the marketing-facing half of this system has visibly changed how the team works, while the sales-facing half is still a promise I have not yet found a way to measure outside of me being a sales nagger. Building that feedback loop, some structured signal of whether the ICP and the proof tiers are actually showing up in real deals, is the next thing on the list; baby steps.
The guardrails
This system is built so that the direction comes from a person of knowledge. AI here is the starter and the accelerator; it surfaces patterns, processes far more than I ever could alone, and hands me a strong first draft, and a machine-made first draft is perfectly fine. A machine-made, unedited, shipped draft is another matter entirely and it applies to my own work every bit as strictly as it applies to anyone else’s, so everything gets edited. Underneath all of it, the voice profile and the banned-phrase list run quietly to strip out the majority of the AI cadence and the puffery before anything reaches a human who might put their name to it, and going by the AI-slop incident, that filter clearly earns its place even on my own hand-written prose.
There is also a “no” built into the system, and it is the guardrail people skip most often. Not every feature deserves to become a headline, and some capabilities matter as proof or context rather than as the main message, which is exactly why the launch tiering skill exists: so that “this deserves a full campaign” becomes a decision with evidence behind it rather than a reflex from whoever built the feature and would like to see it celebrated. When the skill comes back with “Tier 3, no brief”, that is the function doing its job. The friction job. And that might be a whole discussion for another day.
Where this leaves me
The Centre of Excellence started as eighteen pages in Word, written over three days, and it felt like finally writing down what should have been obvious all along. The blueprint made the decisions, and the skills were meant to carry those decisions into every task, every day, so that consistency would not depend on whether anyone felt like following a process on a given Tuesday.
Six months in, the honest scorecard is mixed, and I am fine saying so. The steady work output is genuinely better and faster, and the team has changed how it operates around me, which is the result I actually care about the most. The launch chain proved that I can build the machinery, and it also proved that the machinery is worthless when the story or the details keep on being changed or not communicated effectively until last minute. The sales adoption I built all of this for is still unmeasured, and that one is on me to fix next.
That is the real state of it: not finished and certainly not magic, but working where the foundation is solid and breaking where it is not, exactly as designed, whether I happen to like it or not.
I also want to thank some people who helped me kickstart things where I was confused or in the dark – Andrea Saez, Yi Lin Pein, and Rory Woodbridge as a starting point with others taken into account for their wealth of knowledge: Mary Sheehan, April Dunford, Emma Stratton, and many many more. The more product marketing specialists and brains I can access, the better I become. You will probably never read this blog, but thank you all from the bottom of my heart.