Walk into any marketing organization right now and you'll hear some version of the same ambition: "We're going to use AI agents to do more, and do it faster!" That's a good instinct, because agents really are good at the repetitive, data-heavy work that eats marketing teams alive.
But what usually doesn't come along with the ambition is an operating model to go with it. So a few months in, a lot of teams find themselves with a folder of impressive demos, a pilot channel that's gone quiet, and no clear line back to business value.
I've come to see that as a process gap more than a technology one, which is actually good news, because process is something we already know how to fix. If you treat an AI agent as a tool you buy or a feature you launch, you've skipped the part that matters most. It's much closer to a hypothesis, and hypotheses need somewhere to be tested.
For me, that somewhere is an incubation team: a small, cross-functional pod whose job is to figure out quickly whether a given agent is worth building at all.
Here's the model I use to run that lab, which I call the PROVE framework. It pulls from three disciplines I keep coming back to:
Before any of the framework matters, there's one thing I won't budge on: you establish business value and ROI first, and then you build.
I know how obvious that sounds, but I also know how tempting it is to skip this step. A shiny new model has real gravitational pull, and honestly, prompting is a lot more fun than sitting down to define what "good" would even look like. But if you can't tie a proof of concept back to a metric, you don't really have one.
The model moves through three gates, and each one asks something different of you. Together, they're what keeps a pilot from tipping too far toward ambition on one side or too much caution on the other.
If you scale before you've proven value, you'll burn through the budget. But if you stay in learn-fast mode indefinitely, you'll never ship anything at all. The key to success is knowing which gate you’re in and behaving accordingly.
Inside those gates is a five-stage loop you run over and over. Each stage produces one clear output, and you don't move on to the next one until you have it in hand. I use "PROVE" as the shorthand for:
Start with the customer journey, not the technology. Where in your customer's experience (or in your team's own workflow) is there a moment that's repetitive, judgment-light, data-rich, and expensive in time or money? When you find something with all four of those traits, you've very likely found your first agent.
It helps to know what the capability stack underneath an agent needs to look like before you commit to a candidate, because that's what tells you whether "in weeks rather than quarters" is realistic.
Diverge across candidates, then converge on one by running it through four lenses:
The output of Pinpoint is a single sentence, your value hypothesis: for this user, at this moment, we believe an agent that does X will produce outcome Y, and we'll know we're right when metric Z moves by a set amount.
This is the stage where pilots need the most attention, because an agent is only as capable as the data and context it can reach.
Ready is a deliberate inventory of two things, borrowed from journey-mapping discipline:
For each of those, you're confirming that access is genuinely wired up, that the data is clean and current, that using it squares with your consent and privacy commitments, and that you've been explicit about what this agent isn't allowed to do or say. If the foundation underneath is shaky, that work comes before the agent rather than after it. It's the same sequencing problem I've written about with identity architecture, where skipping rungs produces use cases that quietly underperform.
I call this whole inventory the data and trust spine, because it's what holds up everything else. It decides what your agent is able to do, and it decides whether you can trust what it hands back.
Now you can build, as long as you keep it to the smallest version that can actually test your value hypothesis. You're not building the full vision here, just a slice of it you can put in front of someone.
Orchestrate is where the model, the data connections, the prompts or workflow logic, the tool calls, and the human-in-the-loop checkpoints get wired together. Anything you add beyond what the test requires could potentially slow your learning down. The output is a working agent, in a sandbox, ready to be tried on real tasks.
Put the agent in front of real users doing real work, and measure it against two bars:
Keep in mind, though, that neither bar tells you much if your measurement can only show you what happened at the surface. Getting to analysis that explains why a number moved is what separates a pilot result you can act on from one you have to take on faith.
Expand is the deliberate work of moving from sandbox to production, which means integrating the agent into everyday workflows, hardening governance and monitoring, training the people who'll use it, and setting up ongoing ROI measurement so the value you proved doesn't disappear six months later.
Agents making real decisions inside live workflows are already doing this, and what that looks like in B2B marketing and decisioning is a useful picture of what "in production" means.
Expand also closes the loop. A mature incubation practice always has a next hypothesis queued, so the pod moves from one proven win to the next bet without losing tempo.
In my experience, the right unit for this work is the pod itself, small enough to move and complete enough to ship. One person can't carry it, and a whole department will move too slowly to learn much of anything.
Here's an example lineup to consider against your own org chart:
Above the pod, an executive sponsor protects its time and clears organizational roadblocks while letting the team manage the day to day.
The goal of an incubation team is twofold: driving adoption and innovation, and finding, analyzing, and building the use cases worth scaling. The organizations that win with AI won't be the ones with the most demos. They'll be the ones that built the muscle to prove, through repeatable and collaborative processes, which transformation projects are actually worth it.
I'll be at TAG's //SHIFT AI summit in Atlanta on September 18, speaking about this exact topic, so come join me if you’re in the area! And if you’ve got questions about what your first marketing agent pilot should be, or what your data needs to look like before one can run, we'd love to talk.