Key Takeaways
- Ask what lives in their GitHub repo. An agency that actually uses AI has skills, prompts, tokens, and workflows checked into version control that the whole team can pull down, not a handful of ChatGPT tabs open on someone's laptop.
- Ask who reviews the AI output before it ships. Real adoption has a named senior reviewer (a strategist, a CTO voice, a creative director) and a defined checkpoint, not vibes.
- Ask for the workflow by name. If they can't tell you which tools sit where in the chain (Claude Code, Cursor, Figma, n8n, HubSpot, Webflow, Vercel), and which humans touch which step, the AI is decorative.
- Watch for opportunity thinking. Agencies using AI well are building new deliverables and new business models with the time they saved, not just cutting headcount.
- Beware the volume tell. A sudden jump in output with no jump in performance is a signal the agency is generating, not thinking.
If you're a CMO or founder shopping for a growth agency right now, you already know the pitch deck says "AI-powered, " because every deck says that, and the interesting question isn't whether an agency uses AI, it's whether the AI is doing structural work inside the agency or whether it's a layer of paint on the same brochure work you were buying in 2019.
I run a B2B tech marketing agency and I've spent the last two years rebuilding how we work around agentic tools, so I've had the same conversation from both sides of the table, and I've watched a lot of shops describe "AI capabilities" that turn out to be one person with a Grammarly subscription and a ChatGPT tab. The good news is you can tell the difference in about twenty minutes if you ask the right five or six questions. This post is that list, with the reasoning behind each one, so you leave holding something you can actually use in a pitch meeting.
The pitch tells you nothing. The plumbing tells you everything.
Every agency will say they use AI, and at a minimum most of them have for years already if you count Grammarly and the smart features baked into the Adobe suite. That's true, and it's also useless as a buying signal. What you're actually trying to figure out is whether the agency has restructured how work gets done, or whether they've bolted a chatbot onto the side of the old process and called it a capability.
The way I'd test this is to ask the agency to walk you through one recent deliverable end to end, focusing on the workflow rather than the outcome. Who wrote the brief, what tool ingested it, which model produced the first draft, what skill or prompt system shaped it, which human reviewed it, where does the source of truth live now. Agencies that have actually rewired their process will answer that in specifics and name the tools without hesitation. Agencies that haven't will answer in adjectives.
Question one: what lives in your repo?
This is the single most revealing question you can ask, and almost no buyer asks it. Ask the agency where their prompts, skills, brand tokens, and automations live.
The honest answer at a shop that's done the work is something like: in GitHub, with a readme, versioned, so any team member can pull it down and use it. The pattern to look for is a repo that contains skills for voice, tokens (basically CSS variables for brand), component build rules, page build rules, and a set of Claude MD rules that tell the model how to combine all of it when it builds something new. Marketers don't typically think about GitHub as a home for marketing work, but that's exactly what it becomes once a team starts building skills and workflows it wants everyone using consistently. If an agency's AI work isn't in a repo, it isn't scalable across their team, which means the AI expertise sits with one or two people and you're paying for the rest of the team to catch up on your dime.
The tell here is specificity. Ask them to name a skill. Ask what's in it. Ask who wrote it and who's used it in the last week.
Question two: who is the senior reviewer, and what are they reviewing against?
Generation is cheap, and judgment is the part that still separates work you'd ship from work you wouldn't, which is why the review step matters more than the generation step. The pattern we'd recommend is a named senior reviewer in the loop before anything goes out — sometimes a human, sometimes a Claude project built to act as a senior CTO voice or a senior strategist that flags risks, reorders priorities, and gives a second perspective when there isn't another senior person available on the call. The point isn't the specific tool. The point is that somebody, or something shaped like somebody senior, is checking the output against a defined standard. Ask the agency who reviews AI output and what they're checking for. If the answer is "the account lead reads it, " you're looking at wishful thinking dressed up as a process.
Three workflow tells that show up in the day-to-day
Beyond the repo and the reviewer, there are three workflow tells I look for. None of them require you to be technical to check.
The first is the design and build workflow. Ask how they get from a concept to a live page. The old-world answer is: Figma, then Webflow, then a developer tweaks CSS by hand. The new-world answer sounds like: we design in Figma, feed the design system into Claude Code as a skill, chat components into existence using our tokens, push to GitHub, and deploy to Vercel with a lightweight CMS attached only where we need collections. Both answers are legitimate, but the second one runs at a different speed and produces sites that hit 99 on PageSpeed Insights because there's no application bloat sitting on top, the sites we used to build on WordPress with a page builder would land in the low 60s on the same test, and the jump surprised me the first time I saw it because I'd assumed the ceiling was somewhere in the 80s. The workflow tells you what era the agency is operating in.
The second is the presentation and scenario workflow. Ask how they present strategy. If the answer is Google Slides, fine, but ask if they ever build interactive artifacts, page-by-page tools with sliders and dials that let you run scenarios live in the room. Marketing work is scenario work: if we shift spend here, what happens there. Agencies using AI well are increasingly building these interactive pieces instead of static decks, because it takes hours now instead of weeks.
The third is the source-of-truth workflow. Ask what happens to the HTML and CSS an AI tool produces. Does it get manually rebuilt in HubSpot or Pardot, or do they have skills that convert clean output directly into the source system? The gap between those two answers shows up as a multiple, not a percentage, on production speed.
Question three: efficiency or opportunity?
This one is more about mindset than plumbing, but it's the strongest predictor of whether the agency will still be useful to you in two years.
Ask them what they're doing with the time AI has saved them. Agencies stuck in efficiency thinking will tell you about headcount reduction, faster turnarounds, doing more with less. That's real, and it's also the floor. Agencies doing the interesting work will tell you about new deliverables they've launched because the old ones now take a fraction of the time, new pricing models, new services that weren't possible before, new categories of work they can take on. That's where the compounding happens. If the agency is only pitching you efficiency, you're buying a cost center that will get commoditized. If they're pitching you what they can now do that they couldn't before, you're buying a partner who's going to keep finding new ways to be valuable.
Question four: the volume tell
Here's the counter-signal, and it's worth taking seriously: one of the clearest indicators of shallow AI use is an unprecedented jump in the volume of content an agency produces with no corresponding jump in performance. Fifty blog posts a month, none of them ranking, none of them getting shared, none of them producing pipeline. That's generation without thinking, and it's the failure mode of agencies who adopted AI as a productivity hack without adopting it as a craft.
Ask to see recent output and ask what it produced. If the volume is up and the outcomes aren't, the AI is running the shop instead of the other way around.
When hiring an AI-native agency is the wrong call
I should say this plainly: there are cases where you don't need any of this. If your marketing motion is genuinely simple, if you have one channel that works and you just need someone to run it, a specialist shop with a proven playbook and no AI story at all will probably serve you better than an agency mid-transformation. AI-native agencies earn their keep on complex, multi-channel, content-heavy programs where the compounding matters. On a single-channel paid media buy with clean attribution, the AI question is mostly noise. Buy the specialist.
And if you're evaluating your own team's readiness before you go shopping at all, our AI marketing blueprint walks through the skills, workflows, and org shifts we'd recommend working through first.
What to do with this in your next pitch meeting
The next time you sit down with an agency pitching AI capability, you don't need to become technical to pressure-test them, you need four questions and the patience to listen for specifics. Where does your AI work live, who reviews it, walk me through one recent deliverable end to end, and what have you built with the time you saved. Agencies that have done the work will answer with tool names, workflow steps, and named humans, and they'll do it without flinching, because the plumbing is right there for them to point at. Agencies that haven't will answer with adjectives and case studies from 2022, and you'll feel the answers getting vaguer as the questions get more specific. Take the meeting, ask the four questions, and see which room you're in.
Frequently Asked Questions
Is it a red flag if an agency won't name the specific AI tools they use?
Yes. Agencies with real AI practice name tools without hesitation: Claude Code, Cursor, Figma, n8n, HubSpot, Webflow, Vercel, Supabase, whatever their stack is. Vagueness about tooling almost always means the AI story is thinner than the pitch.
Should I expect an AI-native agency to be cheaper?
Not necessarily, and be suspicious if they are dramatically cheaper. The value of an AI-native agency should show up as more strategic depth, faster iteration, and interactive artifacts you couldn't get before, not the same brochure work at a discount.
How do I know the agency isn't just using ChatGPT and calling it a system?
Ask to see their skills or prompt library. Ask where it's stored. Ask who on the team has contributed to it in the last month. A real system has version control, named contributors, and documentation. A ChatGPT tab has none of those.
What if my in-house team is already using AI, do I still need an agency?
Depends on the gap. If your team has the workflow depth described here, and the capacity, you may not. If your CMO is in Claude on their own laptop but the rest of the team hasn't caught up, an agency that has already built the systems can accelerate you by a year or more. Both patterns are legitimate.
How long should it take an AI-native agency to show real work?
Two to four weeks for meaningful artifacts, not months. If they've genuinely rewired their process, first drafts, interactive strategy docs, and working prototypes come fast. If the first real deliverable is still eight weeks out, the AI is theater.
Sources
- How the Best Marketers Actually Use AI (Hint: It's Not a Prompt)
- What You Need to Know if Your Marketing Agency Uses AI
- Marketing with AI: The Right Way to Use It (and What to Avoid)
- AI in Advertising: Is Your Marketing Agency Using AI?
- What Is an AI Marketing Agency? Benefits and Examples
- How to Tell When a Product Is Truly Powered By AI