Key Takeaways
- A useful B2B marketing agency evaluation scorecard weights four things heavily: proof of AI adoption inside the agency's own workflow (25%), a defensible speed model (20%), senior time genuinely on the account (20%), and a pricing shape that matches your risk (15%), with capability fit, references, and cultural signal splitting the remaining 20%.
- Score every criterion 1 to 5 with a written reason, not a gut feel. A vendor who scores a 4 on AI adoption should be able to show you the repo, the skills, the prompts, or the internal tools, not a slide that says "we use AI."
- The single fastest disqualifier in 2026 is an agency that cannot describe, in specifics, how their own delivery has changed in the last twelve months. If the workflow they sell you is the workflow they used in 2023, you are paying 2023 prices for 2023 output.
- Senior time is the criterion buyers most often overstate and agencies most often understate. Ask for named hours per week on your account from people with a title you would hire directly, and put it in the SOW.
- Fixed monthly retainers hide risk on both sides. Pricing that flexes with scope, or a hybrid of a small retainer plus outcome milestones, is a healthier shape for the current pace of change.
A CMO I was talking to last month walked me through the spreadsheet her team was about to use to pick a new agency. Case studies, team bios, reporting cadence, cultural fit — inherited from a review they'd run in 2019, dusted off, minor edits. None of those are bad things to evaluate. They just wouldn't tell her whether the agency she was about to hire actually knows how to work in 2026. She'd have hired the wrong shop and not known why for nine months.
So I want to give you a scorecard that asks the right questions. Weighted, numbered, with scoring guidance you can hand to a junior on your team and get a consistent read back. I built it because we sit on the other side of these evaluations constantly, and the ones that go well for both sides are the ones where the buyer knew what they were actually measuring.
One thing before we get into it. This is a scorecard for choosing a senior B2B agency partner, not a procurement checklist for a commodity vendor. If you are buying a point service, say paid media execution or a one-off website, the weights below are wrong for you. Use Elevation's checklist or Column Five's dimension list for that kind of buy. What follows is for the strategic partner decision.
The B2B Marketing Agency Evaluation Scorecard
Score each criterion 1 to 5. A 3 is "they can do this competently." A 5 is "they can show me the mechanism behind the outcome." A 1 or 2 means they either could not answer or the answer was a slide. Multiply each score by the weight, sum, and the highest total wins, but pay attention to any 1 or 2 in a category weighted 15% or more. That is a veto, not a rounding error.
1. AI adoption proof (weight: 25%)
This is the criterion that has changed the most and where the gap between agencies is now the widest. You are not scoring whether they "use AI." Everyone will say yes. You are scoring whether AI has restructured their delivery in a way you can inspect.
- Score 5: They can walk you through named tools in their actual workflow (Claude Code, Cursor, n8n, HubSpot, Figma, GitHub, whatever it is), show you skills or prompts they have built in-house, and describe a workflow that did not exist twelve months ago. They can name a thing they used to do in two weeks that now takes two days, and explain the mechanism.
- Score 3: They use ChatGPT and Midjourney. Individuals on the team are experimenting. There is no shared system.
- Score 1: "We are exploring AI" or a slide with logos.
The reason this gets 25% is simple. The gap between an AI-native agency and a traditional one in 2026 is often three to five times faster on the same deliverable, at higher quality, because a senior person is now doing what used to require a team of four. This is the criterion the CMO's 2019 spreadsheet had no column for, and it is the one that would have cost her the most. If you are paying a traditional shop, you are paying for a cost structure that a competitor has already eliminated.
2. Speed model (weight: 20%)
Separate from AI adoption, because speed is a promise about how the work gets to you, not just what tools produced it. Ask them to describe their delivery cycle for a real recent project. How long from kickoff to first live artifact? How many revision rounds are typical? Where does the work sit between rounds?
- Score 5: They can name a cycle time in days, not weeks, for the kind of work you are hiring them for. They can explain the specific practice, for example, working live in the source system rather than in mockups, or shipping components directly from a design system to production, or reviewing in-context rather than in decks, that makes the speed real.
- Score 3: Standard agency cadence, weekly status, two-to-three-week deliverables.
- Score 1: Cannot answer without checking with an account manager.
3. Senior time on the account (weight: 20%)
This is the one buyers ask about and then never enforce, and it is the second place the 2019 spreadsheet fails quietly. The pitch is delivered by the founder and two principals, warm and sharp and exactly the people you want. You sign. The work is delivered by a senior strategist, two juniors, and a project manager, and six months in you cannot remember the last time the founder was in a room. Don't ask whether they promise senior time. Every agency promises senior time. Ask who, how many hours, and whether they'll put it in the SOW. A 5 here means all of that is on paper. A 3 means senior involvement is described but not committed in writing. A 1 means you'll meet the senior team at kickoff and at the QBR. And the best agencies, when you ask why their senior time is so high, will tell you something specific about how AI has collapsed the pyramid so that senior people are now doing the making, not just the reviewing. That answer is the tell.
4. Pricing shape (weight: 15%)
The number matters less than the shape it takes. A flat monthly retainer for eighteen months is a bet by both parties that nothing will change, and something will change. So you are looking at how the pricing flexes. A 5 has a small base plus scope-based or outcome-based components, a clear mechanism to scale up or down without renegotiating the entire contract, and, this is the mature-model tell, the ability to articulate what they will not do at a given tier. A 3 is a flat retainer with a clearly scoped SOW and a review point at six months, which is livable. A 1 is a flat retainer, twelve-month minimum, vague scope, which is the shape that punishes both sides the moment the world moves.
5. Capability fit for your actual next twelve months (weight: 8%)
Score against the specific capabilities you need in the next four quarters, not against their whole capability deck. Your roadmap, not theirs.
6. References who look like you (weight: 7%)
Two references, minimum, from companies within one stage and one category of yours. A great reference from a Fortune 500 does not tell you much if you are a Series B. Score 5 if the references are close analogues and will get on a call. Score 1 if the references are logos on a page.
7. Cultural and communication signal (weight: 5%)
How they behave in the sales process is how they will behave in the work. Do they push back on your brief when it is wrong? Do they answer email in a day or a week? Score what you saw, not what they promised.
How to actually run the scoring
Have two people on your side score independently, then compare. The conversations you have about why one of you gave a 4 and the other gave a 2 are more valuable than the final number. If your scores are within a point on every criterion, one of you is not paying attention.
Weighted totals will usually cluster within ten points across a competitive shortlist. That is fine. The scorecard's job is not to pick the winner for you, it is to force the conversation about the criteria that actually matter and to surface the vetoes. A 1 or 2 in AI adoption, speed model, or senior time is a veto in 2026, even if the total score is competitive. Those are the criteria where the gap between agencies is now structural.
When hiring a senior B2B agency is the wrong call
I should say this plainly because it belongs in an honest scorecard. If you have a strong in-house team, a clear ICP, a working demand engine, and the work you need done is executional, hire freelancers or a specialist shop and skip the senior agency retainer. You will pay less and move faster. A senior agency earns its fee when the strategy is not settled, when the org needs a partner who can operate across brand, demand, product marketing, and web, or when you need to import a way of working your team does not have yet. If none of those are true, the scorecard above will still work, but you are probably solving the wrong problem by running an agency search at all.
The other honest case: if your budget is under roughly fifteen thousand a month, a senior agency is not the right shape. The math does not work for either side, and you will get a junior team pretending to be senior. Better to hire one strong contractor and buy tools.
What to do with the scorecard once you have it
Send it to the agencies before the final meeting. Tell them the weights. The good ones will come prepared to score high on the things that matter, and the ones who cannot will disqualify themselves before you have to. This is a favor to everyone, including the agencies you do not pick. If you want a longer version of how we think about all of this, we published the blueprint we recommend for running modern marketing teams, and it pairs well with the criteria above.
Send it early, tell them the weights, and watch who shows up ready. Score for the mechanism, not the slide.
Frequently Asked Questions
How many agencies should be on the shortlist before scoring?
Three is the right number. Two is not enough contrast to make the scorecard useful, and five means you are running a procurement process rather than a decision. If you have more than five, do a lightweight first pass on AI adoption and senior time commitment alone, and cut to three before you invest in a full scored evaluation.
Should procurement own the scorecard or should marketing?
Marketing should own the criteria and the weights. Procurement can own the process, the paper, and the pricing negotiation. If procurement is setting the weights, the highest-weighted criterion will end up being cost, and you will hire the cheapest agency, which is almost never the right answer for a strategic partnership.
What if an agency refuses to commit senior hours in the SOW?
That is your answer. Every agency will tell you senior people are on the account. The ones who will write it down are a different population from the ones who will not. This is the single cleanest filter in the entire evaluation.
How often should we re-score an incumbent agency?
Once a year, using the same scorecard you used to hire them, with the same weights. If their AI adoption score has not moved up in twelve months, that is a signal worth acting on, because the market is moving. A good incumbent will welcome the exercise and often ask to run it on themselves first.
Is this scorecard different for a content agency versus a full-service agency?
The weights shift slightly. For a content-only engagement, capability fit rises to around 15% and speed model rises to around 25%, because content velocity is the whole game. AI adoption stays at 25%, because content is the category where the AI delta is most visible.
Sources
- Top marketing agencies: how to evaluate them with a scorecard
- B2B Marketing Agency Evaluation Checklist
- How to Compare B2B Content Marketing Agencies
- How to Vet a Marketing Agency: Digital Agency Scorecard
- Marketing Agency Evaluation Scorecard: 20 Criteria Before You Sign
- Anyone here worked with a B2B marketing agency that actually delivered?
- 8 Best B2B Marketing Agencies: Reviewed and Ranked