ai automation build vs buy

Build vs Buy AI Automation for B2B SaaS

By Brian Shelton — Founder of GrowPredictably.com

TL;DR: Build versus buy is the wrong first question for B2B SaaS AI automation, because the answer changes task by task rather than company by company. Sort each workflow by how routine it is and how reversible the consequences are, decide what AI should touch at all, then decide who builds it. You are not choosing a supplier. You are sorting work.

Key Takeaways

  • Every guide ranking for this question sorts the company by size or stage, then hands down one answer for every workflow that company runs.
  • The unit of decision is the task. Some workflows should be bought, some should be built, and a few should not be automated at all.
  • Harvard Business School and BCG found that on work outside AI’s reach, people using AI were 19 percentage points less likely to reach a correct answer than people using none.
  • Automate at the stage that is capping growth, because a partner’s first suggestion can reflect their backlog rather than your constraint.
  • Whoever builds the system, somebody inside your company has to own the weekly routine that keeps it running.

I have watched marketing leaders hire and fire agencies for 15 years, from the agency side and from the in-house side of that table. The pattern that repeats is rarely a bad vendor. It is a company that picked a delivery model before it sorted the work, then applied that single decision to workflows which had nothing in common.

This article walks B2B SaaS marketing and growth leaders through the AI Collaboration Matrix, the Growth Diagnosis, and the twelve-month comparison that settles build, hire, or both.

It is written for the VP of Marketing or the founder who has had two conversations with automation partners, heard a price and a workflow in both, and noticed that neither one asked what was actually capping growth.

What follows is diagnosis before prescription at the automation layer. Rather than asking which supplier to choose, it asks which work you are choosing for.

Why is build versus buy the wrong first question for B2B SaaS?

Build versus buy is the wrong first question because it sorts the company instead of the work. A B2B SaaS team runs dozens of workflows with different stakes and different judgment loads, and one supplier decision cannot be correct for all of them. You are not choosing a supplier. You are sorting work.

The question gets asked that way because that is how the market answers it. In a September 2026 review of the four guides ranking for this term, one was published by a software vendor, two by agencies, and one by a firm selling custom builds. Every one of them resolves toward what its author sells.

They also share a method. Each classifies the company by revenue, headcount or stage, then issues a single verdict. One of them ranks a stage-based decision framework keyed to annual recurring revenue, which is tidy and easy to apply and gives the same answer for lead routing as it gives for positioning work.

That is where the damage starts. A company decides it is a buy company, and the positioning work goes out with the ticket triage. Or it decides it is a build company, and because that verdict applies to everything, the routine reporting that could have been running in two weeks waits nine months behind a hiring plan.

None of this means the ranking guides are dishonest. It means they are answering a different question, which is what a company of your shape usually does, and that is a question about averages rather than about your roadmap.

Sorting the work first is slower for about an afternoon and it changes every decision after it. Some of your workflows belong with a partner. Some belong with your own people. A few belong nowhere near a model, and knowing which is which is worth more than any vendor comparison.

What does the research say about which work AI should touch?

AI capability is uneven in a way that does not track how difficult a task looks to a person. Two tasks can feel equally hard and sit on opposite sides of what the model can actually do. That is why the question of what to automate has to be answered before the question of who builds it.

Harvard Business School and BCG ran a pre-registered experiment with 758 BCG consultants, randomly assigned to work with no AI, with GPT-4, or with GPT-4 plus a short training. On tasks inside AI’s reach, the people using it completed 12.2% more tasks, 25.1% faster, with quality roughly 30% above the control group.

The least experienced participants gained the most.

Then the researchers set one task deliberately outside that boundary, a brand strategy case that required reconciling numbers with softer signals buried in interview notes. The result reversed. Participants with AI access were 19 percentage points less likely to produce a correct recommendation than participants with none at all.

Same people. Same tool. One task improved, one degraded, and nothing on the surface of either said which was which. That is the part worth carrying into a supplier conversation, because it means the risk is not evenly spread across your workflows even when they look equally routine.

A follow-on field study of 244 consultants looked at how people actually worked with the tool rather than whether they used it, and found three distinct modes. Some kept human and machine work separated and got deeper in their own field. Some fused with the tool and got better at working with it.

A small group handed over both the thinking and the doing and got better at neither.

Both studies looked at management consultants rather than B2B SaaS automation programs, so treat them as evidence about how AI changes knowledge work rather than as a verdict on your roadmap. What they establish is enough: capability is jagged, the mode of use matters, and neither can be sorted by instinct.

How do you classify a task before deciding who builds it?

Classify the task on two things: how routine it is, and how reversible the consequences are. Those two questions put every workflow you own into one of four modes, each with a clear rule about what the machine does and what a person does. The classification takes about 30 seconds and it decides everything downstream.

That is the AI Collaboration Matrix, and it sorts the work rather than the tool.

Reversible consequencesConsequential
Routine workAI drafts, a person edits for accuracy and tone. Safe to hand to a partner.A person owns the decision, AI challenges the assumptions. Build the plumbing outside, keep the call inside.
Ambiguous workAI explores options, a person picks and shapes. Mixed, and usually a split.A person leads throughout. AI never frames the problem first. This is the work that should not leave your building.

Copy that grid and put ten of your real workflows into it. Weekly performance summaries, lead routing, support triage, onboarding sequences, pricing pages, positioning, quarterly planning, review responses, proposal drafts, and whatever else eats your team’s hours.

If you are still deciding which workflows are even candidates, start from the ways B2B SaaS teams actually use marketing automation and classify from there.

What each mode means for build versus buy

The pattern falls out quickly. Routine work with reversible consequences is the safest thing you will ever hand to an outside builder, because a mistake is visible and cheap and the specification is stable enough to write down.

Consequential work is different, and the split matters. Somebody can build the mechanism for you. What cannot be handed over is the judgment about what good output looks like, and the prompts that carry your positioning are part of that judgment rather than part of the plumbing.

Three of the four modes do not permit handing the thinking to the model at all. That boundary is the useful part of the framework, because it is the part every vendor conversation skips.

The failure this prevents

The common failure is not a bad classification. It is no classification, which shows up as a team running one mode for everything. They either bring full executive attention to drafting a weekly update, or they let the model frame a pricing decision before anyone has formed a view.

The recovery is the 30 seconds. Classify before you open the tool, and because the mode is fixed in advance, the conversation cannot drift into the wrong one halfway through.

Once ten workflows are sorted, the supplier question answers itself for most of them, which is why this step comes first rather than after a shortlist.

What should a B2B SaaS team automate first?

Automate where the business is actually constrained, which is rarely where the most impressive demo lands. Growth is capped at one stage at a time, and effort spent anywhere else produces activity rather than lift. Find that stage first, then ask what automation would do for it, and only then ask who should build the thing.

That is the Growth Diagnosis: read the stages a buyer moves through alongside the numbers that move at each one, and identify the single stage holding the system back. Its most useful output is the list of things not to invest in this quarter.

Applied to automation, it cuts the roadmap down fast. If your constraint is how long a qualified lead waits for a human reply, then automating reporting is motion. It will produce a dashboard nobody disputes and a pipeline number that does not move.

If your constraint is that new customers never reach a first win, an outbound sequence built by anyone is aimed at the wrong end of the journey.

It also gives you a defense against sequencing you did not choose. A partner’s first suggestion can be the workflow they have built most often, which is a fact about their backlog rather than about your business, and it is worth naming out loud in the first call.

Finding the stage is less involved than it sounds. Walk the journey with your own numbers beside it and look for the point where volume drops hardest between one step and the next, then check whether that drop has been there for more than a quarter.

A one-off dip is noise. A stable cliff is the constraint, and it is usually somewhere nobody on the marketing team is measured on, which is exactly why it survived this long.

Write the answer as one sentence before you talk to anyone: the workflow, the stage it serves, and the number that should move if it works. If you cannot finish that sentence, no supplier decision will rescue it. The same rule applies to choosing which asset to build next, where the gap decides rather than the appetite.

Build or hire: what actually differs?

Both models belong in a working B2B SaaS operation, so the useful comparison is dimension by dimension rather than a verdict. What differs is speed, where the knowledge ends up, how cost behaves, and who is awake when it breaks.

DimensionHiring it outBuilding in-houseWhich way to go
Time to a first working systemWeeks, because they have built this shape beforeMonths, including hiring and rampHire when the workflow is on the constraint and waiting is expensive
Where the knowledge ends upWith them, unless you contract for handoverWith you, by defaultBuild when the workflow encodes something proprietary about how you sell
How cost behaves as volume growsSteps up with scope and change requestsFlat once the person exists, until they leaveHire for bounded scope, build when the work keeps expanding
What happens when the tools changeTheir problem, if the retainer covers itYours, every quarterHire when the surface changes fast and you have no appetite to track it
What happens when it breaks at 3amDepends entirely on what you signedDepends entirely on who you hiredSplit it: whoever builds, someone internal owns the alert
Work carrying your positioningDo notYesBuild, always. This is the judgment the matrix says stays inside

The allocation rule is short. Hire for the bounded, well understood build. Keep in-house the judgment, the prompts that carry your positioning, and ownership of what good output looks like. Split anything that is routine today and consequential next quarter, which is most things that touch a customer.

Everyone in this market says most teams land on a hybrid, and the better guides do name the transition from agency to in-house as the normal path. Almost nobody says which parts go where, and that sentence is the whole value of the answer.

Notice what the table does not do. It never crowns a winner, because for a company running dozens of workflows there is no single winner to crown. What it gives you instead is a verdict per row, which is the only shape of answer that survives contact with a real roadmap.

What does it cost, and how do you price your own?

There is no defensible universal rate for this work, and every published range you will find is the seller’s own pricing presented as market data. The way to decide is to compare the twelve-month total cost of each option using your own numbers. The line items below are what belong in that comparison.

Building in-house costs more than the salary. Count recruiting, the loaded cost of the person, the ramp before they produce anything, tools and infrastructure, the management time to run them, and ongoing maintenance once the first workflow is live.

Hiring out costs more than the retainer. Count discovery and scoping, the build itself, recurring fees, the internal oversight nobody budgets for, change requests when scope moves, and the exit, including whether the handover is contracted or assumed.

Here is the arithmetic with made-up numbers you should replace with your real quotes and your real loaded costs. Say a partner quotes 18,000 dollars to scope and build one workflow plus 3,000 dollars a month to run it, which is 54,000 dollars across the year.

Say the in-house route is a 130,000 dollar loaded salary, four months before the first workflow ships, and 8,000 dollars of tooling, which is 138,000 dollars in year one for output that starts in month five. On those numbers the partner wins year one and the gap closes in year two, and the whole comparison flips if that person is going to run six workflows rather than one.

Two things decide which way your version points. How many workflows the internal hire will actually carry, and how soon the thing needs to be running. Change either one and the answer changes with it.

On the return side, be careful what you count. CoSchedule’s State of AI in Marketing 2025, a survey of 1,005 marketing professionals conducted in December 2024, found 84% reported increased productivity and headlines an average saving of more than five hours a week.

Reported time saved is not recovered cost. Count only the hours you actually reassign to something else, and put the training and process change on the cost side where they belong.

Why do AI automation rollouts never reach the numbers?

The build usually works. What fails is everything around it, because the tools changed and the behavior did not. A system that runs after the person who built it steps back is engineered that way on purpose, and most are not.

Call that discipline Systems That Stick. It has two halves, and both have to be engineered rather than hoped for. The machine half is the automation itself, wired so the numbers update and the follow-up fires without a babysitter.

The human half is the weekly routine, installed as a small designed habit rather than left to willpower. Remove either half and the system decays as soon as attention moves elsewhere.

The evidence says this is an organizational problem rather than a tooling one. Microsoft’s 2026 Work Trend Index, a survey of 20,000 knowledge workers who use AI at work across ten markets fielded between February and April 2026, found that only 19% sit where their own capability and their organization’s support line up.

One in ten are capable people blocked by the absence of a system around them. That is the condition this half of the work responds to rather than proof that any particular method fixes it.

For the build versus buy decision it lands in one place. Whoever builds the thing, somebody inside your company has to own the weekly cadence, or the engagement ends and the system quietly stops.

Owning it is smaller than it sounds, which is the point of designing it rather than hoping for it. One named person, one standing slot in a meeting that already happens, and three things to look at: did the workflow run, did anything land in the exception queue, and did the number it was built to move actually move.

Anchoring that review to a meeting already on the calendar is what keeps it alive through a bad week, because a habit attached to an existing routine survives where a new recurring invite does not.

So add one question to every partner conversation before you sign. What does our team have to do every week for this to keep working, and who trains them to do it? A good partner has an answer ready. A weak one treats the question as somebody else’s scope, which tells you what month four looks like.

The same question is worth asking when you are choosing an operating model for the whole function rather than a single workflow.

Ready to find the one workflow worth automating first?

Sort the work before you choose the supplier. Put ten of your workflows on the two axes, mark which one sits on the stage currently capping growth, and you will have changed the conversation with every vendor you talk to next. The classification takes an afternoon and it survives every tool that comes out after it.

Run the Growth Gap Scan on your own funnel and you will know which stage is holding the system back before you decide what to automate there.

Frequently Asked Questions

Should we build AI automation in-house or hire an agency?

Both, allocated per workflow. Hiring out suits bounded, well understood builds where the specification is stable and a mistake is cheap. Building in-house suits work that encodes something proprietary about how you sell, and anything carrying your positioning. A company running dozens of workflows will not have one correct answer for all of them.

What is the difference between buying an AI tool and hiring someone to build automation?

Buying a tool gives you a capability you configure yourself and maintain forever. Hiring a builder gives you a workflow wired into your own systems, with the knowledge sitting wherever the contract puts it. The decision that matters is not which one you pick, but whether you settled who owns the running of it afterwards.

Which work should never be handed to AI?

Work that is ambiguous and consequential at the same time. Positioning, pricing judgment, strategy calls a customer will interrogate. Harvard Business School and BCG found that on a task built to sit outside AI’s reach, people using AI were 19 percentage points less likely to reach a correct answer than people using none at all.

What should a B2B SaaS team automate first?

The workflow sitting on the stage currently capping growth. Walk your buyer journey with the numbers beside it and find the point where volume drops hardest and has kept dropping for more than a quarter. Automating anywhere else produces activity. A partner’s first suggestion usually reflects their backlog rather than your constraint.

How do you price AI automation against doing it in-house?

Compare the twelve-month total cost of each using your own numbers. In-house means recruiting, loaded salary, the ramp before anything ships, tools and management time. Hiring out means discovery, the build, recurring fees, internal oversight, change requests and the exit. Published market ranges are almost always the seller’s own pricing presented as market data.

Why do AI automation projects stop working after a few months?

The build works and the behavior around it does not. A system that outlasts the person who built it needs two engineered halves: the automation itself, and a weekly routine somebody owns. Microsoft’s 2026 Work Trend Index found only 19% of AI users sit where their own capability and their organization’s support actually line up.

Can you use an agency and an in-house team at the same time?

Yes, and most teams end up there. The useful version is not a vague hybrid but an allocation: the partner builds the bounded plumbing, your people hold the judgment, the prompts that carry your positioning, and the standard for what good output looks like. Split anything routine today that turns consequential next quarter.

Similar Posts