Cloud spend crosses a threshold and someone in finance asks the obvious question: why is this number growing faster than revenue? What follows is usually a scramble — a spreadsheet, a few afternoons in Cost Explorer, a Slack thread that dies after a week.
Eventually the conversation gets serious, and it narrows to three options. Buy a cost management platform. Hire someone whose job is cloud cost. Bring in outside help on a retainer. These get argued as if one is correct and the other two are mistakes. They aren't. They solve different problems, and the right answer depends on facts specific to your organization.
Here is how to think about the decision honestly, including the cases where hiring me would be the wrong move.
What Each Option Actually Does
The three options get compared as if they're substitutes. They mostly aren't.
| Approach | What it gives you | What it doesn't |
|---|---|---|
| Tooling | Visibility. Allocation, anomaly alerts, forecasting, showback and chargeback reporting. | Decisions. Tools identify; they don't negotiate tradeoffs, refactor workloads, or get engineers to change behavior. |
| In-house hire | Continuity and context. Someone who knows your architecture, your team, and your roadmap, present in every planning conversation. | Breadth. One person has seen one company's problems. Ramp time is real and the role is hard to hire for. |
| Consultant / retainer | Pattern recognition and speed. Someone who has seen this failure mode across many environments and knows which levers move first. | Presence. An outsider is not in your standups, and institutional knowledge walks out when the engagement ends. |
Most mature programs end up using all three. The sequencing question is what actually matters: which one do you do first, given where you are now?
The Case for Tooling First
If you cannot currently answer "which team caused last month's increase," buy visibility before you buy anything else. Every other approach depends on data you don't yet have. A consultant's first two weeks would be spent building the allocation model a tool gives you on day one, and you'd be paying consulting rates for it.
Tooling is also the only option that scales without adding headcount. Anomaly detection runs whether or not anyone is paying attention that week.
The failure mode is predictable and extremely common: the dashboard gets bought, it correctly identifies six figures of waste, and eighteen months later that waste is still there. The tool was never the bottleneck. Nobody owned the remediation, and no one had authority to tell a product team their idle staging environment was getting shut down.
A cost tool converts an unknown problem into a known one. That is genuinely valuable and it is not the same as fixing anything.
Choose tooling first when: you lack allocation data, you have engineering capacity to act on findings, and someone credible already owns cost as part of their remit.
The Case for an In-House Hire
The argument for hiring is durability. Cloud cost is not a project that finishes. Architecture changes, teams ship new services, commitments expire, pricing models shift. An embedded person is in the design review before the expensive decision gets made, which is worth more than any number of retrospective audits.
Be honest about the cost, though. A capable FinOps or platform engineer in a US market is a meaningful salary, and loaded cost — benefits, taxes, equity, tooling, management overhead — typically runs substantially above base. Add three to six months before they're fully productive in your environment.
There's a second problem that gets underweighted. The role is genuinely hard to hire for, because it sits between disciplines. You need someone who can read a Terraform module and also hold their own in a conversation with the CFO about amortization. That population is small, and the people in it have options.
The third issue is isolation. A solo FinOps hire sees exactly one environment. They get very good at your problems and have no benchmark for whether your problems are normal. Some of the most expensive mistakes I see come from smart internal people optimizing hard in a direction that was wrong from the start, with nobody positioned to say so.
Choose an in-house hire when: spend is large enough that the role pays for itself several times over, cost decisions are continuous rather than episodic, and you can realistically attract the profile. If you're going to hire, hire early enough that the person shapes architecture rather than cleaning up after it.
The Case for a Consultant or Retainer
The honest argument for outside help is not that consultants are smarter. It's that the first pass through an unoptimized environment is a pattern-matching exercise, and pattern matching improves with volume.
The waste categories repeat with remarkable consistency: orphaned storage and snapshots nobody deleted, over-provisioned instances sized for a load test from two years ago, non-production environments running nights and weekends, commitment coverage that drifted out of alignment with actual usage, cross-AZ data transfer nobody costed at design time, logging retention set to defaults. Someone who has worked through this repeatedly knows which of these is likely to be biggest in your environment before opening the console.
Speed matters here for a reason people miss. Cloud waste compounds. Every month an over-provisioned fleet runs is money that doesn't come back, so a finding delivered in week three is worth materially more than the same finding in month six.
The second real argument is political, and I'd rather say it plainly than dress it up. Telling a senior engineer their service is overspending is easier for someone with no stake in internal politics. Outside recommendations get evaluated on merit more often than internal ones do. That's not a flattering fact about organizations, but it's a consistent one.
The retainer structure exists because the one-time audit model has a known failure: findings decay. An audit produces a report, some of it gets implemented, and the environment drifts back. Ongoing engagement is what prevents the ratchet from slipping.
Choose a consultant when: you need results faster than a hire could ramp, spend doesn't yet justify a full-time role, you've had findings identified but not implemented, or you want a benchmark against how comparable environments are run.
When a Consultant Is the Wrong Answer
Some situations don't call for outside help, and it's worth naming them.
- Spend below roughly $500K/year. The absolute savings often won't clear the engagement cost. Buy a tool, spend two focused engineering weeks on the obvious waste, revisit later.
- No allocation data at all. Fix tagging and account structure first. Analysis on unattributable data produces confident-sounding conclusions built on nothing.
- No capacity to implement. If engineering is fully committed for two quarters, recommendations will sit. Findings nobody can act on are an expensive way to feel productive.
- Mid-migration or mid-replatform. Optimizing an architecture you're about to replace is wasted effort. Wait until the target state is stable.
A Practical Sequence
For most organizations between $2M and $10M in annual cloud spend, the sequence that works looks like this.
- Establish allocation. Tagging standards, account structure, and a tool that reports by team and service. Without this nothing else is measurable.
- Run an intensive first pass. Internal or external, this is where the large one-time reductions live. Our typical engagement lands in the 15–25% range within the first 60 days, and most of that comes from a handful of categories rather than a long tail.
- Establish governance. Anomaly alerts with named owners, cost review in the architecture process, commitment strategy on a review cadence.
- Decide on permanence. Once steady-state, the question becomes whether ongoing work justifies a hire or a lighter retainer. That answer depends on how fast your architecture is changing.
The mistake is treating step two as the whole program. The first pass is the most visible phase and the least durable one. Savings that aren't defended by governance erode — not dramatically, just steadily, as new services ship and defaults reassert themselves.
The Question Worth Asking
Rather than "which of these three should we buy," ask: what is the specific thing currently preventing our cloud costs from being under control?
If the answer is "we can't see where the money goes," that's tooling. If it's "we see it, and nobody has time or authority to act," that's a hire or a retainer, depending on whether the need is continuous. If it's "we've tried and the savings didn't stick," that's a governance problem, and buying another tool will not fix it.
Most organizations know which of those sentences is true. The decision gets difficult when the honest answer is uncomfortable — usually because the real constraint is organizational rather than technical, and no purchase resolves that.
Not sure which applies to you?
A cloud spend review will tell you where your costs actually sit and which of these three approaches fits your situation — including if the answer is that you don't need outside help yet.
Book a cloud spend reviewRelated Reading
- FinOps for Semiconductors — simulation workloads and capex predictability
- FinOps for Fintech — cost per transaction under SOX and PCI-DSS
- CleanTech Cloud Optimization — runway extension without slowing R&D
- HIPAA-Compliant Cloud for Biotech — audit-ready genomics infrastructure
- 5G & Telecom FinOps — multi-region cost under SLA constraints