package/gcp-cost-optimization
GCP Cost Optimization
A model of what your Google Cloud spend should be, with the gap between that and what it is broken into items ranked by saving per unit of effort.
~$14,000 fixed
Starts with Discovery & Scoping
Floors reflect typical engagements. Larger, regulated, or multi-region estates are scoped and quoted after Discovery.
Or 15–20% of realized year-one savings, agreed after read access to billing data.
Where this usually starts
The bill went up again. Finance has asked for an explanation and the honest one is that the platform grew, which is true and unsatisfying. Someone has already looked at the recommendations in the console, applied the easy ones, and found that the number barely moved.
The recommendations engine is good at the local view. It will tell you an instance is oversized. It will not tell you that a third of your fleet is running on-demand while your usage has been stable for eighteen months, that your Windows licensing position would be materially cheaper on sole-tenant nodes, or that a chatty service pair sits in two different zones and you are paying inter-zone egress on every request.
The other reason internal cost work stalls is that it needs someone to hold the whole picture at once — commitment coverage, licensing, cluster packing, storage lifecycle, and network topology — and the people who understand each piece are on different teams with different priorities.
What gets analysed
- Committed use and sustained use discounts. Current coverage against a usage baseline, modelled across term lengths and commitment types. Most estates are either substantially under-committed on stable workloads or committed to the wrong resource shape.
- Sole-tenant BYOL licensing. Whether your Windows Server and SQL Server position is cheaper on dedicated hardware with existing licences than on pay-as-you-go, modelled both ways with the sole-tenant overhead included rather than assumed away.
- Idle and overprovisioned capacity. Instances running with no meaningful utilisation, disks attached to nothing, reserved addresses no longer routed, and environments that were spun up for a project that ended.
- GKE bin-packing. How much of the CPU and memory you pay for is actually requested by running pods. Requests copied from a template are the usual cause, and the gap between requested and used is frequently the largest single line item.
- Storage class and lifecycle. Data sitting in Standard that has not been read in a year, missing lifecycle rules, and multi-region buckets where regional would do.
- Egress patterns. Inter-zone traffic between services that should be co-located, internet egress that should be going through Private Service Connect or private access, and cross-region replication configured once for a reason that no longer applies.
The deliverable is a model, not a list
The output is a savings model: for each opportunity, an estimated annual saving, the effort to realize it, any risk attached, and a dependency order. Some savings are a configuration change. Some require a workload to move. Some require a commitment that is only sensible if a platform decision goes a particular way in the next quarter, and that dependency is stated rather than hidden.
The plan is sequenced by saving per unit of effort, so that the first month of work produces most of the recoverable saving. That ordering matters more than completeness — a list of forty opportunities with no order gets read once and filed.
Two pricing structures, and how to choose
The fixed fee is straightforward: a set price for the analysis and the plan, whatever the analysis finds.
The contingency option is 15 to 20 percent of realized year-one savings. It costs more when the savings are large and nothing when they are not, and it aligns the incentive precisely. It is agreed only after read access to billing data, because committing to a percentage of savings before seeing the spend is a bet rather than a proposal, and a consultant who offers it sight unseen is telling you something about how they price.
Estates with obvious large opportunities usually prefer contingency. Estates that have already been optimised once, where the remaining work is grinding and incremental, usually prefer the fixed fee. Both are on the table and the choice is yours after the numbers are visible.
What is not included
Implementation. This engagement produces the model and the plan; applying the changes is separate, and many clients hand the plan to their own team because most of the items are configuration changes their engineers can make. Where implementation is wanted from Sumech, it is quoted after the analysis, and the items are worth more scrutiny when the same party is proposing and delivering them.
Contract negotiation with Google. Commitment modelling shows what to commit to. Negotiating an enterprise agreement or a private pricing arrangement is between you and Google, and Sumech is not a reseller and takes no margin on your Google Cloud spend.
Non-Google spend. This is a Google Cloud analysis. Third-party SaaS and other clouds are outside it.
Deliverable. Savings model with a prioritized implementation plan.
How the engagement runs
- Week 0 Discovery & Scoping Spend scale, estate shape, and which pricing structure fits. Read access to billing data is arranged here, before any contingency arrangement is agreed.
- Week 1 Baseline Billing export analysed against usage. Spend decomposed by service, project, and workload so that later findings attach to something specific.
- Week 2 Commitment and licensing modelling Committed and sustained use coverage modelled across terms. Sole-tenant BYOL position modelled against pay-as-you-go with dedicated-hardware overhead included.
- Week 3 Capacity, packing, storage, egress Idle and overprovisioned resources identified. GKE requests compared to actual consumption. Storage lifecycle and class reviewed. Egress traced to the topology causing it.
- Week 4 Model, plan, walkthrough Savings model delivered with each item carrying an estimated annual saving, effort, risk, and dependency. Plan sequenced by return per unit of effort, then walked through with your engineering and finance stakeholders.
These figures are starting points, not quotes. Final pricing depends on the size and complexity of your estate and is fixed in writing at the end of Discovery & Scoping.
What this engagement does not cover
Named here rather than discovered later. This is the list that makes the fixed price hold when scope starts moving.
- implementation of recommendations
- contract negotiation with Google
- non-GCP spend
What this is based on
- Sole-tenant node and bring-your-own-licence positions modelled inside a 2,000+ VM migration programme.
- Google Cloud Professional Cloud Architect.
- Cluster bin-packing analysis on GKE running production workloads at high volume.
What buyers ask
We already applied the console recommendations. Is there anything left?
Usually a great deal. The recommendations engine works resource by resource and is genuinely useful at that level. It has no view of your commitment strategy, your licensing position, or your network topology.
The largest items in most models are commitment coverage, GKE request-to-usage gap, and egress caused by workload placement. None of the three appear as a console recommendation, because each requires looking at the estate as a whole.
How does the contingency option actually work?
15 to 20 percent of realized year-one savings, with the percentage and the measurement method agreed in writing before the work starts. Realized means implemented and visible in billing, not modelled.
It is only offered after read access to billing data. Committing to a savings percentage before seeing the spend would be a guess, and it would create an incentive to promise a number rather than find one.
Do you need write access to our billing account?
No. Billing account viewer and the BigQuery billing export are enough. No changes are made to your account, your commitments, or your resources during the analysis.
Will committing to CUDs lock us into decisions we might reverse?
That risk is real and it is modelled rather than glossed. Commitments are recommended against a usage baseline with a stated confidence, and where a platform decision in the near term would change the shape of the workload, the model says so and recommends a shorter term or a smaller commitment.
Over-committing on a workload about to be re-architected is one of the more expensive mistakes available in cost optimization, and it is easy to make when the model only looks backwards.
Can you implement the plan as well?
It is quoted separately after the analysis. Many clients do it themselves, because a large share of the items are configuration changes their own engineers can apply from the plan.
Where implementation touches platform structure — moving workloads to change egress, re-sizing a cluster’s node strategy — the Automation or GKE Platform packages are usually the better fit.
Questions about pricing, terms, and ownership across every engagement are on the FAQ.
Start with Discovery & Scoping
$4,500 fixed, 3–5 days. A current-state review, a gap analysis, a written scope, and a fixed quote for this engagement. Half the fee is credited against the work if you proceed within 60 days.