The Rightsizing Recommendation That Rots in a Backlog

Blog · FinOps

The Rightsizing Recommendation That Rots in a Backlog

By Pintu Sahu7 min read

Short answer

FinOps tools surface waste and hand engineers a to-do — rightsize this instance, reclaim that volume, buy this commitment — then report "realized savings" that only materialize if a human opens the console and acts.

Your cost platform found the waste. It flagged the oversized instance, the orphaned volume, the discount commitment you should have bought last quarter. Then it did the one thing every cost tool does with a finding: it handed it to an engineer as a to-do — and booked the line item as "potential savings." Those savings were never yours. They belonged to a ticket nobody prioritized.

Start with what these tools genuinely earned, because the wedge only lands if the strengths are conceded honestly. A decade ago, cloud spend arrived as a monthly surprise nobody could decompose. The cost-visibility and allocation platforms fixed that. They ingest billing and usage, allocate nearly every dollar to an owner, a team, a product — showback and chargeback that finally made cloud economics legible to the people accountable for them. They compute unit cost, so spend can be read against the business it supports. They detect anomalies before the invoice does. And they generate rightsizing and commitment recommendations that are, more often than not, correct.

Some have gone further, and deserve real credit for it. A few govern cost before deploy — inside the delivery pipeline, at infrastructure-as-code plan time — so a change that blows the budget gets caught in the pull request instead of on the invoice. That shift-left posture is the strongest governance story on the market today. Idle-resource shutdown for non-production, cost-in-the-pull-request, professionalized cloud finance as an operating model — these are advances, not marketing. This post does not dispute any of it.

Here is the uncomfortable part. Every capability above terminates at the same place: a recommendation. The platform observes spend and advises a change. The change itself is implemented by an engineer, in the console or in a separate automation, and the tool tracks the result afterward. That is why the category's own headline metric — "realized savings" — carries a quiet dependency clause: it only materializes if a human opens the console and acts on the finding.

Look at the gap directionally and it is not small. Studies across the category routinely report that a large share of identified savings — frequently the majority — never gets captured in the billing period it was found. The recommendation was right. The instance stayed oversized anyway. The gap isn't an execution failure by tired engineers; it is the architecture working as designed. Three properties make the leak structural:

  • The recommendation has no hands. It is advice rendered next to the estate, not an action taken on it. Someone has to translate the finding into a change, schedule it against a change window, and own the blast radius.
  • The savings number is a forecast, not a fact. "You could save X" is a projection contingent on an action that hasn't happened. It is booked as value the moment it is surfaced, and quietly written down when the ticket ages out.
  • Governance is a rear-view mirror. A dashboard, an ITSM ticket, and a budget alert that fires after the money is already committed. It reports the overspend faster; it does not prevent it.

The vendors know the gap exists, so execution has started to appear — and where it appears, it is deliberately narrow. One pattern is a commitment-only autopilot that automates discount purchases while stating plainly that it "does not make changes to your infrastructure." Another is auto-stop confined to idle, non-production resources. A third ships a separate optimizer where, at the decisive moment, a human still "confidently takes the rightsizing action."

Notice what these have in common. The paths that get automated are the ones with the least blast radius: buying a reservation touches a billing construct, not a workload; stopping an idle dev box breaks nothing anyone is using. The instant execution would touch a running production resource under load, the system reverts to a recommendation and hands the risk back to a person. That is not a roadmap gap to be filled next quarter — it is structural. These platforms sit beside the estate. They were architected to read it and advise on it. Acting on it, under governance, with a way to undo the act, was never in the design.

Entroid closes the gap by moving the fix from a ticket to an action the system itself performs — under governance, and reversibly. This is an architectural property of the fabric, not a demo claim. On a Composable Process Fabric, remediation runs through the same five primitives as everything else: Atomic Agents propose and carry out the change with human-in-the-loop as a first-class control, not a bolted-on approval step; a Deterministic Workflow enforces the budget, authority, and policy check inline before the change is allowed to proceed; Connectors — the only primitive that touches an external system — actually perform the reclaim, the resize, or the commitment change on the cloud estate; and the Operations Toolkits (cluster utilization, congestion, AI recommendations) feed the loop that decides what to act on. The Semantic Ontology already models owner, cost center, and budget, so the action knows whose money and whose authority it is spending.

BESIDE THE CLOUD — observe & advise Ingest billing + usage Allocate · showback Recommend action Budget ALERT + ITSM ticket ← human acts Spend already committed ON THE FABRIC — gate & act Console API IaC INLINE GATE budget · authority · policy Resource created Reclaim / resize reversible Immutable per-action audit what · by whose authority · how to roll back

Because the change executes inside one runtime, the loop closes on itself. The immutable per-action audit records not a projected saving but the actual event: what was done, by whose delegated authority, and — critically — how to reverse it. Reversibility is a design property, not a promise, which is precisely what lets governed remediation reach the running production resource that the beside-the-cloud tools have to leave alone. To be explicit, these are architectural properties of how the fabric is built; any specific reclaim or resize figure would be illustrative, not a delivered ES result.

The same architecture answers the prevention side, and here the shift-left vendors deserve a precise critique rather than a dismissive one. Governing cost in the pipeline is real governance — but it governs exactly one way in. The console click, the raw cloud API call, and the autoscaler all provision resources without ever passing through the pipeline. The tell is that these tools ship a "zero-drift" cleanup loop at all: a loop that continuously hunts down and removes resources that appeared outside the gate is a quiet admission that the gate gets bypassed. And even on the path it does cover, the check is a dollar estimate — "this will cost roughly X" — not a delegated-authority decision about whether this owner may commit this budget at all.

ES binds the gate to the act of creation itself. A budget, authority, and policy check runs inline in a Deterministic Workflow, so an over-budget or non-compliant resource cannot be created on any path — console, API, or infrastructure-as-code all route through the same control. And it denies on authority, not just on a forecast: the ontology knows the owner, the cost center, and the budget, so the answer to "may this be provisioned" is a governed decision with an audit record, not a cost warning a human is free to click past. The overspend becomes un-creatable, and the fix becomes self-executing — a control plane at the moment of action instead of a faster rear-view mirror. This runs over your existing estate through governed Connectors; it is a different control model on the infrastructure you already have, not a rip-and-replace.

A recommendation is a hope with a dollar sign. Governance is the resource that could not be created — and the fix that already ran.

See what this looks like for your enterprise.

Not a demo. A strategic conversation about how your enterprise could operate
when every process runs on one governed fabric.

Start the Conversation