A SaaS Security Company Cuts Cloud Costs by 49% and Streamlines Multi-Cloud Kubernetes Operations with CloudPilot AI
CloudPilot AICloudPilot AIPublishedSep 30, 2026Read13 minA SaaS security company was already familiar with Kubernetes autoscaling. But operating six clusters across AWS and Google Cloud required more than adding and removing nodes.
The platform team also needed to right-size workloads, speed up image delivery, and support I/O-intensive services without stitching together multiple tools.
With CloudPilot AI, the company unified these capabilities on one platform. In the cluster featured in this case study, CPU requests decreased by 65%, memory requests by 54%, and costs by 49%.

Overview
| Industry | Cloud Environments | Region | CloudPilot AI Capabilities |
|---|---|---|---|
| Cybersecurity / SaaS security | Amazon Elastic Kubernetes Service (EKS), Google Kubernetes Engine (GKE) | North America | Image Accelerator, Local SSD ephemeral storage for GKE / EKS, PodMutation, Workload Autoscaler, Node Autoscaler |
CloudPilot AI manages 6 clusters with more than 1,000 cores. The environment spans 4 EKS clusters and 2 GKE clusters, all running Node Autoscaler (NA) and Workload Autoscaler (WA).
About the Customer
The company helps enterprises secure their SaaS applications, identities, and data. Its security posture management and threat detection capabilities help customers identify and respond to risks across their SaaS environments. The company's platform team first connected with CloudPilot AI at its KubeCon Hong Kong booth.
Results
CloudPilot AI helped the company improve the efficiency and reliability of its Kubernetes infrastructure by optimizing container image distribution, ephemeral storage I/O, and resource management. The cost results below cover the cluster featured in this case study and show the changes in resource requests and node costs.
| Area | Customer Need | Deployment and Results | Impact |
|---|---|---|---|
| Image distribution and startup | Reduce duplicate image pulls so new nodes can run workloads sooner | P2P data hit rate above 80%, with approximately 8.65 TiB of image data transferred through P2P | Reuse image content across nodes to cut remote registry downloads and streamline image delivery during scale-out and node replacement |
| Pod ephemeral storage I/O | Improve I/O performance for workloads suited to local storage | Dedicated Local SSD node pools configured across 3 clusters: 2 GKE and 1 EKS | Access local NVMe through standard Kubernetes ephemeral storage interfaces, giving I/O-sensitive workloads high-performance temporary storage |
| Workload placement | Place workloads previously spread across node pools into designated NodePools | Scheduling rules for selected workloads managed centrally through PodMutation | Keep new and recreated Pods in the intended node pools through consistent scheduling policies |
| Resource efficiency and cost | Reduce workload overprovisioning and optimize node costs | CPU requests down 65.17%, memory requests down 53.90%, and node costs down 48.63% in the featured cluster | WA and NA align application resource requests with node capacity, turning lower CPU and memory requests into lower node costs |
The Challenge: Reliable Service Startup as Clusters Scale
The company had no shortage of node autoscaling tools. The team had already used another vendor's autoscaler and explored Karpenter. But for a platform team operating both EKS and GKE, automatically adding and removing nodes addressed only part of the challenge. Keeping services running efficiently and reliably also required fast image delivery, high-performance ephemeral storage, and consistent workload placement.
They needed more than another node autoscaling tool. They needed a unified solution for application resource optimization, container image distribution, storage, and scheduling.
The first challenge was image distribution. During scale-out or node replacement, multiple nodes may need the same image content. Having each node pull that content independently from a remote registry leads to redundant downloads and additional data transfer costs. It also makes service startup more dependent on registry and network performance.
The second was ephemeral storage I/O. For services that use temporary files, disk caches, or container writable layers, sufficient compute capacity alone does not ensure that storage can meet their needs. The platform team needed a consistent way to provide high-performance local storage across GKE and EKS.
The third was workload placement. As the environment evolved, some workloads ended up spread across multiple node pools. The team wanted to consolidate selected workloads into designated node pools and ensure that recreated Pods followed the same scheduling rules, without managing those rules separately for each workload.
Together, these requirements pointed to a broader goal: moving beyond node autoscaling to gain greater control over workload deployment, startup, and runtime performance.
The Solution: Coordinated Image Distribution, Storage, Scheduling, and Autoscaling
Image Accelerator: Reducing Duplicate Image Downloads
CloudPilot AI's Image Accelerator allows nodes to share image content already stored in the cluster. When a new node needs an image, it first attempts to retrieve available content from participating peer nodes. If the content is unavailable or the P2P transfer fails, it falls back to the remote registry. The retrieved content is stored in the node's existing containerd image store and can be shared with other nodes. How Image Accelerator works
As of September 8, 2026, a customer-provided screenshot showed that one cluster had transferred 7,112.42 GiB of image data via P2P, achieved an 86% data hit rate, and completed 718,063 P2P image requests successfully. Reusing image content across nodes reduced repeated downloads from remote registries.
Single-Cluster Results: 7,112.42 GiB Received Through P2P

At the time of the screenshot, 7 nodes were participating in image distribution. The cluster had transferred a cumulative 7,112.42 GiB (approximately 6.95 TiB) of image data via P2P, with an 86% data hit rate and 718,063 successful P2P image requests.
P2P transfers accounted for approximately 86% of all image data delivered to nodes.
Note: Cumulative metrics include contributions from nodes that have since been removed. The totals therefore differ from the sum of the per-node metrics for the 7 nodes shown in the screenshot.
Image Reuse Across Clusters
| Cluster | Cumulative Data Received Through P2P | P2P Data Hit Rate | Cumulative Successful P2P Image Requests |
|---|---|---|---|
| EKS-A | 7,009.96 GiB | 86% | 711,572 |
| EKS-B | 1,712.60 GiB | 89% | 138,910 |
| EKS-C | 83.68 GiB | 79% | 9,532 |
| GKE-A | 32.60 GiB | 71% | 862 |
| GKE-B | 16.27 GiB | 82% | 9,968 |
| Total | 8,855.11 GiB (about 8.65 TiB) | 870,844 |
Supporting Faster Startup and Reliable Image Delivery
- Fewer duplicate remote downloads: The single-cluster total of 7,112.42 GiB received through P2P and the historical total of approximately 8.65 TiB across five clusters show that image content was reused across nodes.
- Faster image preparation: During scale-out or node replacement, new nodes can retrieve image content already available in the cluster. This reduces repeated requests to remote registries and the time spent pulling images during Pod startup.
- Alternative image sources: Peer-to-peer distribution, with fallback to a remote registry, provides alternative sources for image content during scale-out, node replacement, and recovery.
Local SSD Ephemeral Storage: Up to 55× Random Read Performance in GKE Benchmarks
CloudPilot AI's Local SSD ephemeral storage feature uses local NVMe SSDs on compatible GKE and EKS nodes to provide Kubernetes ephemeral storage. This gives workloads a local storage path for high-throughput, low-latency reads and writes of temporary data. Platform teams can configure node pools to use this storage, while applications access it through standard Kubernetes ephemeral storage interfaces.
For disk-backed emptyDir volumes, container writable layers, runtime data, images, logs, and other supported uses, applications can access the Local SSD-backed node filesystem directly. Teams do not need to manage raw disks or handle disk formatting and mounting themselves. They also do not need hostPath volumes or privileged containers to access the local storage.
The company has configured dedicated Local SSD node pools in 2 GKE clusters and 1 EKS cluster.
CloudPilot AI's GKE Benchmark Results
In CloudPilot AI's GKE Standard benchmark tests, each fio operation was run 3 times using a Pod's emptyDir volume and direct I/O. The first test compared two otherwise identical n2-standard-4 nodes. The baseline node used a 50 GB pd-balanced boot disk. The other node had the same configuration plus a 375 GB NVMe Local SSD.
| fio Test | Boot Disk Baseline | Local SSD | Performance Gain |
|---|---|---|---|
| Sequential write, 1 MiB | 154.26 MiB/s | 392.12 MiB/s | 2.54× |
| Sequential read, 1 MiB | 155.36 MiB/s | 704.60 MiB/s | 4.54× |
| Random write, 4 KiB | 3,241 IOPS | 99,971 IOPS | 30.84× |
| Random read, 4 KiB | 3,255 IOPS | 179,317 IOPS | 55.09× |
A second GKE test compared c3-standard-4 with c3-standard-4-lssd, which includes a 375 GB Local SSD. Relative to the baseline, the Local SSD configuration delivered 2.54× the sequential write performance, 4.57× the sequential read performance, 30.81× the random write performance, and 55.25× the random read performance. The two tests evaluated both separately attached Local SSDs and instance types that include Local SSDs, showing particularly strong results for temporary data workloads with intensive random I/O.
Platform teams can therefore size node pools for both compute demand and the ephemeral storage performance required by I/O-sensitive services.
PodMutation: Consistent Scheduling for Selected Workloads
The company wanted to place workloads previously spread across different node pools into designated NodePools. CloudPilot AI's PodMutation centralizes the scheduling policies for these workloads and applies the relevant selectors, tolerations, or affinity rules when Pods are created.
The platform team can update scheduling constraints for selected workloads without changing every Deployment or StatefulSet individually. When Pods are recreated during a controlled rollout or normal operation, these rules ensure that they are scheduled onto the intended node pools and continue to apply to future replacements.
These policies keep workload placement consistent across deployments and Pod replacements, making platform changes easier to manage and track.
Note: Placing workloads in a designated node pool does not mean placing all replicas on a single node. Existing replica-distribution and availability-zone constraints still apply. PodMutation does not relocate running Pods; migration requires controlled Pod recreation or a rolling update.
Workload Autoscaler and Node Autoscaler: Nearly 49% Lower Hourly Node Costs in One Cluster
One cluster managed by CloudPilot AI illustrates how application and node optimization work together to reduce overprovisioning and lower node costs.
| Metric | Before Optimization | After Optimization | Change |
|---|---|---|---|
| CPU request | 197.51 Cores/h | 68.79 Cores/h | −65.17% |
| Memory request | 674.73 GiB | 311.06 GiB | −53.90% |
| Total hourly node cost | $9.470/hour 2026-08-15 06:00 | $4.865/hour 2026-09-08 06:00 | −48.63% $4.605 less per hour |
In this cluster, CPU requests fell by nearly two-thirds, memory requests fell by more than half, and hourly node costs dropped by nearly 49% between the two snapshots.
Workload Autoscaler: Right-Sizing Applications
Workload Autoscaler analyzes application resource requirements and adjusts CPU and memory requests to reduce overprovisioning. In this cluster, CPU requests decreased from 197.51 to 68.79 Cores/h, and memory requests decreased from 674.73 to 311.06 GiB, enabling further optimization of node capacity.

CPU: Original requests, current requests, and actual usage

Memory: Original requests, current requests, and actual usage

Node Autoscaler: Translating Application Right-Sizing into Lower Node Costs
Once application requests more closely reflect workload demand, Node Autoscaler manages the capacity of the nodes that run those workloads. Coordinating application and node optimization allows reductions in resource requests to translate into lower node costs.

The screenshots show that total hourly node costs fell from $9.470 at 06:00 on August 15 (Spot: $1.363; On-Demand: $8.107) to $4.865 at 06:00 on September 8 (Spot: $1.663; On-Demand: $3.202)—a 48.63% reduction between the two snapshots.
Before optimization: Total node cost of $9.470 per hour

After optimization: Total node cost of $4.865 per hour

Beyond Node Autoscaling: A Better Fit for Service Operations
These improvements addressed the company's core requirements when evaluating solutions: reducing node costs while ensuring that application resources, image distribution, and ephemeral storage worked together to support its services. For the platform team, choosing CloudPilot AI meant bringing these capabilities together in one platform, without having to select, integrate, and maintain separate components around a node autoscaler. The Head of Infrastructure explained the reasoning behind the decision:
"We had already explored Karpenter. It handles node autoscaling, but our needs went further. Karpenter itself does not include a Workload Autoscaler to continuously tune workload CPU and memory requests based on actual usage, nor does it provide P2P image acceleration. We would still need to select, integrate, and maintain separate tools for those capabilities.
We chose CloudPilot AI because it brings coordinated application and node optimization, image acceleration, and Local SSD ephemeral storage together in one platform. We wanted to reduce not only resource waste, but also the delays caused by repeated image pulls during scale-out and the burden of maintaining multiple infrastructure components. What matters most to us is getting services ready faster, giving I/O-sensitive workloads suitable storage, and keeping reliability our top priority."
— Henry, Head of Infrastructure
With CloudPilot AI, the team expanded its approach beyond node autoscaling to include workload optimization, image delivery, and storage. Lower costs were one outcome, but the team also gained broader operational capabilities and greater control over its infrastructure.
Reliability Comes First
For the company, cost optimization must account for service requirements. Node capacity can change, but image preparation, ephemeral storage, and workload scheduling need to be considered together.
CloudPilot AI addresses these requirements in a single infrastructure management platform rather than focusing only on overprovisioning. For a SaaS security company, reliability comes first: every efficiency gain must preserve service performance and availability.
What's Next
In the next phase, the teams will extend optimization to batch workloads and bring more clusters under CloudPilot AI management, expanding the partnership beyond the services already on the platform. Two priorities will guide this work:
- Optimize Job NodePools for batch workloads. Adjust resource configuration and node management for node pools running batch job Pods, based on the customer's batch-processing requirements. The objective is to reduce idle capacity and resource waste and improve utilization.
- Extend CloudPilot AI to additional clusters. Bring more clusters under CloudPilot AI management so additional workloads can benefit from automated optimization and infrastructure management.
About CloudPilot AI
CloudPilot AI is a cloud infrastructure company based in San Francisco, specializing in autonomous Kubernetes optimization for enterprises operating at scale.
Its mission: Autoscaling Kubernetes for the most demanding teams.
The platform continuously right-sizes workloads, scales node capacity, accelerates container image delivery, automates workload placement, and enables high-performance local storage for I/O-intensive applications. Together, these capabilities help SRE and platform teams reduce cloud costs, improve application performance, and maintain reliability across multi-cloud environments—without manual tuning.
Trusted by more than 100 enterprises worldwide, it helps customers achieve average savings of 67%.
Product References

Related articles
How a Global Gaming Company Cut EKS Node Costs by 47% with Workload, Node, and Spot Optimization
A global gaming company tripled CPU utilization and cut EKS node costs by 47%—without changing application code—by coordinating CloudPilot AI's Workload Autoscaler, Node Autoscaler, and Spot Automation.
CloudPilot AI·9 minHow a Generative AI Company Cut AWS EKS Costs by 55% with Full-Stack Intelligent Autoscaling
A leading generative AI company reduced its real AWS EKS costs by more than 50% and sharply improved cluster resilience—without changing a single line of code. See how CloudPilot AI's full-stack autoscaling rightsized workloads and reshaped nodes on top of EKS Auto Mode.
CloudPilot AI·8 min
Netvue Achieves 52% Reduction in GPU Costs using Automation
Discover how Netvue reduced GPU inference costs by 52% using CloudPilot AI. Learn how Kubernetes-based scheduling, Spot instance automation, and multi-cloud flexibility helped build a scalable, high-performance AI infrastructure.
CloudPilot AI·7 minThe Kubernetes Autoscaling Brief
Once a month: what we see in the Spot market, one real cluster torn down, and what shipped.
No spam. Unsubscribe anytime.