Skip to content

Migration playbook: EC2 / virtual machines → Amazon EKS

EC2 to Amazon EKS Kubernetes migration

Containerize EC2 server fleets and move them to an autoscaling Amazon EKS cluster, so you stop paying for idle capacity sized for peaks.

Migration drivers

Why teams make this move

Driver 01

Idle capacity

Instances running around the clock at low average CPU to cover occasional peaks.

Driver 02

Manual patching and drift

Configuration drifting between servers that are maintained over SSH and ad-hoc scripts.

Driver 03

Slow scaling

Auto Scaling groups that take minutes to bring new instances into service.

Execution sequence

How the migration runs

Each phase ends with a check you can verify — data parity, error rates, latency — and the rollback path is agreed before any traffic moves.

  1. 01Phase

    Containerization and hardening

    Packaging applications as minimal, standardized container images.

  2. 02Phase

    EKS cluster and VPC in Terraform

    Provisioning a multi-AZ Amazon EKS cluster with Karpenter for node autoscaling.

  3. 03Phase

    Helm charts and GitOps

    Defining deployments declaratively and syncing them with Argo CD.

  4. 04Phase

    Traffic cutover

    Shifting traffic from the EC2 load balancers to ingresses managed by the AWS Load Balancer Controller.

Risk prevention

Pitfalls that derail this migration

Risk 01

Missing resource requests and limits

Not measuring CPU and memory needs, which leads to node memory pressure and OOM kills.

Risk 02

Local disk dependencies

Applications writing uploads to local VM disks instead of object storage (S3) or EFS.

Risk 03

No readiness or liveness probes

Traffic reaching new pods before their database connection pools are ready.

Before and after

What we measure

We take a baseline before any change and report the same numbers after cutover, from your own tools. They are the evidence of whether the migration worked — not figures promised in advance.

Compute cost
Monthly EC2 spend vs cluster spend for the same traffic
Utilization
CPU and memory requested vs used, per node
Scale-up time
Pending pod to ready, measured under load

Questions

Frequently asked migration questions

It packs several services onto fewer, right-sized nodes and lets you run interruptible workloads on spot capacity. The saving depends on how idle your current instances are, which we measure first.

Kubernetes marks the node unreachable and, after a timeout (five minutes by default, and tunable), reschedules its pods onto healthy nodes. Running several replicas across Availability Zones keeps the service up while that happens.

Rehearse the cutover before the real one

Tell us about your data volume, traffic and timeline. An engineer will reply within one business day to set up a call about the migration plan and its rollback path.