01 / 05

Discuss spot instances and spot fleet

Difficulty: 6/10
cost optimization, capacity management, interruption handling

Spot Instances are spare EC2 compute capacity available at up to 90% discount compared to On-Demand prices, while a Spot Fleet is a collection of Spot Instances (and optionally On-Demand Instances) that launches and maintains capacity from multiple pools to optimize cost, performance, or capacity.

Spot Instances leverage AWS's unused EC2 capacity, offering significant cost savings (up to 90% off On-Demand prices) in exchange for potential interruption when AWS needs the capacity back. Spot Instances are ideal for fault-tolerant, stateless, or flexible workloads such as batch processing, big data analytics, CI/CD pipelines, containerized workloads, and development/test environments. Unlike On-Demand or Reserved Instances, Spot capacity can be reclaimed by AWS with a two-minute interruption notice.

A Spot Fleet is a collection of Spot Instances (and optionally On-Demand Instances) that launches and maintains a target capacity from multiple Spot Instance pools, each defined by instance type, Availability Zone, and price [citation:2]. Spot Fleet automatically requests capacity from the most cost-effective pools based on your allocation strategy, and can automatically replace instances that are interrupted or become unhealthy [citation:2].

Scenario Questions

0-2 years experience

  1. 1You need to run a batch data‑processing job that can tolerate interruptions and you want to minimize cost. How would you use Spot Instances and Spot Fleet to achieve this?
  2. 2If a Spot Instance in your fleet gets terminated, what steps would you take to ensure your job continues without data loss?
  3. 3What happens if your Spot Fleet's target capacity can't be met because spot prices exceed your bid?

2-5 years experience

  1. 1Your service currently runs on on‑demand instances. Management asks you to cut compute costs by 30% using Spot Fleet. What trade‑offs would you evaluate and how would you migrate?
  2. 2During a deployment you notice that Spot Fleet is not achieving the desired capacity and some tasks are failing. How would you debug the issue?
  3. 3Explain how you would configure a Spot Fleet to balance cost versus availability across multiple instance types and AZs.

5-8 years experience

  1. 1Design a highly available, auto‑scaling web tier that uses Spot Fleet for the majority of capacity but falls back to on‑demand when spot capacity is insufficient. What components and policies would you put in place?
  2. 2At scale, how would you handle spot instance interruptions to avoid cascading failures in a distributed system?
  3. 3Discuss the performance and cost implications of using the 'capacity‑optimized' allocation strategy versus a simple bid‑price strategy in a large‑scale fleet.

8+ years experience

  1. 1Your organization wants to migrate a legacy monolith to a microservices architecture on AWS, and you are tasked with defining the compute strategy. How would you incorporate Spot Fleet across multiple teams while ensuring SLA compliance and operational consistency?
  2. 2What governance and monitoring framework would you establish to manage Spot Fleet usage, cost allocation, and risk across the entire enterprise?
  3. 3If a new AWS region launches with limited spot capacity, how would you redesign your global deployment to maintain resilience without over‑provisioning on‑demand instances?

Follow-up Questions

  • How would you monitor spot instance interruptions in production?
  • What metrics would you use to decide when to fall back to on‑demand instances?
  • Can you give an example where using Spot Fleet would be a poor choice?
Share

Share via WhatsApp, X, Facebook, LinkedIn or copy link. Open Graph preview enabled.