Skill 24 · Architect For Startups
Subchapter 24.12
references/ec2.mdMarkdown5 KBView on GitHub
Almost never as your first choice. Start with Lambda or ECS Fargate. EC2 makes sense for startups only when:
The hidden cost: EC2 requires patching, AMI management, monitoring agents, and capacity planning. At a 3-person startup, that’s 10-20% of an engineer’s time — your scarcest resource.
Running instances 24/7 in dev/staging: A t3.medium costs ~$30/month. Three dev environments left running = $90/month doing nothing nights/weekends. Use Auto Scaling scheduled actions or Lambda-triggered stop/start. Savings: 65% on non-prod instances.
Elastic IPs not attached to instances: $3.60/month per unused EIP. Teams allocate them “for later” and forget. Check monthly.
gp2 volumes still in use: gp2 costs more than gp3 and performs worse at small sizes (gp2 scales IOPS with size; gp3 gives 3000 IOPS baseline regardless). Convert all gp2 → gp3 immediately, it’s free and non-disruptive.
Savings Plans purchased too early: Don’t buy Compute Savings Plans until you have 3+ months of stable EC2 usage data. Startups pivot — a 1-year commitment on instance types you abandon in 3 months is wasted money.
Data transfer between AZs: $0.01/GB each way. A chatty microservices architecture across AZs can accumulate $100s/month in cross-AZ transfer that doesn’t show up obviously in Cost Explorer. Keep tightly-coupled services in the same AZ or use ECS/EKS service mesh for efficient routing.
t3.small or t3a.small ($15/month) — burstable is fine for dev/stagingcapacity-optimized for batch/ML — diversify across 10+ instance typesA single EC2 instance is a valid architecture. For internal tools, admin panels, or low-traffic services that don’t justify container orchestration — one t3.small with a Docker Compose setup and an AMI-based backup strategy works fine. It’s not “production best practice” but it’s pragmatic until revenue justifies HA.
Don’t build for multi-AZ until you have paying customers who’d notice. Multi-AZ doubles your compute baseline cost. If your SLA is informal and your recovery plan is “redeploy from AMI in 10 minutes,” that’s fine at pre-PMF.
Spot instances for production stateless workloads is fine. The standard advice is “never use Spot in production.” For startups: if your service handles graceful shutdown and you have On-Demand fallback in your ASG, the 70% savings funds an extra engineer-month every few months.
| Signal | Why EC2 |
|---|---|
| Monthly Fargate spend > $10K with >80% utilization | EC2 + Savings Plans is 30-50% cheaper |
| Need GPU instances | No Fargate GPU support |
| Need instance store NVMe for caching | Fargate has no local storage option |
| Running workloads requiring privileged containers | Fargate doesn’t support privileged mode |
p3.2xlarge costs ~$2,200/month. Use Spot for training and be deliberate about GPU hours.