Week 2: Using Cloud Resources for Horizontal Scaling

DSAN 6000: Big Data and Cloud Computing
Fall 2026

Class Sessions
Author
Affiliation

Jeff Jacobs

Published

Tuesday, September 8, 2026

Open slides in new tab →

What Makes Data “Big”? Why Do We Need A “Cloud”?

Internet of Things (IoT)

Nearly 100 billion devices generating data:

  • Smart home devices (Alexa, Google Home, Apple HomePod)
  • Wearables (Apple Watch, Fitbit, Oura rings)
  • Vehicles (not just self-driving cars!)
  • Industrial sensors
  • Smart city infrastructure
  • Medical devices, remote patient monitoring

Slide Author: Amit Arora!

Smartphone Location Data

Slide Author: Amit Arora!

Real-World Examples

Netflix

  • Scale: 260+ million subscribers generating 100+ billion events/day
  • AI Use: Personalization, content recommendations, thumbnail generation
  • Stack: S3 → Spark → Iceberg → ML models → Real-time serving

Uber

  • Scale: 35+ million trips per day, petabytes of location data
  • AI Use: ETA prediction, surge pricing, driver-rider matching
  • Stack: Kafka → Spark Streaming → Feature Store → ML Platform

OpenAI

  • Scale: Trillions of tokens for training, millions of queries/day
  • AI Use: GPT models, DALL-E, embeddings
  • Stack: Distributed training → Vector DBs → Inference clusters

Slide Author: Amit Arora!

Data Architecture / Infrastructure Demo

How would you make X, the Former Bird App, scale to a billion users!?!

Figure 2-1 from Kleppmann and Riccomini (2026): A “standard” relational approach

“Back of the envelope” calculation demo

“The Cloud” = Renting Computers from a Big Company

(…That’s it, that’s the whole definition)

Our Toolbox From Now Until December 😎

Big Data Handling
You could technically run all these on your laptop!
…But in this class they’ll run on remote EC2

Cloud Services

  • Virtual Machines: EC2
  • Object Storage: S3
  • Elastic Map-Reduce: EMR
IDEs
These should be run on your laptop
…But just to connect to EC2 (edit/run remote files)!

AWS “Core”

  • Remotely-Accessible Computers with EC2
  • Remotely-Accessible Storage with S3

EC2: Elastic Cloud Compute

HW1:

  • Create your own remote EC2 computer instance
  • Use it to download (one-way) from remote storage bucket

S3: Simple Storage Service

HW2:

  • Create your own remote storage instance
  • Read and write to it from your EC2 instance

Athena x Glue: Query S3 Like a Database

HW5:

  • Auto-ingest data into an S3 bucket, along with metadata
  • Tell Athena the schema of this data using Glue
  • Run SQL queries on S3
    • You never manually make a DB!
    • You never manually set up a server to execute SQL queries or to save their results!

HW2: EC2 🤝 S3

  • In Class-Demo: Getting EC2 to talk to S3

References

Kleppmann, Martin, and Chris Riccomini. 2026. Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. 2nd edition. Santa Rosa, CA: O’Reilly.
Loukides, Mike. 2010. What Is Data Science? O’Reilly Media. June 2, 2010.
Mell, Peter, and Timothy Grance. 2011. The NIST Definition of Cloud Computing.” National Institute of Standards and Technology, Special Publication 800 (2011): 145.