Week 2: Using Cloud Resources for Horizontal Scaling
DSAN 6000: Big Data and Cloud Computing
Fall 2026
Class Sessions
What Makes Data “Big”? Why Do We Need A “Cloud”?
Internet of Things (IoT)
Nearly 100 billion devices generating data:
- Smart home devices (Alexa, Google Home, Apple HomePod)
- Wearables (Apple Watch, Fitbit, Oura rings)
- Vehicles (not just self-driving cars!)
- Industrial sensors
- Smart city infrastructure
- Medical devices, remote patient monitoring

Slide Author: Amit Arora!
Smartphone Location Data

Slide Author: Amit Arora!
Real-World Examples
Netflix
- Scale: 260+ million subscribers generating 100+ billion events/day
- AI Use: Personalization, content recommendations, thumbnail generation
- Stack: S3 → Spark → Iceberg → ML models → Real-time serving
Uber
- Scale: 35+ million trips per day, petabytes of location data
- AI Use: ETA prediction, surge pricing, driver-rider matching
- Stack: Kafka → Spark Streaming → Feature Store → ML Platform
OpenAI
- Scale: Trillions of tokens for training, millions of queries/day
- AI Use: GPT models, DALL-E, embeddings
- Stack: Distributed training → Vector DBs → Inference clusters
Slide Author: Amit Arora!
Data Architecture / Infrastructure Demo
“The Cloud” = Renting Computers from a Big Company
(…That’s it, that’s the whole definition)
Our Toolbox From Now Until December 😎
Big Data Handling
You could technically run all these on your laptop!
…But in this class they’ll run on remote EC2
- Python
- Apache Arrow
- Apache Hadoop \(\leadsto\) Apache Spark
Cloud Services
AWS “Core”
- Remotely-Accessible Computers with EC2
- Remotely-Accessible Storage with S3
EC2: Elastic Cloud Compute
HW1:
- Create your own remote EC2 computer instance
- Use it to download (one-way) from remote storage bucket
S3: Simple Storage Service
HW2:
- Create your own remote storage instance
- Read and write to it from your EC2 instance
Athena x Glue: Query S3 Like a Database
HW5:
- Auto-ingest data into an S3 bucket, along with metadata
- Tell Athena the schema of this data using Glue
- Run SQL queries on S3
- You never manually make a DB!
- You never manually set up a server to execute SQL queries or to save their results!
HW2: EC2 🤝 S3
- In Class-Demo: Getting EC2 to talk to S3
References
Kleppmann, Martin, and Chris Riccomini. 2026. Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems. 2nd edition. Santa Rosa, CA: O’Reilly.
Loukides, Mike. 2010. “What Is Data Science?” O’Reilly Media. June 2, 2010.
Mell, Peter, and Timothy Grance. 2011. “The NIST Definition of Cloud Computing.” National Institute of Standards and Technology, Special Publication 800 (2011): 145.





