First Cohort

Lessons and interview talk-throughs — pick up a week, or walk a scenario out loud.

16 lessons Interview AWS · OpenAI · NVIDIA

Continue learning

Start with Week 1 — understanding the ‘Chat’ in ChatGPT: context windows, memory, and why a stateless model can still hold a conversation.

Open Week 1 →
Course path

Lessons & interviews

Week 1 · Chat
Understanding the ‘Chat’ in ChatGPT

Why a stateless model can still hold a conversation — context windows, short-term memory, summarization, and long-term memory.

Read →
Week 1 · Generative
What Does the “G” in ChatGPT Mean?

How LLMs generate text one token at a time — next-token prediction, Prefill, Decode, and KV Cache.

Read →
Week 1 · Pre-trained
What Does the “P” in ChatGPT Mean?

How pre-training teaches language, knowledge, and next-token prediction from enormous amounts of data — before you ever start chatting.

Read →
Week 1 · Questions
Ten Questions on ChatGPT

Say Chat, Generative, Pre-trained, and Transformer out loud — ten interview-style questions, with no answers on the page.

Try →
Week 1 · System Design
System Design - Part 1

What System Design means — requirements, scalability, availability, DNS, and load balancers — before choosing a specific technology.

Read →
Week 1 · Python
Python Basics - Part 1

Five questions on variables in memory, typing, Big-O, and how a Python list works — with no answers on the page.

Try →
Week 2 · Kernel
Understanding user space vs kernel space in the Age of GenAI

How Chrome, Python, PyTorch, and an LLM still ask the kernel for CPU, memory, disk, and GPU — user space, kernel space, and system calls.

Read →
Week 2 · Scheduler
Process Scheduler

Which task gets the CPU, for how long, and why R, S, D, and Z still show up in Linux interviews and GPU-serving incidents.

Read →
Week 2 · Memory
Linux memory management

Virtual vs physical memory, pages, the MMU, page cache, swap, OOM — and why RAM and VRAM are different problems in GenAI.

Read →
Week 2 · I/O
Linux I/O management

How a read() travels through VFS, page cache, the block layer, and the NVMe driver — and why model loading is an I/O problem before it is a GPU problem.

Read →
Week 2 · Network
Linux networking

Follow one packet from a socket through TCP, IP, routing, Netfilter, and the NIC — and why a healthy GPU can still look slow.

Read →
Week 2 · Processes
Quick Linux Process Troubleshooting

Four first-pass commands: which process is using CPU, memory, disk I/O, and the network.

Read →
Week 2 · System Design
System Design from Scratch: Scaling from 100 to 100,000 Users

Start with one application and one database, then ask what breaks next: load balancers, database scaling, replication, and caching.

Read →
Week 3 · GPU
Introduction to GPU

GPUs were built for games, then became the foundation of modern AI — CPU vs GPU, SMs, Tensor Cores, VRAM, bandwidth, NVLink, and cooling.

Read →
Week 3 · Memory
Understanding GPU Memory Requirements for LLM Inference

A practical memory-sizing walkthrough for a 70B model — weights, KV cache, runtime overhead, and why the total can reach about 1 TB.

Read →
Week 3 · Kubernetes
Kubernetes on GPU

How a cluster discovers GPUs, exposes nvidia.com/gpu, schedules them, and lets a container actually use the card.

Read →
Week 3 · Designing an AI Inference System
Designing an AI Inference System

Gateway, router, and GPU workers — prefill, decode, KV cache, continuous batching, and how a request is streamed back.

Read →
Week 3 · Vera Rubin
Inside NVIDIA Vera Rubin

336 billion transistors, two dies, eight HBM4 stacks, 896 Tensor Cores — and why even NVIDIA’s most powerful chip spends most of its time waiting.

Read →
Week 3 · Regular Expressions
Regular Expressions for Python & DevOps

Beginner-friendly patterns for phone numbers, emails, URLs, and Apache access logs — then extract them in Python.

Read →
Week 3 · Python os and subprocess
Python os and subprocess module

Check paths, file sizes, and search for files with os and pathlib, then run Linux commands from Python with subprocess.

Read →
Week 3 · sys, platform, getpass and paramiko
Python for DevOps using sys, platform, getpass and paramiko module

Detect the operating system, accept a hidden password, and read command-line arguments with platform, getpass, and sys. Run remote Linux commands over SSH with Paramiko.

Read →
Week 3 · Getting Started with Boto3
Python for DevOps: Getting Started with Boto3

Sessions, resources, and services — and how to confirm which AWS account and identity a script is using before automation runs.

Read →
Week 4 · Amazon Bedrock for Beginners
Amazon Bedrock for Beginners: One AWS Service, Many AI Models

One AWS entry point for many foundation models — token-based billing, InvokeModel vs Converse, the request flow, a first boto3 call, and what to monitor in production.

Read →
Week 4 Bedrock Mantle + Lambda + API Gateway Build a GenAI Q&A API on AWS
Week 4 · vLLM on Amazon EKS
Run Your Own AI Model on Kubernetes

A beginner's guide to serving an LLM with vLLM on Amazon EKS — build the cluster, pre-pull the image, run Qwen3-0.6B behind an OpenAI-compatible API, test it, and learn from what went wrong.

Read →
Week 4 · System Design
System Design: Highly Available AWS Architecture

Build a highly available AWS web app step by step — Amplify, compute choices, APIs, Aurora, RDS Proxy, ElastiCache, queues, and microservices as demand grows.

Read →
Interview
Interview talk-through

Verbal SRE / Cloud scenarios by company — structure first, no AI coach.

Practice →