First Cohort
Lessons and interview talk-throughs — pick up a week, or walk a scenario out loud.
Lessons & interviews
Why a stateless model can still hold a conversation — context windows, short-term memory, summarization, and long-term memory.
How LLMs generate text one token at a time — next-token prediction, Prefill, Decode, and KV Cache.
How pre-training teaches language, knowledge, and next-token prediction from enormous amounts of data — before you ever start chatting.
Say Chat, Generative, Pre-trained, and Transformer out loud — ten interview-style questions, with no answers on the page.
What System Design means — requirements, scalability, availability, DNS, and load balancers — before choosing a specific technology.
Five questions on variables in memory, typing, Big-O, and how a Python list works — with no answers on the page.
How Chrome, Python, PyTorch, and an LLM still ask the kernel for CPU, memory, disk, and GPU — user space, kernel space, and system calls.
Which task gets the CPU, for how long, and why R, S, D, and Z still show up in Linux interviews and GPU-serving incidents.
Virtual vs physical memory, pages, the MMU, page cache, swap, OOM — and why RAM and VRAM are different problems in GenAI.
How a read() travels through VFS, page cache, the block layer, and the NVMe driver — and why model loading is an I/O problem before it is a GPU problem.
Follow one packet from a socket through TCP, IP, routing, Netfilter, and the NIC — and why a healthy GPU can still look slow.
Four first-pass commands: which process is using CPU, memory, disk I/O, and the network.
Start with one application and one database, then ask what breaks next: load balancers, database scaling, replication, and caching.
GPUs were built for games, then became the foundation of modern AI — CPU vs GPU, SMs, Tensor Cores, VRAM, bandwidth, NVLink, and cooling.
A practical memory-sizing walkthrough for a 70B model — weights, KV cache, runtime overhead, and why the total can reach about 1 TB.
How a cluster discovers GPUs, exposes nvidia.com/gpu, schedules them, and lets a container actually use the card.
Gateway, router, and GPU workers — prefill, decode, KV cache, continuous batching, and how a request is streamed back.
336 billion transistors, two dies, eight HBM4 stacks, 896 Tensor Cores — and why even NVIDIA’s most powerful chip spends most of its time waiting.
Beginner-friendly patterns for phone numbers, emails, URLs, and Apache access logs — then extract them in Python.
Check paths, file sizes, and search for files with os and pathlib, then run Linux commands from Python with subprocess.
Detect the operating system, accept a hidden password, and read command-line arguments with platform, getpass, and sys. Run remote Linux commands over SSH with Paramiko.
Sessions, resources, and services — and how to confirm which AWS account and identity a script is using before automation runs.
One AWS entry point for many foundation models — token-based billing, InvokeModel vs Converse, the request flow, a first boto3 call, and what to monitor in production.
A beginner's guide to serving an LLM with vLLM on Amazon EKS — build the cluster, pre-pull the image, run Qwen3-0.6B behind an OpenAI-compatible API, test it, and learn from what went wrong.
Build a highly available AWS web app step by step — Amplify, compute choices, APIs, Aurora, RDS Proxy, ElastiCache, queues, and microservices as demand grows.
Verbal SRE / Cloud scenarios by company — structure first, no AI coach.