A YouTube-first AI/ML engineering roadmap: from Python and the math foundations, through data handling with NumPy/Pandas, classical machine learning and evaluation, feature engineering, deep learning with PyTorch, computer vision and NLP, transformers, LLMs, embeddings and vector databases, RAG, AI agents, FastAPI serving, MLOps with MLflow, Docker, deployment and production AI — ending in a serious end-to-end capstone.
Track your progress, earn XP, and pick up where you left off.
Milestone 01 (Python + Data Foundations) begins here. Get comfortable writing Python without leaning on tutorials: syntax, data structures, functions, OOP, files, errors and clean code.
Python Foundations
Syntax, variables, types, conditionals, loops and functions.
~12h · 5 resources
Collections & Comprehensions
Lists, tuples, dicts, sets and comprehensions.
~8h · 8 resources
OOP, Modules, venv & Errors
Classes, packages, virtual environments and exception handling.
~12h · 10 resources
Only the linear algebra, calculus, probability and statistics needed to understand ML — no more.
Linear Algebra
Vectors, matrices, dot products, transpose, eigen-things and dimensionality.
~8h · 3 resources
Calculus & Gradients
Derivatives, partial derivatives, gradients and gradient descent.
~6h · 3 resources
Probability & Statistics
Probability, Bayes, distributions, expectation, variance and hypothesis testing.
~10h · 2 resources
Arrays, broadcasting and vectorization — the numerical backbone of ML.
Arrays, Indexing & Broadcasting
Dimensions, shapes, slicing, broadcasting and vectorized math.
~8h · 2 resources
Load, clean, transform and visualize data — the daily work of an ML engineer.
Series, DataFrames & Cleaning
Loading CSV/JSON, missing values, duplicates, filtering and sorting.
~10h · 2 resources
Grouping, Merging & Visualization
Groupby, aggregation, joins, pivot tables, Matplotlib and Seaborn.
~10h · 4 resources
Turn raw data into model-ready features — and learn to avoid data leakage.
Encoding, Scaling & Missing Values
Categorical/numerical handling, encoding, normalization and standardization.
~8h · 2 resources
Pipelines & Avoiding Data Leakage
Feature selection/extraction, outliers, and fitting preprocessing on training data only.
~8h · 2 resources
Milestone 02 (First ML Model). Supervised vs unsupervised, the train/val/test split, and the bias–variance tradeoff.
ML Concepts & Workflow
What ML is, learning types, features/labels, parameters/hyperparameters.
~8h · 4 resources
Overfitting, Regularization & Validation
Train/test split, cross-validation, overfitting/underfitting, bias/variance and regularization.
~8h · 2 resources
Regression and classification algorithms with intuition, strengths and limitations.
Regression Algorithms
Linear, polynomial, ridge, lasso and elastic net.
~10h · 3 resources
Classification Algorithms
Logistic regression, KNN, Naive Bayes, SVM, decision trees and random forests.
~12h · 6 resources
Clustering, dimensionality reduction and anomaly detection.
Clustering
K-Means, hierarchical clustering and DBSCAN.
~8h · 2 resources
Dimensionality Reduction & Anomaly Detection
PCA and anomaly detection.
~6h · 1 resources
Metrics, tuning, thresholds, class imbalance and explainability.
Metrics & the Confusion Matrix
MAE/MSE/RMSE/R², accuracy/precision/recall/F1, ROC-AUC and PR-AUC.
~8h · 2 resources
Tuning, Thresholds & Explainability
Cross-validation, GridSearchCV/RandomizedSearchCV, threshold selection, imbalance and SHAP.
~8h · 3 resources
Milestone 03 (Complete ML Project). Ensembles, boosting and production-quality tabular pipelines.
Ensembles & Boosting
Bagging, boosting, Gradient Boosting, XGBoost, LightGBM concepts and stacking.
~10h · 4 resources
Query and analyze data where it lives, using PostgreSQL.
SQL & PostgreSQL
SELECT/WHERE/GROUP BY/HAVING/ORDER BY, JOINs, subqueries, CTEs and window functions.
~12h · 5 resources
Turn skills into a reproducible, documented portfolio project.
Reproducible Project Workflow
Datasets, business framing, EDA, baselines, experiment tracking and reproducibility.
~10h · 6 resources
Milestone 04 (Deep Learning Model). Neurons, forward/backprop, losses, optimizers and regularization.
Neural Networks & Backpropagation
Neurons, tensors, forward propagation, loss functions and backpropagation.
~10h · 4 resources
Optimization & Regularization
SGD, Adam, learning rate, batch size, epochs, dropout and batch normalization.
~8h · 4 resources
The primary deep-learning framework for this roadmap.
Tensors, Autograd & Training Loops
Tensors, datasets, dataloaders, nn.Module, autograd, optimizers and training/validation loops.
~14h · 5 resources
Milestone 05 (CV/NLP Project). Images, CNNs, augmentation and transfer learning.
Images, OpenCV & CNNs
Image representation, preprocessing, convolution, pooling and feature maps.
~12h · 4 resources
Detection & Advanced CV
Transfer learning, face detection and object detection (YOLO optional).
~10h · 3 resources
Text preprocessing, classic representations and text classification.
Text Preprocessing & Classification
Tokenization, stopwords, stemming/lemmatization, BoW, TF-IDF, n-grams and embeddings.
~12h · 4 resources
Milestone 06 (Transformer Application). Attention, the transformer architecture and Hugging Face.
Attention & Transformer Architecture
Self-attention, encoder/decoder, positional encoding, BERT and GPT concepts.
~10h · 5 resources
Hugging Face Transformers
Pipelines, the model hub and fine-tuning.
~10h · 3 resources
LLMs, tokens, context windows, prompting, structured outputs and tool calling.
LLMs & Prompting
Foundation models, tokens, temperature, inference and prompt engineering.
~8h · 5 resources
LLM APIs & Tool Calling
OpenAI, Gemini and Anthropic APIs; structured outputs and function/tool calling.
~8h · 4 resources
Represent meaning as vectors and search it fast.
Embeddings & Vector Search
Embeddings, semantic similarity, cosine similarity and nearest-neighbor search with FAISS/Chroma/Qdrant.
~10h · 6 resources
Milestone 07 (RAG Application). Ground an LLM in your own documents with citations.
RAG Architecture & Pipeline
Ingestion, parsing, chunking, metadata, retrieval, reranking, context construction and citations.
~14h · 5 resources
Milestone 08 (AI Agent). Agent loops, tools, planning, memory and multi-step workflows.
Agents, Tools & LangGraph
Agent loops, function calling, planning, memory/state, tool routing and evaluation.
~12h · 6 resources
Serve models and AI features behind a clean, validated API.
FastAPI & Model Inference Endpoints
Routes, request/response models, validation, async endpoints, auth basics and docs.
~10h · 4 resources
Experiment tracking, reproducibility, model/data versioning and registries.
MLflow, DVC & the ML Lifecycle
Experiment tracking, model registry, data versioning and pipelines.
~10h · 3 resources
Package models and services into reproducible containers.
Images, Containers & Compose
Dockerfile, env vars, volumes, networking and Docker Compose.
~10h · 4 resources
Milestone 09 (Production ML API). Ship a model as a public API with config, secrets and health checks.
Deploying Inference Services
Local/API/cloud deployment, environment config, secrets, logging and health checks.
~10h · 3 resources
Latency, batching, caching, monitoring, drift and retraining.
Serving, Monitoring & Drift
Model serving, latency, batching, caching, monitoring, data/model drift and observability.
~10h · 3 resources
Privacy, prompt injection, key security, bias, fairness and explainability.
Securing & Governing AI Systems
Data privacy, prompt injection/jailbreaks, insecure endpoints, API-key security, rate limiting, output validation, bias and fairness.
~8h · 2 resources
Architect production AI systems: serving, batch vs real-time, RAG, caching, queues, scaling and cost.
Designing Production AI Architecture
Model serving, batch vs real-time, vector DB and RAG architecture, caching, queues, async processing, GPU workloads and cost optimization.
~10h · 3 resources
Consolidate Python, SQL, ML/DL, NLP/CV, LLMs/RAG, MLOps and system design for interviews.
Interview Drills & Portfolio
Explain algorithms and projects, debug models, justify metrics and answer system-design questions.
~10h · 4 resources
Milestone 10 (Final Capstone). Build ONE serious end-to-end AI product: ingestion → cleaning → EDA → features → training → evaluation → tracking → serving → FastAPI → Docker → deployment → monitoring. Customize it and explain every major decision.
End-to-End Capstone
Combine everything into a deployed, monitored AI product with documentation.
~40h · 5 resources