Course Introduction
Course overview and expectations.
- Workshop
- Python
- Presenter
- Dr. Babak Hossein Khalaj
Ordered session timeline with workshops and presenters.
Course overview and expectations.
Data-science in production and buisness contexts.
Foundations of data representation and modeling.
Relational databases, querying, and data acquisition.
Principles of exploratory and explanatory visualization.
Data collection, cleaning, preprocessing, and feature preparation.
Review on probability and statistics; based primarily on Tibshirani et al.
Probability, estimation, uncertainty, and elementary inference for statistical learning.
Regression taught from a statistical-learning perspective, based primarily on Tibshirani et al.
Classification, regularization, model selection, and related topics; based primarily on Tibshirani et al.
Clustering, dimensionality reduction, and unsupervised learning; semi-supervised learning included if time permits. Based primarily on Tibshirani et al.
Causal concepts and related material from the original syllabus.
Evaluation protocols, validation, metrics, and error analysis.
Deep-learning foundations based on Sergey Levine’s course materials.
Core neural architectures and optimization, based on Sergey Levine’s course materials.
Replaces the original Classic NLP, Modern NLP Architecture, and LLM survey sequence; covers attention and transformer fundamentals.
Hands-on construction of a small LLM, using Andrej Karpathy’s video lectures as the main practical reference.
Foundations of generative modeling, based on the principal/original research papers.
Diffusion and score-based generative modeling, based on the principal/original research papers.
Moved after transformers and generative models to use the required neural-sequence foundations; based on recent transformer-based time-series papers.
Implementation, reproducibility, and integration of trained models.
End-to-end machine-learning pipelines.
Deployment monitoring, drift, reliability, and lifecycle management.