track · predict
Machine Learning Engineer Track
Prepare for Machine Learning Engineer roles focused on data preparation, classical ML algorithms, feature engineering, evaluation, tuning, interpretation, and portfolio-ready projects.
Before you start
- Python, SQL, Git, and Linux basics
- Algebra, probability, statistics, and optimization fundamentals
- Ability to complete weekly coding projects and maintain a GitHub portfolio
- Comfort with notebooks, source-code files, documentation, and reproducible environments
18-week roadmap
| Weeks | Focus | Portfolio checkpoint |
|---|---|---|
| Weeks 1-3 | Python, SQL, Git, notebooks, data cleaning | Clean EDA repository |
| Weeks 4-6 | Math, statistics, probability, optimization | From-scratch regression exercise |
| Weeks 7-11 | Supervised learning and feature engineering | Two business prediction projects |
| Weeks 12-15 | Unsupervised learning, anomaly detection, evaluation | Segmentation or anomaly project |
| Weeks 16-18 | Portfolio polish, documentation, interview practice | Final GitHub portfolio review |
Detailed curriculum
› stage/01/foundationsPython, SQL, and Data Handling Foundations
Project: build a clean data ingestion, validation, and EDA pipeline for a messy customer, sales, or operations dataset.
- Python functions, OOP basics, type hints, virtual environments
- NumPy and pandas: data cleaning and joins
- SQL joins, aggregations, CTEs, window functions
- Git workflow, Linux shell basics, pytest basics
› stage/02/math-statsMath and Statistics for Machine Learning
Project: implement linear regression with gradient descent from scratch and validate it against a library implementation.
- Linear algebra: vectors, matrices, dot product, matrix multiplication
- Probability: distributions, conditional probability, Bayes' rule
- Statistics: mean, standard deviation, correlation, sampling, confidence intervals
- Optimization: gradients, loss functions, learning rate intuition
› stage/03/feature-engineeringData Preparation and Feature Engineering
Project: create a reusable preprocessing workflow for a tabular business dataset and document every transformation.
- Missing-value handling, outlier handling, duplicate checks
- Encoding: one-hot, ordinal, target encoding awareness
- Scaling: StandardScaler, MinMaxScaler, RobustScaler
- Train/validation/test split and leakage prevention
› stage/04/regressionSupervised Learning: Regression
Project: predict price, cost, delivery time, or customer value using structured data.
- Regression problem framing and metric selection
- Regularization and the bias-variance trade-off
- Residual analysis and error analysis
- Business interpretation of numeric predictions
› stage/05/classificationSupervised Learning: Classification
Project: build a churn, loan-risk, lead-scoring, fraud-risk, or support-ticket classification model.
- Binary and multiclass classification workflows
- Class imbalance, threshold selection, calibration awareness
- Confusion matrix, precision, recall, F1, ROC-AUC, PR-AUC
- Feature importance and classification error analysis
› stage/06/unsupervisedUnsupervised Learning and Anomaly Detection
Project: build customer segmentation and anomaly detection for transaction, usage, or operational data.
- Distance metrics and clustering assumptions
- Cluster validation and cluster profiling
- Dimensionality reduction for visualization
- Outlier detection for suspicious or unusual records
› stage/07/evaluationModel Evaluation, Tuning, and Interpretation
Project: tune and interpret a tabular ML model for a real business-style decision problem.
- Cross-validation and validation strategy
- Hyperparameter tuning without leakage
- Metric selection by business objective
- Interpreting important drivers and common failure cases
› stage/08/portfolioPortfolio, Communication, and Interview Readiness
Project: polish three ML repositories into a job-ready portfolio with clear problem statements and results.
- Writing clear READMEs and project summaries
- Explaining metrics to technical and non-technical audiences
- Answering ML interview questions with trade-offs
- Maintaining clean, reviewable GitHub repositories
Capstone portfolio projects
Business Prediction System
A complete classical ML project for churn, risk, lead scoring, customer value, price, or cost prediction, from EDA through a tuned final model.
Segmentation and Anomaly Analysis
Unsupervised learning used to segment users, products, or transactions and flag unusual records worth investigating, with business recommendations.
Choose Your Format
Self-paced learners who want a personalized schedule.
- Full curriculum
- Weekly 1:1 sessions
- Paced to you
Learners who want peer accountability on a fixed schedule.
- Full curriculum
- Structured batch sessions
- Peer projects
Learners who want full support through to an offer.
- 1:1 or Cohort format
- Full placement support
- Interview prep
Includes full placement support and interview prep.