Live Cohort 2 · Starts 29th Aug

From Data Analyst to Data Engineer in 10 weeks.

You do not watch this one. You build it.

A 10 weeks intensive LIVE cohort to go from writing queries to building and running production-grade data pipelines.

Inner Circle US$630 US$840 after 25th Aug
This cohort includes Data Engineering Bootcamp 1.0 worth US$240
Proof, before the pitch

Analysts who actually shipped with us.

Real reviews from Cohort 1. They did not finish a course, they finished a portfolio.

Mansi Katarmal
Mansi Katarmal Technology Analyst

"This cohort has made a real impact on my learning journey. The curriculum is well-structured and aligned with industry practices, making it easy to connect concepts with real-world implementation. What stands out the most is the combination of the live sessions, learning resources, and hands-on assignments. Many concepts feel like a black box before the session, but the live classes break them down in a structured way."

Jaideep Gupta
Jaideep Gupta Tableau Developer

"I really enjoyed the LIVE Data Engineering Bootcamp for Analysts. The best part of the course was definitely Harun. His depth of knowledge is incredible, and he doesn't just teach how things work, he explains why they work that way. That made it much easier to understand concepts that initially felt overwhelming. Another thing I appreciated was the Codebasics team."

Naga Durga Srivani
Naga Durga Srivani Business Analyst

"If you are looking for a data engineering course that actually teaches you how things work in the real world, this is it. Harun is an excellent instructor — he knows his subject inside out and never makes you feel rushed when you have a question. The live sessions are well structured and interactive, and the Codebasics team is always present and quick to help whenever you get stuck."

What learners say

Real feedback from Data Engineering learners.

This setup really benefits from having two instructors-one to lead the class and another to review the questions, identify the most important ones, and pause the session when needed. The approach worked very well. The coordination and synchronization between both of you were excellent.
Kiran B
Wow..!! The narrative of analytics to DE with example and the roadmap is shown is amazing.
Vishwas
The business scenario and assignments is a great way to learn.Thanks Hemanand.
Shaurya Vashistha
I joined DE Live bootcamp and the live sessions are so good - you guys really made sure everyone understood foundations and concepts very well. Glad I joined.
Keerthana V
The use of whiteboard to explain concepts was useful to me. Seeing concepts visually first before code helps!
Rucha
Thank you team, appreciate the efforts!
Shalini Goutam
I have to drop from here. Thanks Haroon its was eye opening sessic through!!! So much to learn
Ashish Babaria
Understood the missing parts of my databricks knowledge. Thanks Harun.
Naman Garg
thank you Harun, it was a great session.
Indrani
1.5M+ YouTube Subscribers
5 ★ 7000+ reviews
721K+ Learners
By Dhaval Patel, Hemanand & Team
Learn from people who ship

Four practitioners who run this in production.

Dhaval Patel Dhaval Patel
Founder, Codebasics · Ex-NVIDIA
1 Million+ YouTube subscribers · 659K+ learners · Built AI products at NVIDIA

Programme Creator: Designs the curriculum, decides what gets taught and in what order, sets the capstone brief, and holds the quality bar across every module.

Hemanand Vadivel Hemanand Vadivel
Co-founder, Codebasics · Ex-Edgewell
10+ years in analytics and leadership · Built and scaled teams across international markets at Edgewell

Teaches Orchestrate and Distribute modules: productivity systems, stakeholder management, LinkedIn strategy, personal branding.

Harun Harun
Lead Faculty
10+ years building enterprise data platforms using Azure, Databricks, and modern cloud technologies for large-scale organisations. Brings that same production mindset into every live session.

Teaches: Advanced SQL, Data Modelling & Warehousing, Azure Data Factory, Azure Databricks, PySpark, Fabric, Delta Lake, Airflow Orchestration & End-to-End Data Pipelines.

Kiran Kiran
Content Curator & Program Manager, Codebasics
Analytics Engineer driving curriculum development and learner success through practical, project-based learning experiences.

Teaches: Capstone Support, Project Guidance, Assignments, Practice Sessions & Doubt Resolution.

The stack

Every tool real Data Engineers use in 2026.

Advanced SQL Advanced SQL
Python & pandas Python & pandas
PySpark PySpark
Databricks & Delta Lake Databricks & Delta Lake
Microsoft Fabric Microsoft Fabric
ADLS Gen2 ADLS Gen2
dbt Core & Cloud dbt Core & Cloud
Apache Airflow Apache Airflow
Azure Data Factory Azure Data Factory
Apache Kafka Apache Kafka
Azure Event Hubs Azure Event Hubs
GitHub Actions GitHub Actions
CI/CD CI/CD
Azure Key Vault Azure Key Vault
Power BI Power BI
Week by week

10 weeks. 19 live sessions. One production pipeline.

Session 1 - SQL for Modern Data Engineering, Part 1
Joins beyond INNER and OUTER · Join algorithms: nested loop, hash, merge · Anti-joins · Window functions deep dive · Partition, order and frame · Ranking and running totals
Output: A ranking and running totals query pack on a sample warehouse
Session 2 - SQL for Modern Data Engineering, Part 2
Recursive CTEs · MERGE statements · Incremental loading patterns · CDC concepts · Query execution plans · Warehouse optimisation
Output: An incremental load built with MERGE and tuned from the execution plan
Session 3 - Data Modelling & Warehouse Engineering
OLTP vs OLAP · Star and snowflake schema · Fact vs dimension tables · SCD Type 1 and Type 2 · Partitioning strategies · Medallion architecture · Data contracts
Output: A dimensional model with SCDs on a Bronze, Silver, Gold layout
Session 4 - Production Python for Data Engineers
Modular Python architecture · OOP for pipelines · Config-driven frameworks · Logging and exception handling · Retry mechanisms · Environment management · Secrets handling
Output: A config-driven Python pipeline framework with logging and retries
Session 5 - Advanced Python Data Processing
APIs and ingestion patterns · Async processing · Parallel execution · File streaming · Memory optimisation · Testing with pytest · Packaging basics
Output: A tested async ingestion job that streams large files without blowing memory
Session 6 - PySpark Deep Dive
Spark architecture · Executors and DAGs · Lazy evaluation · Partitioning · Broadcast joins · Shuffle optimisation · Spark UI analysis · Caching strategies
Output: A Spark job you have tuned yourself using the Spark UI
Session 7 - Delta Lake & Lakehouse Engineering
Delta internals · ACID transactions · OPTIMIZE and ZORDER · Time travel · Schema evolution · Change Data Feed · Incremental ETL · Bronze, Silver, Gold
Output: An incremental ETL pipeline running on Delta Lake
Session 8 - Azure Data Engineering Stack
ADLS Gen2 · Event Hubs · Key Vault · Managed identities · Integration Runtime · Networking basics · Synapse vs Databricks vs Fabric
Output: A secured Azure data landing zone using Key Vault and managed identities
Session 9 - Enterprise Data Pipelines
Azure Data Factory · Fabric Pipelines · Databricks Workflows · Metadata-driven pipelines · Config-based orchestration · Parameterisation · Reusable frameworks
Output: A metadata-driven, parameterised orchestration pipeline
Session 10 - Microsoft Fabric Engineering
OneLake · Lakehouse and Warehouse · Fabric Data Factory · Eventstream · Real-Time Intelligence · DirectLake · Fabric governance
Output: An end-to-end Fabric solution with Direct Lake reporting on top
Session 11 - Connecting the Dots
End-to-end view of the stack so far · How the pieces fit together · Architecture recap · Trade-offs between platform options · Concept clarity and doubt clearing
Output: One architecture diagram of the full stack, explained in your own words
Session 12 - dbt Core Fundamentals
Models and sources · refs() · Materialisations · Snapshots · Incremental models · Tests · Documentation
Output: A tested, documented dbt project with incremental models and snapshots
Session 13 - CI/CD & Reliability Engineering, plus Capstone Introduction
Git branching strategies · GitHub Actions · Automated testing · Deployment pipelines · Monitoring and freshness checks · Cost optimisation · Incident management · Capstone brief and assessment criteria
Output: A CI/CD workflow with automated tests and freshness checks, plus your capstone brief
Session 14 - Stakeholder Management & Personal Branding
Stakeholder communication · Framing requirements · Explaining technical work to non-technical audiences · Personal branding · Building online credibility as a data engineer
Output: A stakeholder-ready update on your own work and a refreshed professional profile
Session 15 - Capstone Jamming Session
Implementation questions on your capstone · Design review · Debugging together · Unblocking pipeline issues · Peer feedback
Output: Your capstone unblocked, with a clear next step
Session 16 - Advanced Analytics & Apache Airflow Engineering
Macros and Jinja · Semantic layer · MetricFlow · SQLFluff · Lineage · Data quality frameworks · Governance · DAG architecture · Dynamic DAGs · Sensors · XCom · Scheduling and monitoring · Retry patterns · Failure handling
Output: A semantic layer with lineage, plus scheduled Airflow DAGs that survive failures
Session 17 - Streaming Data Engineering
Kafka fundamentals · Event-driven architecture · Structured Streaming · Watermarking and windowing · Event-time processing · CDC streaming · Event Hubs integration
Output: A streaming pipeline that handles late-arriving data with watermarks
Session 18 - Interview Prep & System Design
SQL interview rounds · PySpark interview questions · Data modelling rounds · System design · Resume transformation · LinkedIn optimisation · Mock interviews
Output: A DE-positioned resume and profile, tested in a mock interview
Session 19 - Final Demo & Graduation
Capstone demo covering API ingestion · Lakehouse architecture · PySpark transformations · dbt modelling · Airflow orchestration · CI/CD · Power BI reporting · Monitoring
Output: One production-grade pipeline repo covering all eight layers, ready for a hiring manager to read in five minutes
The Codebasics promise

Premium data engineering training, without the premium price.

Other paths can work. Here is exactly what you get from each.

What you get Self-study Other live bootcamps Our live cohort
Live, mentor-led sessions
Real projects shipped to GitHub Rarely Sometimes 1 pipeline repo
The current 2026 lakehouse stack On your own Often outdated
Job assistance
InvestmentYour time₹2,00,000+US$630
Investment

The earlier you join, the less you pay.

Inner Circle learners enroll at a lower price until 25th Aug. Standard pricing applies later. Cohort begins on 29th Aug.

Save US$210 · until 25th Aug
US$630 US$840
One-time payment · EMI available · Standard price after 25th Aug

What's included

  • 10 weeks live cohort, 19 live sessions
  • One end-to-end pipeline repo, all 8 layers, on your GitHub
  • The full 2026 data engineering stack
  • Job assistance & interview prep
  • 1 year access to all session recordings
  • Lifetime version access + certificate
EMI available

Pay in monthly instalments instead of one payment. Choose your EMI plan at checkout, no extra paperwork.

Secure checkout · UPI, cards, net banking · Razorpay · No-questions refund policy
The same curriculum, 19 live sessions, capstone, and faculty for everyone.
Everything you'd reasonably want to know

Questions?

When does the bootcamp officially start? +
The bootcamp officially commences on Saturday, 29th August 2026.
When do I get access after enrolling as an Inner Circle member? +
You are securing your seat now. Full bootcamp access opens on 29th August 2026.
What is the Inner Circle, and how is it different from regular enrollment? +
The Inner Circle is early enrollment, open until 25th August 2026. Inner Circle members enroll at a reduced price and get a direct say in the curriculum before the bootcamp launches on 29th August. You will receive a short feedback form asking which tools, topics and gaps matter most in your work, and your inputs shape what this cohort covers. You are not just enrolling early, you are helping shape what gets built.
Do I also get the Data Engineering Bootcamp 1.0? +
Yes. Every enrollment includes full access to the Data Engineering Bootcamp 1.0 at no extra cost. It includes Job Assistance, Live Problem Solving, and a Virtual Internship. You get both for the price of one.
When are the live sessions? +
Saturdays and Sundays, 4 to 7 PM IST. Sessions are fully live and interactive with hands-on labs and real-time Q&A. Recordings are available for revision.
What if I miss a live session? +
All live sessions are recorded and available within 24 hours. You can catch up at your own pace, though live attendance is strongly recommended as the labs and discussions are where most of the real learning happens.
What happens after the 10 weeks? Do I lose access? +
No. You keep access to all session recordings for 1 year from the bootcamp start date.
Do I need prior data engineering experience? +
No. This bootcamp is built for working data analysts who want to cross into data engineering. If you have at least 1 year of analyst experience and are comfortable with SQL, you are ready.
I am a fresher with no work experience. Can I join? +
We strongly advise against it. This bootcamp moves fast and assumes analyst-level SQL fluency and data literacy. If you are starting from zero, the Codebasics Data Analytics Bootcamp is the right first step. Build that foundation and come back.
Who is this bootcamp designed for? +
Working data analysts, BI developers, and business analysts with 1 to 4 years of experience who want to own the full data stack, not just the dashboard layer. If you already work as a data engineer, this bootcamp is likely below your current level.
How do I get help if I am stuck? +
Every enrolled learner gets access to the Discord community where you can ask questions, connect with fellow learners, share progress, and learn from each other throughout the bootcamp. The mentor team also provides weekly hands-on lab support.
Is there job assistance? +
The Data Engineering Bootcamp 1.0 included with your enrollment has dedicated job assistance. The bootcamp itself focuses on building your skills, shipping a production-grade capstone on GitHub, and preparing you for system design interviews with the core faculty.
I used a subsidy and now want to refund this Bootcamp itself. What happens? +
You get back the amount you actually paid for this Bootcamp. Your original purchase stays intact and you keep access to it.
I used a subsidy (my existing Data Engineering Bootcamp or individual course purchase). Can I refund my original purchase after enrolling? +
No. Once your existing purchase is applied as a subsidy to reduce your price, that original purchase becomes non-refundable.
What is the refund policy? +
Full refund, no questions asked, if you request it on or before 31st August 2026. That is after the first two live sessions (29th and 30th August), so you can see exactly how the bootcamp runs before deciding.
What is the Inner Circle price and when does it close? +
Inner Circle enrollment is open until 25th August 2026. After that, Standard pricing opens from 26th August 2026 until the bootcamp starts on 29th August 2026.
I already own the Data Engineering Bootcamp 1.0. What do I pay? +
The amount you paid for the Data Engineering Bootcamp 1.0 is fully adjusted and deducted from your enrollment fee.
I bought only some individual courses from the Data Engineering Bootcamp 1.0. What do I pay? +
The amount you paid for those individual courses is deducted from your enrollment fee.
Cohort 2 · Starts 29th Aug

The best time was yesterday. The second best is now.

US$630 US$840 after 25th Aug
Enroll now →
Talk to us Chat with us