|
via Udemy |
Go to Course: https://www.udemy.com/course/apache-iceberg-data-lakehouse-engineering/
Welcome to Data Lakehouse Engineering with Apache Iceberg: From Basics to Best Practices - your complete guide to mastering the next generation of open table formats for analytics at scale.As the data world moves beyond traditional data lakes and expensive warehouses, Apache Iceberg is rapidly becoming the cornerstone of modern data architecture. Built for petabyte-scale datasets, Iceberg brings ACID transactions, schema evolution, time travel, partition pruning, and compatibility across multiple engines - all in an open, vendor-agnostic format.In this hands-on course, you'll go far beyond the basics. You'll build real-world data lakehouse pipelines using powerful tools like:PyIceberg - programmatic access to Iceberg tables in PythonPolars - lightning-fast DataFrame library for in-memory transformationsDuckDB - local SQL powerhouse for interactive developmentApache Spark - for large-scale batch and streaming processingAWS S3 - cloud-native object storage for Iceberg tablesAnd many more: SQL, Parquet, Glue, Athena, and modern open-source utilitiesWhat Makes This Course Special?Hands-on & Tool-rich: Not just Spark! Learn to use Iceberg with modern engines like Polars, DuckDB.Cloud-Ready Architecture: Learn how to store and manage your Iceberg tables on AWS S3, enabling scalable and cost-effective deployments.Concepts + Practical Projects: Understand table formats, catalog management, schema evolution, and then apply them using real datasets.Open-source Focused: No vendor lock-in. You'll build interoperable pipelines using open, community-driven tools.What You'll Learn:The why and how of Apache Iceberg and its role in the data lakehouse ecosystemDesigning Iceberg tables with schema evolution, partitioning, and metadata managementHow to query and manipulate Iceberg tables using Python (PyIceberg), SQL, and SparkReal-world integration with DuckDB, and PolarsUsing S3 object storage for cloud-native Iceberg tablesPerforming time travel, incremental reads, and snapshot-based rollbacksOptimizing performance with file compaction, statistics, and clusteringBuilding reproducible, scalable, and maintainable data pipelinesWho Is This Course For?Data Engineers and Architects building modern lakehouse systemsPython Developers working with large-scale datasets and analyticsCloud Professionals using AWS S3 for data lakesAnalysts or Engineers moving from Hive, Delta Lake, or traditional warehousesAnyone passionate about data engineering, analytics, and open-source innovationTools & Technologies You'll Use:Apache Iceberg, PyIceberg, Spark,DuckDB, Polars, Pandas, SQL, AWS S3, ParquetIntegration with Metastore/Catalogs (REST, Glue)Hands-on with Jupyter Notebooks, CLIBy the end of this course, you'll be able to design, deploy, and scale data lakehouse solutions using Apache Iceberg and a rich ecosystem of open-source tools - confidently and efficiently.