Data Lake and Data Warehouse are important solutions for storing and managing data, and they play a crucial role in data management, data analysis, and decision-making. In ASF, there are various projects about Data Lake and Data Warehouse, for example: Apache Hive, Apache Hudi, Apache Iceberg, Apache Paimon, Apache Cassandra, Apache HBase, Apache Cloudberry (Incubating) etc. In this topic, you will get the latest status of data lake and warehouse, best practices the companies use them in the production, and the roadmap of these projects.
Data Lake & Data Warehouse
- Track Chairs
- Lidong Dai Shaofeng Shi Zongtang Hu Jean-Baptiste Onofré Huaxin Gao
- Sessions
- 16
- Room
- MainRoom - YiHe Hall
Agenda
2026-08-07
Full schedule →
- 14:00GMT+8
- 14:30GMT+8
- 15:00GMT+8
- 15:45GMT+8
- 16:15GMT+8
-
16:45GMT+8
The Anatomy of Iceberg Failures: Lessons from Real-World EscalationsNoémi Pap-Takács, Boglárka Egyed
2026-08-08
Full schedule →
-
14:00GMT+8
Apache Iceberg V3 in Production: Lessons from a Large-Scale DeploymentYuming Wang, Fei Wang
- 14:30GMT+8
- 15:00GMT+8
-
15:45GMT+8
Evolving a real-time lakehouse: Stability and performance breakthroughs at scaleZhuojun Jiang, Wenling Zhang
- 16:15GMT+8
-
16:45GMT+8
Challenges of Implementing Iceberg Features in a C++ Query EngineZoltán Borók-Nagy, Péter Rózsa
2026-08-09
Full schedule →
- 13:30GMT+8
- 14:00GMT+8
- 14:30GMT+8
- 15:15GMT+8