From Data Ingestion to Data Lake: Building a Modern Lakehouse with Apache SeaTunnel

Lidong Dai

Chinese Session 2026-08-08 14:30 GMT+8  (ROOM : WanChun Hall) #datalake

Building a data lake is no longer just about choosing Iceberg, Hudi, or Paimon. In real-world systems, the biggest challenge often lies one step earlier: how data reliably, efficiently, and continuously enters the lake. In this session, we will explore how Apache SeaTunnel serves as a unified data ingestion and integration layer for modern data lake architectures. Starting from common pain points—multi-source data, Batch + CDC coexistence, schema evolution, and operational complexity—we will walk through how SeaTunnel simplifies data movement into data lakes and lakehouse systems. Through real production scenarios, you will see how SeaTunnel connects transactional databases, message queues, and file systems into Iceberg- or lakehouse-based storage, enabling scalable, maintainable, and evolvable data platforms. The talk focuses on practical architecture decisions, not vendor-specific solutions.

The session will cover:

  • Typical data lake architecture evolution and common pitfalls
  • The role of data integration in lake and lakehouse systems
  • Apache SeaTunnel architecture and design principles
  • End-to-end ingestion examples: databases, CDC, and streaming data into data lakes
  • Operational considerations and best practices
  • Roadmap of SeaTunnel in the data lake ecosystem

Speakers:


Lidong Dai: WhaleOps Technology co-founder

Apache Incubator Mentor, Apache DolphinScheduler PMC member & Apache SeaTunnel PMC member