Deep Dive to Flink CDC Architecture and Large-Scale Production Practices
Yanquan Lv
Chinese Session 2026-08-09 13:30 GMT+8 (ROOM : YuanMing Hall) #streamingSyncing hundreds of databases and tens of thousands of tables, tracking schema evolution,connecting heterogeneous systems—these are common pain points in enterprise real-time data integration. Flink CDC introduces a new Pipeline architecture that enables declarative multi-table synchronization, automatic schema evolution tracking, and one-stop data transformation. Beyond the Pipeline architecture, Flink CDC 3.6 adds support for Flink 2.2 and delivers seamless integration with mainstream databases (MySQL, PostgreSQL, Oracle) and lakehouse systems (Iceberg, Paimon, Hudi). This session will dive into the Pipeline architecture design and share performance tuning and production deployment experiences from actual production environment large-scale real-time data lake ingestion practices, helping attendees build reliable, maintainable enterprise-grade real-time data pipelines.
Speakers:

Yanquan Lv: Apache Flink Committer
I am an Apache Flink Committer and the primary maintainer of the Flink CDC project. Currently working on the Open Source Big Data team at Alibaba Cloud, I focus on real-time data synchronization and stream-batch unified processing. I have helped numerous enterprises implement large-scale real-time data lake ingestion solutions, with extensive experience in distributed systems and real-time computing.