Building AI's Data Artery: Architecture and Practices of Unified Multimodal Data Pipelines

Xiaochen Zhou

Chinese Session 2026-08-09 13:30 GMT+8  (ROOM : Mtn WanShou Hall) #dataops

Abstract: In the GenAI era, the massive flow of multimodal data demands a robust infrastructure, yet fragmented data pipelines have become a critical bottleneck for enterprises. At Tongcheng Travel, we historically operated 4 disjointed data pipeline services (Offline Sync, Real-time Lake Ingestion, legacy Sqoop, and a standalone SeaTunnel service). This fragmentation caused extremely high maintenance costs and hindered unified data governance.

This session details how we successfully architected a unified “Data Artery” through platformization. We will explore how we consolidated the data entry points and built a true “Batch-Stream Unified” foundational architecture based on Apache SeaTunnel, comprehensively supporting data flows from traditional data warehouses to modern AI scenarios.

Key Content:

  1. Breaking Data Silos: A deep dive into designing a unified multimodal data pipeline service powered by the Apache SeaTunnel engine, smoothly replacing and consolidating 4 legacy integration systems to achieve complete architectural standardization.
  2. Compute Enhancement & AI Multimodal Empowerment: Exploring how to deeply integrate SeaTunnel’s Transform mechanism with real-time stream processing capabilities to efficiently execute complex data cleaning and dynamic transformations. We will highlight hardcore support for AI workloads, including real-time parsing of unstructured data and Embedding preprocessing for LLMs.
  3. Zero-Downtime Migration & Strict Validation: Sharing enterprise-grade practices on migrating massive legacy tasks. We will detail the “Dynamic Task Conversion and Bi-directional Data Reconciliation Mechanism” we designed to ensure zero data loss and a seamless transition for the business during the underlying architecture upgrade.
  4. Future Cloud-Native Evolution: Looking ahead at the blueprint for multimodal unified data pipelines. We will discuss cloud-native containerized deployments on Kubernetes for elastic scaling, and how to deeply integrate with the LLM ecosystem to build a robust Data+AI foundation.

Speakers:


Xiaochen Zhou: Data Engineer @ Tongcheng Travel | Apache SeaTunnel Committer

Xiaochen Zhou is a Data Engineer at Tongcheng Travel and an active Apache SeaTunnel Committer. In his current role, he specializes in designing, building, and optimizing high-performance data pipelines. Within the open-source community, he is deeply involved in the core development and technical evolution of Apache SeaTunnel. Recently, his focus has shifted to the intersection of Data and AI, where he is dedicated to architecting unified multimodal data pipelines for the GenAI era.