Building a Multimodal AI Lakehouse with Apache Gravitino in China Mobile Wutong Data Platform
Xiaojing Fang, Xintong Jiang
Chinese Session 2026-08-07 15:00 GMT+8 (ROOM : YuanMing Hall) #aiinfraAs AI workloads become a core part of modern data platforms, enterprises need to manage not only structured tables, but also unstructured data, vector data, AI models, and AI functions in a unified way. Traditional lakehouse architectures are often centered on tabular data only, which makes it difficult to build a consistent metadata and control plane across heterogeneous data and AI assets.
In this session, we will share how China Mobile is building a multimodal AI lakehouse in the Wutong Data Platform with Apache Gravitino. In this architecture, Gravitino serves as the unified metadata and control plane to manage structured and unstructured data together with AI models and AI functions. This enables consistent governance, discovery, and management across different types of assets in the platform.
We will also discuss how the platform combines Apache Iceberg and Lance to support unified analytics across tabular and multimodal data. In addition to introducing the relevant capabilities of Apache Gravitino, this talk will present practical experience from China Mobile’s production platform, including architecture design, metadata organization, integration patterns, governance considerations, and lessons learned from real-world deployment.
This session is intended for data platform engineers, lakehouse architects, and AI infrastructure practitioners who are exploring open architectures for multimodal data and AI asset management.
Speakers:

Xiaojing Fang: Apache Gravitino Committer
Apache Gravitino PPMC, architect at China Mobile, focusing on data and AI infrastructure.

Xintong Jiang: Software R&D Engineer at ChinaMobile
Boasting more than a decade of expertise in big data and AI, spearheads the development of PB-scale multimodal data lakes at China Mobile Digital Intelligence Business Unit.Core competencies lie in distributed computing and hybrid retrieval.