From Lakehouse to Multimodal Data Lake: Rethinking Data Infrastructure for AI
Zheng Yubin, Lili Ma
Chinese Session 2026-08-07 14:00 GMT+8 (ROOM : JingMing Hall) #dataaiThe rapid rise of generative AI and machine learning workloads is fundamentally changing how data platforms are designed. Traditional data lakes and lakehouse architectures — built around large-scale scans and structured analytics — are increasingly challenged by new requirements such as vector search, multimodal data processing, and feature engineering pipelines.
In this talk, we explore how data infrastructure is evolving in the AI era. Modern lakehouse table formats like Apache Iceberg have made schema evolution a well-solved problem. But AI workflows introduce a new challenge — Data Evolution: efficiently backfilling embeddings, recomputing features, and adding multimodal encodings without full table rewrites.
We will also discuss emerging approaches for multimodal data management, including new table formats like Lance — built on the Apache Arrow type system — designed for AI workloads that unify structured data, embeddings, images, audio, and video in a single system.
Finally, we will look at real-world architectural patterns such as Netflix’s Media Data Lake and explore how open ecosystems are enabling a new generation of AI-native data platforms.
Speakers:

Zheng Yubin: AWS, Senior Developer Advocate
Zheng Yubin, with over 20 years of experience in the ICT industry and digital transformation. Currently, serving as a Senior Developer Evangelist at AWS, specializing in Cloud Native, Cloud Security, and Generative AI. As the first female technical evangelist at AWS China, engaged with the developer community. As an architect with 18 years of experience, have provided consulting and technical implementation of solutions, including data center construction and software-defined data centers, for the finance, education, manufacturing, and high-tech industry. Leveraging industry expertise, offering technical guidance to developers, seeking mutual success.

Lili Ma: AWS, Senior Data Specialist Solutions Architect
Lili Ma is a Data Specialist Solutions Architect at Amazon Web Services with over a decade of experience in data infrastructure research and product innovation. She started with Hadoop and Hive during her academic years, then moved through IBM DB2, the MPP data warehouse Greenplum, the compute-storage decoupled Apache HAWQ, and cloud-native databases Amazon Aurora and ElastiCache. She is a PMC member of the Apache HAWQ project and an early member of the Greenplum team. She has published multiple academic papers at international conferences including SIGMOD, GCC, SKG, and PDCAT, and holds several international patents. She has presented at ApacheCon, DTCC, KCD, and Greenplum community events on topics ranging from distributed database scalability to cloud-native architectures.