Empowering Agentic AI: How Apache Projects Can Shape the Next Open Data and AI Stack
As AI systems evolve from passive assistants into more autonomous, goal-driven agents, the need for open, reliable, and interoperable infrastructure becomes increasingly important. This panel explores how Apache projects can help empower the next wave of agentic AI by providing the core building blocks for data access, processing, governance, streaming, orchestration, and execution.
From Apache Kafka for real-time event streams, Apache Spark and Apache Flink for large-scale data processing, Apache Iceberg for modern lakehouse storage, to metadata and governance projects such as Apache Gravitino, the Apache ecosystem already provides much of the foundation needed for intelligent agents to interact with enterprise data safely and effectively. The discussion will examine where open-source data infrastructure is already enabling agentic workflows, what technical gaps still remain, and how Apache communities can collaborate to define the next generation of AI-native data systems.
This panel is both a reflection on Apache’s long-standing role in shaping the modern data stack and a forward-looking conversation about how open communities can help build trusted, governed, and extensible foundations for agentic AI.
Speakers:

Junping Du: Found & CEO
Founder and CEO of Datastrato, Ex-Chairperson of LF AI & DATA, ASF Member, Committer and PMC for Apache Hadoop, Co-founder of Apache Ozone, YuniKorn, etc.

William Guo: Apache Software Foundation Member
- Apache Software Foundation Member
- Apache IPMC Member
- PMC of Apache DolphinScheduler
- Mentor of Apache SeaTunnel(incubating)
- Founder of ClickHouse China Community
- Track Chair of Workflow/Data Governance of Apache Con Asia 2021/2022
William used to be the CTO of Analysys and the Senior Big Data Director of Lenovo, general manager of bid data in Wanda. He worked as Big Data Director/manager at CICC, IBM, and Teradata. He has more than 20 years of experience in big data technology and data management.

Jerry Shao: Datastrato, CTO
Jerry Shao is the co-founder and CTO of Datastrato, focused on open source Big Data are for more than 10 years. He is an Apache member, committer and PMC member of Apache Spark and Apache Inlong, the original creator of Apache Gravitino.

Tom Tan: Datastrato advisor
Advisor Datastrato & Kwaai AI lab Head of AI, data and infra, SmartNews VP of engineering, cloudwalk Director, Apple data and ML platform

Mark Hoerth: Product Lead, Datastrato
Mark Hoerth is Product Lead and Solutions Architect at Datastrato, where he works with enterprise data teams to turn federated governance requirements into shipped product, most recently Iceberg REST Catalog federation and multi-cloud credential vending in Gravitino 1.3. He previously spent seven years at Dremio, since acquired by SAP, ending as Principal Product Manager for core technology including Apache Iceberg and Dremio’s Polaris-based open catalog. He holds two degrees from Stanford University and is based in the San Francisco Bay Area.

Attila Turóczy: Senior Director of Engineering at Cloudera
Apache Hive, Impala and Big Data enthusiasm at Cloudera