Building an AI-Powered Spark Diagnosis Assistant Across the Apache Data Stack

Tianhang Li

Chinese Session 2026-08-07 16:45 GMT+8  (ROOM : JingMing Hall) #dataai

Modern Spark platforms generate a huge amount of operational signals, but diagnosing performance bottlenecks, resource waste, and recurring production issues still depends heavily on expert experience. In this session, we will share how we built an AI-powered diagnosis assistant that connects multiple Apache big data components, including Gravitino, Spark History Server, YARN ResourceManager, Celeborn, and Ranger, to provide intelligent analysis and actionable recommendations for Spark workloads.

The assistant helps engineers troubleshoot failed or slow Spark jobs, identify optimization opportunities, and deliver governance insights across performance, resource usage, storage efficiency, and operational stability. By combining metadata, runtime metrics, scheduling signals, shuffle behavior, and access control context, the system can generate practical suggestions for SQL tuning, resource sizing, skew mitigation, shuffle optimization, data governance, and daily on-call troubleshooting.

We will cover the overall architecture, data collection and reasoning pipeline, real-world diagnosis scenarios, and the measurable impact on compute cost, storage cost, and engineering efficiency. This talk is intended for data platform engineers, Spark practitioners, and open source users interested in applying AI to observability, operations, and optimization in the Apache ecosystem.

Speakers:


Tianhang Li: “Big Data Development Engineer at Bilibili | Apache Gravitino Contributor | Expert in Metadata Management & Spark Optimization”

Li Tianhang is a Big Data Development Engineer at Bilibili, where he specializes in metadata management and Spark computing engine optimization for large-scale data scenarios