Impala 5.0: Where lakehouse tables meet lower latency and better operability

Quanlong Huang

Chinese Session 2026-08-07 16:45 GMT+8  (ROOM : Mtn BaiWang Hall) #olap

Apache Impala is a native query engine implemented using a massively parallel processing (MPP) architecture for open data and open table formats.

In this session, we will share updates from the Impala community over the past year, including highlights from the upcoming 5.0 release:

  • Deeper Apache Iceberg integration, including progress toward Iceberg v3 capabilities such as row lineage and deletion vectors; REST catalog support; and broader improvements to metadata handling and table maintenance.
  • A major milestone for the Calcite-based query planner, including assorted optimizations and correctness fixes for edge cases, with the goal of matching or outperforming the legacy planner where it matters most.
  • Catalog scalability and observability improvements, including warm failover in catalogd HA, HMS incremental event processing, and enhancements to local-catalog mode.
  • Execution-side improvements such as intermediate result caching, late materialization for arrays, more accurate memory estimation, and related optimizations.
  • Additional work across workload management, admissiond, OpenTelemetry integration, and other operational and ecosystem-facing features. We will also briefly outline ongoing efforts targeted at future releases, e.g., History-based Optimizer, PIVOT/UNPIVOT support, AI Query Profile Analyzer, etc.

Speakers:


Quanlong Huang: Cloudera, Senior Staff Engineer

Quanlong Huang is a software engineer at Cloudera. He has been contributing to the Apache Impala project for the past 8+ years. He is a committer and PMC member of Apache Impala, also a committer of Apache ORC and contributor of some other open-source projects like Apache Hive, Apache Hadoop, Apache Thrift, etc.