Efficient metadata caching in Apache Impala

Csaba Ringhofer, Noémi Pap-Takács

English Session 2026-08-07 15:00 GMT+8  (ROOM : Mtn BaiWang Hall) #olap

How to cache enough info in memory to plan queries on > 1 million file tables? Our talk will describe Apache Impala’s metadata caching solution that can deal with massive Iceberg and Hive tables efficiently.

Caching table and file metadata is key to high query throughput by allowing query planning without reaching out to other services. For huge tables (e.g. >1M files) the metadata cache becomes a massive memory hog and expensive to maintain. Impala’s current fine-grained caching and invalidation is effective for Hive tables, but doesn’t fully translate to Iceberg’s architecture. The focus of the talk is a new, cleaner and more efficient caching solution for Iceberg tables.

Attendees can learn about:

  • metadata caching problem in SQL engines
  • caching quirks for classic Hive and Iceberg tables
  • the effect of file system: cost of caching in Hadoop vs object stores
  • minimizing memory footprint in Java
  • benchmarking the overhead of updating the cache

Speakers:


Csaba Ringhofer: Software engineer at Cloudera

Csaba Ringhofer has been working on Apache Impala since 2017 at Cloudera. He is a Member of the Apache Impala PMC. He studied at the Budapest University of Technology and Economics.


Noémi Pap-Takács: Apache Impala Committer

Noémi Pap-Takács is a software engineer at Cloudera and a committer on the Apache Impala project. Her focus lies in performance optimization and the integration of Apache Iceberg into Impala.