Efficient metadata caching in Apache Impala
Csaba Ringhofer, Noémi Pap-Takács
English Session 2026-08-07 15:00 GMT+8 (ROOM : Mtn BaiWang Hall) #olapHow to cache enough info in memory to plan queries on > 1 million file tables? Our talk will describe Apache Impala’s metadata caching solution that can deal with massive Iceberg and Hive tables efficiently.
Caching table and file metadata is key to high query throughput by allowing query planning without reaching out to other services. For huge tables (e.g. >1M files) the metadata cache becomes a massive memory hog and expensive to maintain. Impala’s current fine-grained caching and invalidation is effective for Hive tables, but doesn’t fully translate to Iceberg’s architecture. The focus of the talk is a new, cleaner and more efficient caching solution for Iceberg tables.
Attendees can learn about:
- metadata caching problem in SQL engines
- caching quirks for classic Hive and Iceberg tables
- the effect of file system: cost of caching in Hadoop vs object stores
- minimizing memory footprint in Java
- benchmarking the overhead of updating the cache
Speakers:

Csaba Ringhofer: Software engineer at Cloudera
Csaba Ringhofer has been working on Apache Impala since 2017 at Cloudera. He is a Member of the Apache Impala PMC. He studied at the Budapest University of Technology and Economics.

Noémi Pap-Takács: Apache Impala Committer
Noémi Pap-Takács is a software engineer at Cloudera and a committer on the Apache Impala project. Her focus lies in performance optimization and the integration of Apache Iceberg into Impala.