Starburst Workshop: Iceberg Ingestion and Table Optimization

Ingesting high-volume data streams into open lakehouse formats is only half the battle. Without intentional ingestion pipelines and continuous table maintenance, Apache Iceberg deployments can rapidly suffer from the “small file problem,” bloated metadata snapshots, and degrading query speeds that inflate compute costs.
Join Starburst for this hands-on technical workshop, “Iceberg Ingestion and Table Optimization.” Designed for data engineers and lakehouse architects, this session cuts through high-level theory to demonstrate how to build efficient, scalable ingestion pipelines directly into Apache Iceberg tables while applying automated optimization techniques to keep analytical workloads running at peak performance.
Through step-by-step demonstrations, you will explore best-practice ingestion patterns, learn how to prevent table bloat, and master native optimization operations that ensure your lakehouse remains fast, clean, and cost-effective as data volumes expand.
What You’ll Learn:
- High-Throughput Ingestion Pipelines: Implement reliable, scalable data ingestion patterns into Apache Iceberg without creating file fragmentation.
- Tackling the Small File Problem: Leverage data compaction routines to merge small files into optimal layouts for Trino and Iceberg query execution.
- Automated Table Maintenance: Configure routine snapshot expiration, orphan file cleanup, and manifest rewrites to prevent metadata sprawl.
- Storage and Query Optimization: Fine-tune partition layouts, sorting strategies, and metrics collection to accelerate downstream read performance and slash cloud compute bills.
- Real-World Architectural Patterns: Apply production-tested operational patterns across hybrid and multi-cloud lakehouse environments.
Whether you are designing a greenfield Iceberg architecture or tuning existing lakehouse tables for high-concurrency BI and AI workloads, this workshop delivers the practical engineering strategies required to ingest efficiently and optimize at scale.