Apache Iceberg Lakehouse Engineering
Service description
I build open data lakehouses on Apache Iceberg so your analytics gets real warehouse behavior directly on object storage, without a proprietary engine and without locking your data to one vendor. I design the table layout, set up the catalog, and turn raw Parquet files into governed, versioned tables that any engine can read. The goal is simple: ACID transactions, reliable schema evolution, and predictable performance while your data stays in open formats you fully own.
We start from your current state — raw files, a Hive mess, or an expensive warehouse you want to unwind. I model the tables, choose partitioning and sort orders that match how you really query, and enable hidden partitioning so analysts never handle partition columns by hand. I configure schema evolution so adding or renaming columns never rewrites history, and I set up time travel and snapshot rollback so a bad load is a one-line fix, not an incident. On the engine side I integrate Spark for batch and maintenance, Trino for interactive SQL, and Flink for streaming — all on the same tables through a shared catalog. I document the design and coach your team to run it without me.
— Iceberg table design, partitioning and sort strategy
— Catalog setup, migration from Hive or raw Parquet
— Schema evolution, time travel and snapshot rollback
— Spark, Trino and Flink integration on shared tables
We start from your current state — raw files, a Hive mess, or an expensive warehouse you want to unwind. I model the tables, choose partitioning and sort orders that match how you really query, and enable hidden partitioning so analysts never handle partition columns by hand. I configure schema evolution so adding or renaming columns never rewrites history, and I set up time travel and snapshot rollback so a bad load is a one-line fix, not an incident. On the engine side I integrate Spark for batch and maintenance, Trino for interactive SQL, and Flink for streaming — all on the same tables through a shared catalog. I document the design and coach your team to run it without me.
— Iceberg table design, partitioning and sort strategy
— Catalog setup, migration from Hive or raw Parquet
— Schema evolution, time travel and snapshot rollback
— Spark, Trino and Flink integration on shared tables
Contact the freelancer
Order the service or ask the freelancer a question.
Freelancer contacts
E-mailShow
