Skip to content

What is Data Lakehouse?

Data & Analytics, explained by the engineers who build it. Definition, how it works, use cases and common questions.

Data Lakehouse definition

A data lakehouse is a data architecture that combines the low-cost, flexible storage of a data lake with the reliability, performance and governance features of a data warehouse. It stores data in open file formats on object storage and adds a table layer, such as Delta Lake, Apache Iceberg or Apache Hudi, providing ACID transactions, schemas and fast SQL.

How does a data lakehouse work?

A lakehouse keeps data as files, typically Parquet, in cloud object storage. An open table format sits on top and tracks which files belong to each table version in a transaction log or metadata tree. This allows multiple engines to read and write safely, with ACID transactions, schema enforcement and evolution, time travel to earlier versions, and efficient updates and deletes, which are hard to achieve with plain files in a traditional data lake.

Query engines such as Databricks SQL, Spark, Trino, Snowflake, BigQuery and Amazon Athena can read these tables directly. A governance layer, such as Unity Catalog or a similar catalog service, manages permissions, lineage and discovery, so the same storage supports BI dashboards, SQL analytics, data science notebooks and machine learning training without copying data between separate systems. That single copy simplifies security reviews too.

The medallion architecture

Lakehouses commonly organize data into three quality layers, an approach popularized by Databricks as the medallion architecture. Each layer builds on the previous one, and transformations between layers are automated, tested and monitored so problems are caught early. The names matter less than the discipline of separating raw, cleaned and business-ready data. Gold tables should map to real business questions.

  • Bronze: raw data ingested from sources, kept for replay and audit.
  • Silver: cleaned, validated and conformed data joined across sources.
  • Gold: aggregated, business-level tables for reporting and ML features.

Benefits of a data lakehouse

The main benefit is one platform instead of two. Organizations that previously copied data from a lake into a separate warehouse maintained duplicate pipelines, inconsistent numbers and two sets of costs. A lakehouse reduces that duplication. Open formats also reduce lock-in, because the same tables can be read by different engines, and storage stays on inexpensive object storage. Data science and BI teams work from the same governed data, which shortens the path from analysis to production models.

Trade-offs and challenges

Lakehouses require solid data engineering. Teams must manage file sizes, compaction, partitioning, table maintenance and catalog configuration, which managed warehouses largely hide. For small, mostly structured datasets used only for BI, a conventional cloud warehouse may be simpler and cheaper to operate. Choosing between table formats and catalogs also matters, though interoperability between Delta Lake and Iceberg has improved as vendors have converged on open standards. Budget time for table maintenance jobs from the start.

Example of a lakehouse

An ecommerce company ingests clickstream events, orders and product catalog data into bronze Delta tables. Silver tables join sessions to orders and clean product attributes. Gold tables feed revenue dashboards in Power BI and a recommendation model trained in the same platform, both using identical customer definitions. Nexzem designs lakehouse platforms on Databricks and open table formats when clients need analytics and machine learning on the same governed data. Adding a new gold table takes hours rather than a new pipeline project.

Data Lakehouse: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What is the difference between a data lakehouse and a data warehouse?

A data warehouse traditionally stores data in a proprietary format managed entirely by the warehouse engine. A lakehouse stores data in open file and table formats on object storage, readable by many engines, and supports unstructured data and machine learning alongside SQL. The gap is narrowing as warehouses add support for open table formats.

Is Databricks a data lakehouse?

Databricks is the company most associated with the lakehouse concept and offers a lakehouse platform built on Delta Lake, Apache Spark and Unity Catalog. Other vendors, including Snowflake, Google, AWS and Microsoft, also support lakehouse architectures, particularly through Apache Iceberg and their own storage and catalog services.

What is the difference between Delta Lake and Apache Iceberg?

Both are open table formats that add ACID transactions, schema evolution and time travel to files in object storage. Delta Lake originated at Databricks and is tightly integrated with Spark. Iceberg originated at Netflix and has broad multi-engine support. Both are widely adopted, and tools increasingly allow reading one format through the other.

Keep exploring the data & analytics glossary

Need Data Lakehouse in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.