A data lake can hold years of transactions, documents and machine logs cheaply, yet still be difficult to use reliably. A lakehouse adds the structures needed to manage and analyse that information: tables, transactions, metadata, access controls and efficient computation. The ambition is to support engineering, analytics and AI around shared data while reducing the number of copies and disconnected systems a company must maintain.
The investment question
The investment question is who captures value when customers keep data in relatively open storage but buy sophisticated services around it. The storage provider, query engine, catalogue, governance layer and application platform can all earn revenue. They do not necessarily have equal bargaining power.
Our view is that openness expands the addressable market but makes operational execution more important. A supplier cannot rely solely on customers being unable to read their data elsewhere. It must offer better performance, reliability, development workflows or governance. The strongest proposition combines flexibility with a managed experience that customers can operate economically.
How a lakehouse works
Object storage holds files. A file format such as Parquet describes how values are organised within a file. A table format coordinates the files and metadata that represent a table over time. Apache Iceberg’s project documentation describes capabilities including schema evolution and reliable table changes. These are distinct from the query engine that executes SQL against the table.
Delta Lake likewise provides a storage layer supporting transactional reliability and batch and streaming processing. A lakehouse platform combines such foundations with compute, orchestration, security and user tools. Choosing a table format does not by itself select a complete platform or guarantee that every engine supports every feature identically.
A catalogue helps engines find tables and interpret their metadata. In commercial platforms, catalogue services may also coordinate permissions, discovery and lineage. Data maintenance remains necessary: small files, outdated snapshots and unsuitable layouts can increase costs. Managed automation can conceal that work from users, but it does not make the underlying work disappear.
Market structure and competitive advantage
| Layer | Core responsibility | Potential source of value |
|---|---|---|
| Object storage | Durable storage of data and metadata | Scale, resilience and cloud integration |
| Table format and catalogue | Consistent table state and discovery | Interoperability, coordination and governance |
| Compute engines | SQL, transformation, streaming and machine learning | Performance and workload efficiency |
| Managed platform | Integrated development and operations | Productivity, trust and reduced complexity |
Databricks is a prominent integrated platform, while Snowflake and hyperscaler products increasingly support open lake data alongside other storage approaches. Independent engines and open-source projects offer additional choices. The market is therefore broader than a two-vendor contest, and many customers use several platforms for distinct workloads.
Competitive advantage can come from the optimiser, runtime, workload scheduling and management tools. Databricks’ Unity Catalog documentation also illustrates the importance of unified discovery, permissions and lineage. Such controls become valuable when they are consistently used across production workloads, rather than merely available as features.
Economics: cheap storage is only the starting point
The complete bill includes object storage, compute, orchestration, table maintenance, networking, governance and engineering. Retaining more raw data can be inexpensive relative to repeatedly processing it. Poor layouts, unnecessary refreshes and duplicated transformations can overwhelm savings from a lower storage price.
Consider an illustrative organisation saving £20,000 annually by removing duplicate storage while adding £30,000 of compute and operational work to support the new architecture. Its direct saving has become a £10,000 cost increase. Conversely, if the shared platform removes £80,000 of repeated engineering work, consolidation can be worthwhile. These figures demonstrate the accounting logic, not typical project economics.
Serverless services shift capacity management to the provider. Customers still need cost attribution, workload priorities and controls over expensive jobs. A service can feel operationally simple while being financially unpredictable if many teams or agents submit work without budget boundaries.
Supplier margins depend on the spread between customer charges and the infrastructure and support needed to execute workloads. Faster engines can increase that spread, reduce customer bills, or do both. Analysts should distinguish revenue growth driven by more valuable applications from growth driven mainly by inefficient processing or temporary migration duplication.
AI and hyperscalers: shared data with different requirements
Lakehouses are attractive for AI because engineering, training preparation, evaluation and analytics can use related datasets. Large volumes of text, images and logs can remain in object storage while tables organise associated metadata and structured results. Reusing governed information can reduce reconciliation between separate environments.
However, a training dataset and an online application have different requirements. Training may tolerate a fixed snapshot; an agent checking inventory needs fresh operational state. A lakehouse does not automatically replace a low-latency transactional database or a specialised serving system. Production architecture must connect these components with explicit freshness and access guarantees.
Hyperscalers benefit from the object storage and compute consumed beneath many lakehouse products. They also compete with the software platforms using that infrastructure. An independent vendor can therefore be both a major cloud customer and a rival for the enterprise relationship. Commercial partnerships do not remove that underlying tension.
Current market debates — September 2026
The first debate is whether open table support will make compute engines interchangeable. Snowflake’s Iceberg documentation describes multiple catalogue arrangements and associated capabilities. Google’s managed Iceberg tables in BigQuery combine open-format data with managed maintenance. This convergence strengthens customer choice, but read support, write support, policy enforcement and feature compatibility still require separate evaluation.
The second debate is whether platform breadth improves economics. Adding SQL, machine learning, application development and governance can let vendors capture more of a customer’s spending. It can also create a complex buying proposition, overlapping products and larger operational responsibilities. Cross-selling is valuable only when the integrated product reduces friction or solves an additional paid problem.
The third debate is how much control shifts to the catalogue. If multiple engines can access the same tables, the service defining discovery, permissions and trusted metadata becomes strategically important. Yet catalogue openness and federation can also limit that control. The investment question is which services remain indispensable when customers exercise their theoretical freedom to use another engine.
Structural debates: openness has several dimensions
Open files, open table specifications, accessible catalogue APIs and portable application logic are different forms of openness. A platform may support some well while retaining proprietary dependencies elsewhere. A credible portability assessment asks whether a customer can move permissions, scheduled jobs, notebooks, semantic definitions and operational procedures, not just query a sample table.
There is also a trade-off between assembling best-of-breed components and buying a managed platform. Assembly can improve choice and reduce selected licence costs, but the customer assumes more integration and incident responsibility. A commercial platform earns its premium when its automation and support cost less than the complexity they replace.
Finally, consolidation can reduce duplicated data while concentrating operational risk. Shared compute, metadata or policy failures can affect many teams. Isolation, recovery and clear ownership therefore become more important as a lakehouse becomes central to the business. A broad architecture needs disciplined boundaries as well as shared services.
What to watch
Watch active production workloads, cost per completed job, data maintenance overhead and evidence of successful interoperability. Test whether different engines produce consistent results under the required permissions and table features. Customer references are most useful when they describe ongoing operations after migration, including staffing and incident handling.
For AI, follow the conversion of experiments into repeatable training, evaluation or serving workflows, and the resulting consumption and retention. The durable opportunity lies in making diverse data useful with fewer operational burdens. Open storage creates the foundation; dependable execution determines which suppliers capture the value above it.
Explore this sector
Data platforms & analytics — sector overview
Related sectors: Semiconductors & chipmaking · AI infrastructure & data centres