decryptingtech

Technology. Business models. Market debates.

Browse this section

Data platforms & analytics

A company can own years of customer records and still struggle to answer a basic question consistently. Sales may use one definition of revenue, finance another, and an AI assistant may retrieve an outdated document with no understanding of either. Data platforms turn scattered records into information that applications and people can use. Their value depends on accuracy, context, access and operational reliability, not simply the amount of data stored.

The investment question

The investment question is which suppliers become essential as businesses use more data and AI. Databases, warehouses, lakehouses, pipelines, governance and business intelligence solve different problems, but vendors increasingly combine them. The resulting competition is about control of workloads and customer relationships as much as individual features.

Our view is that durable value accrues where a platform makes data dependable and useful while reducing operational work. AI increases the importance of fresh, well-governed information, yet it can also commoditise interfaces and intensify price competition. More data does not automatically produce more supplier revenue if processing becomes cheaper or customers consolidate tools.

How the data stack works

Operational databases record transactions and application state. Pipelines move or expose changes, transform data and coordinate processing. Warehouses organise analytical workloads, while lakehouses apply database-like management to data in lake storage. Governance defines ownership, meaning, permissions and evidence of how data has been used. Business intelligence turns governed information into metrics, analysis and decisions.

These are functional distinctions, not rigid product boundaries. A warehouse vendor can offer AI services, a lakehouse platform can add operational databases, and a cloud provider can bundle the entire workflow. Databricks’ Lakebase documentation describes a managed Postgres service integrated with its wider platform, illustrating the convergence between transactional applications and analytical data.

The architecture should follow the workload. Processing a payment requires different guarantees from scanning years of sales history. Searching documents for an AI assistant differs from calculating a financial metric. Combining tools can simplify operations, but using one product does not remove the need to understand these different requirements.

Market structure and competitive advantage

Subsector Core job Main competitive question
Databases Record and serve application state Reliability, performance and switching costs
Warehouses Run governed analytical queries Cost, concurrency and operational simplicity
Lakehouses Share managed lake data across workloads Interoperability and execution quality
Data pipelines Deliver timely, usable data Connector reliability and workflow ownership
Governance Establish meaning, access and accountability Coverage and enforceable policy
Business intelligence Support consistent decisions Trusted metrics and user adoption
The six subsectors describe customer needs; individual vendors increasingly operate across several of them.

Hyperscalers compete through integrated services and existing customer commitments. Independent platforms such as Snowflake and Databricks compete through cross-cloud reach, workload capabilities and a more specialised data experience. Database, integration, governance and BI specialists add further depth. An independent platform can also be a major customer of the cloud infrastructure provider with which it competes.

The strongest switching barriers are often operational: production queries, data models, permissions, engineering skills and business processes. Storing data in an open format can reduce one form of dependence without making the full application portable. The catalogue, processing engine, governance rules and surrounding workflow may still require substantial migration effort.

Economics: consumption and customer value

Revenue models include subscriptions, user licences, capacity commitments and consumption-based charges. Consumption can align spending with use, but makes near-term revenue sensitive to optimisation and workload timing. A vendor can improve performance for customers and initially reduce the resources billed for an unchanged workload.

Illustratively, if query volume grows 40% but the resources required per query fall 25%, consumption rises only 5% before price and mix effects. This is arithmetic, not a forecast. It explains why workload growth, data growth and revenue growth can diverge even when the product is becoming more useful.

The customer’s total cost includes infrastructure, vendor charges, engineering labour, duplicated data, maintenance and the consequences of bad information. A cheaper query engine may not be cheaper overall if it requires more operational effort. Conversely, a convenient integrated platform may become expensive if customers cannot control consumption or if unrelated workloads compete for shared capacity.

For investors, expansion within existing accounts is valuable but needs interpretation. Net retention can reflect new workloads, price changes and increased consumption, while also being affected by optimisation. Gross margins should be read alongside cloud infrastructure costs, support requirements and the economics of AI features. Adjusted profitability does not remove the importance of stock-based compensation and dilution.

AI and hyperscalers: the data advantage

Enterprise AI needs information beyond a model’s pretraining. Retrieval can supply relevant documents or records at inference time, while tools can query operational systems. Fine-tuning and other model-development workflows create additional data requirements. None of these approaches automatically resolves stale records, inconsistent definitions or inappropriate access.

This creates opportunities throughout the stack. Databases serve application and agent state; pipelines keep data fresh; warehouses and lakehouses support preparation and analysis; governance controls use; BI supplies trusted metrics and business context. The incremental revenue depends on which functions customers pay for separately and which become bundled features.

Hyperscalers have an advantage because data, compute and model services can be consumed in one environment. Microsoft Fabric’s architecture illustrates the integrated approach across ingestion, analytics, databases and reporting. Independent platforms must justify their additional layer through better operations, capabilities or flexibility across environments.

Current market debates — September 2026

The immediate debate is whether AI creates substantial new paid workloads or primarily improves existing data tools. Snowflake’s 2 September 2026 filing reported product revenue of $1.49 billion, up 37%, and management attributed acceleration to both its core platform and AI. The reported growth is evidence of business momentum; the attribution remains management’s assessment and does not reveal the economics of every AI feature.

Open table formats are another active competitive issue. Google’s Iceberg managed-table documentation describes managed analytics over an open-format foundation. This can improve interoperability, while shifting competition towards catalogues, query performance, governance and the services operating around the data.

The constructive independent-platform case is that customers want a consistent data layer across complex environments. The hyperscaler case is that integrated services reduce procurement and operational friction. Evidence should come from production workload wins, customer expansion and total cost, rather than simply counting features that appear on competing product pages.

Structural debates: openness, consolidation and trust

The first structural debate concerns where dependence moves as storage formats open up. Customers may gain the ability to read the same files with several engines while continuing to rely on one catalogue or policy system. True portability requires compatible semantics, permissions, performance and operating procedures as well as readable data.

Second, consolidation can reduce integration work while weakening specialist vendors. A platform that bundles adequate pipelines, governance and BI may capture spending previously spread across several suppliers. Specialists remain attractive where the customer’s needs exceed bundled capabilities or span many platforms. The outcome depends on measurable workflow value, not a universal preference for fewer tools.

Third, AI increases the cost of poor data context. A confident answer based on the wrong definition can be more damaging than a failed query that clearly stops. Trusted metrics, lineage and access enforcement therefore become part of the product experience. The winning interface may be conversational, but the underlying business logic still has to be correct.

What to watch

Track production workloads, consumption after optimisation, customer expansion, gross margins and platform consolidation. Examine whether open-format support works across actual engines and policies. Distinguish experimental AI usage from repeated, paid business processes.

This section follows the path from a recorded event to a trusted decision. Durable supplier value comes from making that path reliable, economical and easier to operate as the number of applications and users grows.

Explore this sector

Data platforms & analytics — sector overview

Related sectors: Semiconductors & chipmaking · AI infrastructure & data centres