A retailer may process each purchase in milliseconds but still need hours to reconcile sales across stores, websites and returns. A data warehouse brings those records together for analysis. Its job is to answer large, repeatable questions reliably: which products make money, which customers return, and which numbers finance should trust. AI expands the audience for those answers while raising the cost of getting them wrong.
The investment question
The investment question is whether a warehouse becomes the enduring centre of analytical work or a replaceable engine that scans data held elsewhere. Suppliers earn money when customers process information, but performance improvements, open formats and competing cloud services can reduce spending on existing workloads.
Our view is that durable value comes from a combination of query performance, trusted data, operational simplicity and integration into decisions. Storage scale alone provides limited evidence of competitive advantage. The strongest platforms make many departments productive without forcing each department to build and maintain its own analytical infrastructure.
How a warehouse works
Warehouses are designed for analytical queries that scan and aggregate large datasets, rather than mainly updating individual application records. Column-oriented storage, compression, partitioning and query optimisation can reduce the work needed to answer these questions. Data modelling establishes consistent relationships between facts such as sales and dimensions such as customer, product or time.
Modern cloud architectures commonly separate storage from compute, allowing capacity to change without copying the entire dataset. Terminology needs care: a Snowflake virtual warehouse is a compute resource, not the whole data platform. Its warehouse documentation distinguishes resizing a cluster from adding clusters to handle concurrency, and describes automatic suspension of idle resources.
Ingestion, transformation, access policies and workload management surround the query engine. A fast query over incomplete data remains a bad answer. Equally, a perfectly modelled dataset is less useful if analysts queue behind a large overnight processing job. Successful platforms balance freshness, correctness, responsiveness and cost across different users.
Market structure and competitive advantage
| Competitive position | Main attraction | What customers must evaluate |
|---|---|---|
| Independent cloud data platform | Consistent experience across supported clouds | Integration depth and total consumption cost |
| Hyperscaler warehouse | Cloud identity, procurement and service integration | Portability and dependence on the wider cloud |
| Enterprise software ecosystem | Existing business applications and reporting relationships | Modernisation complexity and product overlap |
| Lakehouse analytical engine | Analysis close to open-format lake data | Operational simplicity and workload performance |
Snowflake, Google BigQuery, Amazon Redshift and Microsoft Fabric compete alongside Databricks and established enterprise platforms. Google’s BigQuery overview describes a managed analytics platform extending beyond traditional SQL reporting. Microsoft similarly positions Fabric Data Warehouse within its broader analytical environment. Competition increasingly spans the workflow around the engine.
Switching costs arise from SQL transformations, permissions, scheduling, dashboards, shared datasets and staff knowledge. A technically portable table is only one component of that system. However, customers can introduce a second engine for new workloads without immediately migrating everything, so incumbents must keep winning incremental spending.
Economics: efficiency can help customers and challenge revenue
Warehouse charges may depend on compute time, capacity reservations, data scanned, storage and additional services. The bill must be assessed against workload volume and required response time. A lower headline compute rate can be offset by inefficient queries, excessive data movement or substantial engineering overhead.
For illustration, suppose a workload needs 1,000 compute-hours at £4 per hour. Spending is £4,000. A faster engine that completes the same work in 600 hours at £5 costs £3,000. A higher hourly rate can therefore deliver a lower workload cost. This simplified example excludes storage, support, commitments and other charges; it is not a vendor comparison.
The supplier benefits if lower cost encourages additional analysis, applications or departments to adopt the platform. It loses near-term revenue if customers simply keep the savings. This tension explains why technical improvements can temporarily restrain consumption growth while improving the product’s competitive position.
Assess gross margin alongside cloud infrastructure costs and product mix. Some AI services have different underlying costs from established SQL workloads. Commitments and remaining performance obligations provide information about contracted demand, but recognition depends on contractual terms and usage patterns. They should not be treated as interchangeable with current consumption.
AI and hyperscalers: from dashboards to questions and actions
AI interfaces let users ask business questions without manually writing SQL. That can increase query demand, but it also exposes ambiguities that dashboards previously concealed. A request for revenue must specify currency, period, returns, recognition rules and access rights. A model needs a governed semantic context to produce a meaningful answer.
Warehouses can also supply training datasets, features, evaluation records and business context for retrieval or agents. Some processing can occur close to stored data, reducing unnecessary movement. Yet proximity does not eliminate inference costs, model evaluation or the need to restrict which records an agent can access.
Hyperscalers combine warehouses with model services, object storage, networking and enterprise contracts. Independent platforms can offer a consistent data experience across supported clouds and models. The competitive question is whether customers value that flexibility enough to offset the convenience and commercial benefits of buying an integrated cloud stack.
Current market debates — September 2026
The first debate is whether AI adds substantial paid work to the core analytical platform. Snowflake’s second-quarter fiscal 2027 release, published on 2 September 2026 for the quarter ended 31 July, reported product revenue of about $1.49 billion, up 37%, and net revenue retention of 126%. Management attributed momentum to both core and AI products. These disclosures do not isolate the revenue that would have occurred without AI.
The second debate is warehouse versus lakehouse. Open table formats and broader workload support make the labels less useful as a strict technical divide. Buyers increasingly ask which platform delivers governed SQL, engineering and AI workloads with the least friction. Vendors may win the same customer while serving different departments or processing stages.
The third debate concerns autonomous optimisation. Better query planning, workload sizing and data layout can lower bills and reduce operational effort. The bull case is that easier, cheaper analytics expands the market. The countercase is that customers consolidate workloads and maintain budget discipline, leaving suppliers with more activity but weaker revenue per unit of activity.
Structural debates: owning data versus owning the workflow
Open storage and table formats can reduce the cost of introducing another compute engine. That puts more pressure on proprietary execution, governance, collaboration and developer experience to justify a premium. Openness can also attract customers who would otherwise avoid committing valuable data to a platform.
The warehouse’s position as a trusted source of business information is potentially stronger than its role as a fast scanner of bytes. Metric definitions, policies and organisational adoption are difficult to replicate quickly. However, the semantic layer can also sit in an independent BI or modelling tool, limiting the warehouse vendor’s control over the full workflow.
There is a further tension between consumption revenue and customer confidence. Unpredictable bills can discourage experimentation precisely when vendors want more users. Budget controls, transparent attribution and reliable performance can therefore be competitive features, even when they constrain short-term consumption.
What to watch
Watch production workload additions, expansion among existing customers, query efficiency, concurrency and spending after optimisation. Evaluate customer outcomes using equivalent data, freshness, security and service levels. A benchmark with a warm cache or favourable query selection may say little about the customer’s complete workload.
For AI, follow repeat usage of governed analytical assistants, verified answer quality and measurable reductions in time to decision. Separate trials and enabled accounts from sustained paid consumption. The enduring opportunity is to become the trusted operating environment for business analysis as access broadens beyond specialist data teams.
Explore this sector
Data platforms & analytics — sector overview
Related sectors: Semiconductors & chipmaking · AI infrastructure & data centres