Two teams can use the same dataset and disagree about what it means. An AI assistant can retrieve a technically accurate document that its user should never have seen. Data governance addresses these problems by establishing ownership, definitions, quality expectations, access rules and evidence of use. It connects the technical handling of information with the business decisions about who can trust it and for what purpose.
The investment question
The investment question is whether governance becomes an operational necessity with recurring value, or remains an expensive documentation project that employees work around. AI increases the consequences of poor data handling, but a larger theoretical risk does not automatically translate into a successful software business.
Our view is that durable suppliers connect discovery and policy to everyday workflows and enforceable controls. A catalogue full of stale descriptions has limited value. A platform that helps people find approved information, obtain appropriate access and understand downstream consequences can reduce delays while making the organisation more accountable.
How governance works
A data catalogue inventories assets and records metadata: where information resides, who owns it, what it means and how it is used. A business glossary defines terms such as customer or revenue. Classification identifies sensitive or important information. Quality checks test expectations about completeness, consistency or validity, while stewardship assigns responsibility for resolving problems.
Lineage records how data flows through systems and transformations. OpenLineage provides an open framework for collecting metadata about jobs, runs and datasets. Lineage can help explain which reports or models may be affected by a change. Its usefulness depends on coverage and detail; a partial technical graph is not a complete account of business meaning.
Policy management and policy enforcement are related but distinct. Recording that a dataset is restricted does not necessarily prevent unauthorised queries. Enforcement must occur through the relevant database, storage, application or access service. Governance products need clear integrations, and organisations need explicit responsibility for gaps between the documented rule and actual behaviour.
Market structure and competitive advantage
| Product position | Main responsibility | Source of differentiation |
|---|---|---|
| Enterprise catalogue and governance | Shared definitions, discovery and stewardship | Cross-system coverage and organisational adoption |
| Platform-native governance | Controls within a data environment | Integration with actual execution and permissions |
| Data quality and observability | Detect defects and operational problems | Useful detection, diagnosis and remediation |
| AI governance | Inventory, oversight and evidence for AI use | Connection between policy, models, agents and runtime activity |
Specialist products such as Collibra compete alongside broader data-management platforms and native cloud capabilities. Microsoft Purview describes an enterprise data-governance environment, while AWS Lake Formation manages access to supported lake data and analytical integrations. Their scope and enforcement boundaries should be evaluated directly rather than assumed to be identical.
The moat can come from integrations, accumulated metadata, business definitions and embedded approval processes. However, those assets become liabilities if they require extensive manual maintenance. Successful products keep metadata connected to changing systems and make useful governance easier than bypassing it.
Economics: measure saved effort and improved execution
Governance spending includes software, implementation, integration, stewardship and ongoing remediation. Subscription pricing may depend on users, assets, capacity or product modules. A software-only comparison can understate the organisational work required to define ownership and resolve disagreements about data.
For illustration, suppose 200 analysts each save one hour weekly searching for trusted datasets. At an assumed £50 per hour over 45 working weeks, the time has a nominal value of £450,000 annually. This is potential capacity released, not automatically a cash saving. Real value depends on adoption, what people do with the time and the full cost of the programme.
Avoided incidents can be valuable but are difficult to attribute precisely. A credible business case should include observable improvements such as faster access approval, fewer repeated reconciliations, quicker change-impact analysis and reduced time to resolve quality defects. Broad claims about eliminating all risk are neither measurable nor operationally realistic.
For suppliers, services intensity matters. A platform that needs large bespoke deployments may grow revenue while struggling to scale efficiently. High retention can reflect deep usefulness, but it can also reflect expensive implementation and customer inertia. Renewal quality is stronger when users actively depend on the product and expand its operational coverage.
AI and hyperscalers: permissions and meaning become machine inputs
AI assistants need more than access to raw data. They need definitions, provenance, freshness and restrictions on permitted use. An agent that can act across systems also requires an identity, bounded permissions and records of the actions it takes. Governance must connect those requirements to the systems that actually serve data and execute work.
Retrieval introduces derived copies such as chunks and embeddings. Changes to source permissions, corrections or deletion need to reach these copies. A model response may combine several sources, making provenance and evaluation more complex than a traditional dashboard. Good metadata can help, but it does not by itself guarantee a correct or appropriate answer.
Hyperscalers have an advantage inside their own environments because identity, storage and compute are closely integrated. Enterprises spanning several clouds and business applications may still need a broader view. The independent vendor’s opportunity is to coordinate meaning and evidence across those boundaries without pretending that every platform enforces policy in the same way.
Current market debates — September 2026
The first debate is whether governance platforms can extend from documenting AI projects to overseeing operating agents. Collibra’s current AI Governance documentation describes a central repository for agents, models and use cases. The commercial opportunity is broader oversight; the practical test is whether inventory stays current and connects to actionable controls and evidence. Registration alone is not runtime enforcement.
The second debate is consolidation into enterprise application ecosystems. Salesforce completed its acquisition of Informatica on 18 November 2025. This combines a major data-management business with a large enterprise application platform. Integration may improve distribution and AI workflows, while customers still need to assess cross-platform support and the independence of their data architecture.
The third debate is automated metadata generation. AI can draft descriptions, propose classifications and identify relationships, reducing repetitive work. Yet plausible descriptions can be wrong, and incorrect labels can propagate widely. Durable value requires review mechanisms, confidence signals, ownership and a way to correct the underlying catalogue when business meaning changes.
Structural debates: central standards and distributed ownership
Central governance can establish common definitions and minimum controls, but central teams rarely understand every dataset as well as its business owners. Distributed ownership brings expertise closer to the data while risking inconsistency. Effective operating models combine shared standards with named local responsibility and escalation paths for disagreements.
There is also a trade-off between a neutral enterprise layer and native enforcement. A broad catalogue can see across systems but may depend on connectors and local policy engines. A native platform can enforce controls deeply within its environment but provide less complete coverage elsewhere. Customers may need both, with explicit division of responsibility.
Finally, governance has to support useful access. Excessive friction encourages unofficial copies and workarounds that reduce visibility. Well-designed approval workflows, clear classifications and reliable self-service can improve both productivity and control. This makes user experience an economic requirement, rather than a cosmetic addition to a compliance-oriented product.
What to watch
Watch active usage, metadata freshness, ownership coverage, approval times and the share of critical data with functioning quality checks and traceable lineage. Evaluate enforcement through practical access tests. Catalogue size is a weak success measure when assets are stale, duplicated or disconnected from business use.
For AI, track registered production systems, permission propagation, evidence of model and agent activity, and the time required to investigate an incorrect answer or action. The strongest governance businesses make trusted data easier to use at scale while preserving clear accountability for how it is used.
Explore this sector
Data platforms & analytics — sector overview
Related sectors: Semiconductors & chipmaking · AI infrastructure & data centres