Most governance artefacts are accurate on the day they are published. A definition gets agreed, written into a catalog, and then the business moves. A system is replaced, a rule is revised, the person who knew what the cryptic field meant leaves. Nothing in the catalog notices. Active metadata is the response to that: metadata re-derived from its sources rather than typed in once and maintained by hand.
This page explains what active metadata means, how it differs from a conventional catalog entry, what "active governance" does and does not cover, and where the term gets oversold.
Conventional metadata is a description someone typed. A table gets a name, an owner, a sensitivity label and a short explanation, and all of it is accurate at the moment of entry. Active metadata inverts the direction: instead of a person describing the asset, the system derives the description from the asset and the documents around it, and does it again on a schedule.
The second half of the term matters as much as the first. Metadata, strictly, is names, types, tags and lineage. That is enough to find a table and not nearly enough to use it correctly. The information that decides whether an answer is right lives one level up: the rule that says which transactions are excluded, the calculation behind a KPI, the process the field is written by, the role accountable for the definition. Calling that layer metadata stretches the word, which is why the parenthetical exists.
"Active" is contested, and the differences are practical. At least three products go by the name:
Metadata enriched by query logs and access telemetry: who uses what, how often, which assets are abandoned. Useful for prioritisation, silent on meaning.
Lineage and profiling computed rather than documented. Strong on technical structure, still silent on business logic.
Definitions, rules and mappings re-extracted from the documents and systems that describe them. This is the one that touches meaning, and the one this page is about.
They are complementary rather than competing, and most estates end up with more than one. Asking a vendor which of the three they mean will save a lot of time.
The pattern is consistent enough to be predictable. A catalog goes in, a first wave of documentation gets done during the project, and coverage climbs. Then the project ends, and maintenance becomes an ongoing ask on domain experts who have a day job. Coverage flattens. The entries written during the push are never revisited, so accuracy decays quietly behind a number that still looks respectable.
This is an incentive problem before it is a tooling problem. The person who knows what a field means gets nothing from writing it down. The cost of the entry falls on them, and the benefit falls on someone in another function who will consume it a year later. Every governance programme runs into this, and no amount of workflow tooling changes the arithmetic.
Underneath the coverage number, a second thing is happening. The definitions and the data drift apart. A rule gets revised in a document that never reaches the catalog. A field is deprecated and a replacement appears. The catalog entry still reads correctly, because it describes a world that used to exist. Nothing flags the divergence, because nothing is checking.
Where the term gets oversold. "Active" is sometimes marketing for running a scanner on a schedule. Re-extraction narrows the gap between the description and the reality; it does not close it. Something still has to decide whether a proposed change is correct, and that something is a person who knows the domain. A vendor claiming governance without a review step is describing an unaccountable system: when the definition turns out to be wrong, no one approved it. The honest version of the pitch is narrower. Automation can carry the extraction, the comparison and the routing. It cannot carry the judgement.
Not every estate needs this. If you run a single system, your definitions are stable, one team owns them, and nobody is disputing what a term means, a maintained catalog entry or a wiki page is proportionate and will cost you far less. Re-extraction earns its keep where the sources are many and disagree with each other, where a definition has more than one claimant, or where the gap between the document and the catalog entry goes unnoticed because nothing is checking. And if the pain you are feeling is pipeline reliability rather than meaning, an observability tool is the right purchase.
Step 1
Agents re-read the sources on a schedule instead of waiting for someone to update an entry.
Step 2
What comes back is compared against what is approved, and divergence is surfaced with the evidence.
Step 3
Differences reach the domain owner who can settle them, rather than a central backlog.
Traditional governance is curated by committee, agreed in workshops and out of date before the catalog is adopted. Metagem inverts the order: agents do the extraction in the background and propose what changed, so experts spend their time on the conflicts and the judgement calls rather than on data entry.
Extraction runs whether or not anyone has time this quarter, so the baseline keeps up rather than freezing on the day the project ended.
When a re-extraction disagrees with an approved definition, both versions appear with the source each came from, so the reviewer is judging rather than investigating.
Changes reach the people who know the answer instead of a central team who have to go and ask them.
Conflicting definitions and broken rules surface while the context is being built, which is usually the first time anyone has seen them listed in one place.
Data Quality AuditPoint us at one domain and the sources that describe it. We re-extract, compare against whatever you have documented, and show you where the two disagree.
Talk to our team