EXPLAINER · ACTIVE METADATA

What is active metadata, and what makes governance "active"?

Most governance artefacts are accurate on the day they are published. A definition gets agreed, written into a catalog, and then the business moves. A system is replaced, a rule is revised, the person who knew what the cryptic field meant leaves. Nothing in the catalog notices. Active metadata is the response to that: metadata re-derived from its sources rather than typed in once and maintained by hand.

This page explains what active metadata means, how it differs from a conventional catalog entry, what "active governance" does and does not cover, and where the term gets oversold.

Diagram following one catalog entry through three change events: a schema change, a revised definition, and the departure of its owner. A static catalog entry, typed once, goes from accurate to drifting, then stale, then orphaned, because nothing in the catalog notices the change. Active metadata re-derives the same entry from its sources at every event, so it stays current; each refresh is proposed for review rather than applied automatically. Agents propose. People decide.

Key takeaways

  • Active metadata is re-derived from its sources, rather than entered once and maintained by hand.
  • It also covers business logic, not just asset descriptions: rules, calculations, processes and ownership, where conventional metadata stops at names, types, tags and lineage.
  • The term is used loosely. Some vendors mean usage telemetry, some mean automated lineage, some mean re-extraction of meaning. They are different products, and it is worth asking which one you are being shown.
  • Automation moves the work; it does not remove the judgement. Re-extraction changes who spends time on what, not whether a human decides what is correct.

What active metadata actually means

Conventional metadata is a description someone typed. A table gets a name, an owner, a sensitivity label and a short explanation, and all of it is accurate at the moment of entry. Active metadata inverts the direction: instead of a person describing the asset, the system derives the description from the asset and the documents around it, and does it again on a schedule.

The second half of the term matters as much as the first. Metadata, strictly, is names, types, tags and lineage. That is enough to find a table and not nearly enough to use it correctly. The information that decides whether an answer is right lives one level up: the rule that says which transactions are excluded, the calculation behind a KPI, the process the field is written by, the role accountable for the definition. Calling that layer metadata stretches the word, which is why the parenthetical exists.

"Active" is contested, and the differences are practical. At least three products go by the name:

Usage-driven

Metadata enriched by query logs and access telemetry: who uses what, how often, which assets are abandoned. Useful for prioritisation, silent on meaning.

Automation-driven

Lineage and profiling computed rather than documented. Strong on technical structure, still silent on business logic.

Source-derived

Definitions, rules and mappings re-extracted from the documents and systems that describe them. This is the one that touches meaning, and the one this page is about.

They are complementary rather than competing, and most estates end up with more than one. Asking a vendor which of the three they mean will save a lot of time.

  • Conventional metadata is a description someone typed. Active metadata is a description something derived.
  • Names, types and tags find you an asset. Rules, calculations and ownership tell you whether you can trust it.
  • Three different products share the name. Usage telemetry, automated lineage, and source re-extraction.
  • Nothing here fixes data. It describes and governs it; remediation happens at the source.

Why governance coverage plateaus

The pattern is consistent enough to be predictable. A catalog goes in, a first wave of documentation gets done during the project, and coverage climbs. Then the project ends, and maintenance becomes an ongoing ask on domain experts who have a day job. Coverage flattens. The entries written during the push are never revisited, so accuracy decays quietly behind a number that still looks respectable.

This is an incentive problem before it is a tooling problem. The person who knows what a field means gets nothing from writing it down. The cost of the entry falls on them, and the benefit falls on someone in another function who will consume it a year later. Every governance programme runs into this, and no amount of workflow tooling changes the arithmetic.

Underneath the coverage number, a second thing is happening. The definitions and the data drift apart. A rule gets revised in a document that never reaches the catalog. A field is deprecated and a replacement appears. The catalog entry still reads correctly, because it describes a world that used to exist. Nothing flags the divergence, because nothing is checking.

Where the term gets oversold. "Active" is sometimes marketing for running a scanner on a schedule. Re-extraction narrows the gap between the description and the reality; it does not close it. Something still has to decide whether a proposed change is correct, and that something is a person who knows the domain. A vendor claiming governance without a review step is describing an unaccountable system: when the definition turns out to be wrong, no one approved it. The honest version of the pitch is narrower. Automation can carry the extraction, the comparison and the routing. It cannot carry the judgement.

When this is not the right tool

Not every estate needs this. If you run a single system, your definitions are stable, one team owns them, and nobody is disputing what a term means, a maintained catalog entry or a wiki page is proportionate and will cost you far less. Re-extraction earns its keep where the sources are many and disagree with each other, where a definition has more than one claimant, or where the gap between the document and the catalog entry goes unnoticed because nothing is checking. And if the pain you are feeling is pipeline reliability rather than meaning, an observability tool is the right purchase.

HOW METAGEM DOES IT

Re-extract. Compare. Route.

Step 1

Re-extract

Agents re-read the sources on a schedule instead of waiting for someone to update an entry.

  • Documents and systems, structured and unstructured
  • Configured per domain, not a blanket crawl

Step 2

Compare

What comes back is compared against what is approved, and divergence is surfaced with the evidence.

  • The old text, the new text, and the source
  • Conflicting definitions shown side by side

Step 3

Route

Differences reach the domain owner who can settle them, rather than a central backlog.

  • Approve, reject, or correct
  • Each decision recorded with who made it and when
Traditional governance is curated by committee, agreed in workshops and out of date before the catalog is adopted. Metagem inverts the order: agents do the extraction in the background and propose what changed, so experts spend their time on the conflicts and the judgement calls rather than on data entry.
MetagemWhat's different here

Where it shows up

Coverage that does not depend on goodwill

Extraction runs whether or not anyone has time this quarter, so the baseline keeps up rather than freezing on the day the project ended.

Divergence surfaced with evidence

When a re-extraction disagrees with an approved definition, both versions appear with the source each came from, so the reviewer is judging rather than investigating.

Ownership stays with the domain

Changes reach the people who know the answer instead of a central team who have to go and ask them.

Quality findings as a by-product

Conflicting definitions and broken rules surface while the context is being built, which is usually the first time anyone has seen them listed in one place.

Data Quality Audit

Frequently asked questions

A catalog is where the answers live, and most catalogs are populated by hand. The difference is the direction of maintenance: a conventional entry is written once by a person and decays from that moment, while an actively maintained one is re-derived from the source and re-proposed when the source changes. Metagem is a business catalog by shape. What differs is how it gets filled and how it stays current.

It is scheduled per source rather than continuous, and the cadence is a configuration choice: document repositories that change weekly do not need the same frequency as a schema that changes twice a year. There is a real cost to running extraction, so the sensible setting is tied to how fast a given source actually moves.

No, and the distinction matters. This tracks drift in meaning: definitions, rules, mappings and ownership. It does not profile data flows, raise operational alerts, or tell you a nightly load failed. Those are jobs for a data observability tool, and if unreliable pipelines are your actual problem then that is the category to look at, not this one.

Domain owners approve, reject or correct what the agents propose, and every decision is recorded with the person and the timestamp. Metagem writes back metadata only. It never modifies your business data. Where it finds a quality issue, it surfaces it with a suggested fix at the source; the correction is made by you or by a downstream tool.

See what has already drifted.

Point us at one domain and the sources that describe it. We re-extract, compare against whatever you have documented, and show you where the two disagree.

Talk to our team