Ask how a number reached the board deck, and the answer arrives as a graph: nodes, edges, column-level arrows tracing the figure back through pipelines to source. It looks like evidence. It is actually an estimate — every lineage graph in this category is a reconstruction, assembled by crawlers that parse what they can parse and infer the rest. The difference between a reconstruction and a record seems academic until an auditor, a regulator, or an AI incident asks you to swear by it.
The map got genuinely good
Start with the credit, because it is owed. Automated, column-level lineage across a sprawling, heterogeneous estate is one of the genuinely hard problems the metadata platforms have made real progress on. Crawlers that parse many SQL dialects, read pipeline definitions and job configs, unpack BI-tool models, and stitch the inferred relationships into a navigable end-to-end graph represent serious, hard-won engineering. A question that used to be tribal knowledge — what breaks if we deprecate this column? — became a query.
And the category keeps compounding the capability: more parsers, broader source coverage, shorter scan cycles, graphs that refresh closer and closer to the moment pipelines change. For impact analysis, migration planning, and root-causing a broken dashboard, crawled lineage is materially useful. None of what follows argues otherwise.
The method defines the boundary
But look at how the graph is made, because the method defines the boundary. Every lineage crawler in the category runs the same architecture: connect to sources, parse the artifacts it finds, infer relationships between them, and assemble the edges into a map. Three properties follow from that method — structurally, regardless of any vendor's roadmap:
- Coverage is bounded by parsers. A crawler can only see the paths it has a parser for. The scheduled pipeline is visible; the analyst's notebook, the stored procedure in an unusual dialect, the one-off script, the manual export into a spreadsheet, the engine a team adopted last quarter — invisible until someone ships the parser. And in the graph, silence about a path looks identical to no data flows here.
- Freshness is bounded by the scan. The graph shows the estate as of the last crawl. Between scans, pipelines change and run; whatever executed in that window gets reconstructed later — or never.
- Edges are inferred, not witnessed. A column-level edge through dynamic SQL or templated code is a probability ranked by heuristics, and the graph rarely distinguishes what it observed from what it guessed.
None of this is negligence. It is the physics of standing beside a runtime and reconstructing what it did after the fact. A survey can be superb. It is still a survey of a territory that changes beneath it — and the territory does not file a change request when it grows a new road.
The agent trace inherits the gap
The category's AI-era pitch extends the same graph toward models and agents: trace the answer back to its sources, show which governed tables fed the response. The ambition is exactly right — provenance for AI is about to be a board-level demand. But the extension inherits the architecture, so the trace covers what the crawler could see. The retrieval pipeline that filled an agent's context window is a data movement. So is the embedding job, the cached context, the file an agent pulled from a share, the answer pasted into a downstream system. These are the newest, fastest-changing paths in the enterprise, they are frequently not SQL at all, and they are precisely the roads the map has never heard of.
It is worth being precise about the strongest architecture in this category, because it deserves respect. Query-time access control — attribute-based policy, row- and column-level security, dynamic masking applied inside the data platforms they front, with every query authorized and logged — is real enforcement of the read, not documentation. Its query log is genuinely witnessed provenance of access. But it is provenance of reads, at the doors it fronts. The moment data leaves that perimeter — an extract, a cache, a context window — the witnessed record ends. And the business action the data went on to drive — the posting, the approval, the record change — was never in the log's scope at all. The read was witnessed. The action never was.
Emission: the territory reports itself
There is a second way to get lineage, and it is not a better crawler. In Entroid's architecture, the Semantic Ontology is not a catalog beside the data — it is the governed model the runtime executes on. Every process on the Composable Process Fabric is composed from five primitives — Deterministic Workflows, Intelligence Orchestration, Atomic Agents, Functions, and Connectors — and every data action any of them takes passes through the same inline gate: classification, entitlement, quality, and residency evaluated at the moment of the read, the write, or the act.
That placement changes what lineage is. When the gate passes an action, the runtime writes an immutable per-action record — the input, the actor (human or agent), the policy that was evaluated, and the output — at execution time, as a byproduct of executing at all. This is an architectural property of the design, not a feature claim: there is no parser, because nothing is being reconstructed; there is no scan interval, because there is no scan; there is no inferred edge, because every edge is the action that created it. Completeness does not depend on parser coverage. On the fabric, provenance is not a map of the territory — it is the territory reporting itself, action by action.
The distinction is survey versus telemetry. A crawled graph answers: how does data probably move around here? An emitted record answers: what did this specific action do, to which data, under which policy, on whose authority? The first is cartography. The second is testimony.
Governed execution does not abolish disagreement
Now the fair print. The by-construction claim applies to actions on the fabric. Entroid runs over the estate you already have — Connectors are the only primitive that touches external systems, and a Connector's action is gated and recorded like any other — but a job that never touches the fabric emits nothing, for the same reason no architecture can testify to work it did not execute. If your critical flows live entirely outside any governed runtime, an emission architecture does not retroactively illuminate them.
What changes is the direction of travel. Closing a crawler's gap means writing parsers forever, chasing an estate that adds engines, notebooks, and agents faster than parsers can ship. Closing a runtime's gap means moving processes onto the fabric — and every process moved converts a probabilistic path into a witnessed one, permanently. Notice what the unwitnessed paths are: the script, the export, the side-channel feeding an agent — the same ungoverned routes an enterprise should be narrowing for a dozen other reasons. During the transition, the two approaches coexist: keep the crawled map for the surrounding estate. Just be deliberate about which of the two you would hand a regulator as evidence.
Four questions that separate map from territory
When a vendor presents automated, column-level, real-time, end-to-end lineage, four questions locate the architectural boundary quickly:
- Coverage. Which movement paths in our estate have no parser today — notebooks, scripts, manual exports, retrieval pipelines feeding agent context — and how much decision-critical data travels them?
- Freshness. What is the lag between a pipeline changing and the graph reflecting it — and what executed inside that window?
- Provenance of the edge. For a given column-level edge, can the tool show whether the relationship was observed or inferred, and how the two are distinguished?
- Evidentiary weight. When the auditor asks how this number was produced, are we presenting a record of execution or a well-researched estimate?
For discovery and impact analysis, an estimate is enough — and the category provides a good one. For governing what data actually did — the question auditors, regulators, and AI-governance frameworks are converging on — only a record will do. And a record of execution can only come from the thing that executed.
Crawled lineage is a map that chases the territory. Emitted lineage is the territory testifying to what it did.
See what this looks like for your enterprise.
Not a demo. A strategic conversation about how your enterprise could operate
when every process runs on one governed fabric.
