Your visual-inspection model can now show you exactly which pixels it read. It can box the defect, highlight the extracted field on the invoice, cite the frame in the video feed. That is provenance of the read — and it is genuinely valuable. But when the auditor, the regulator, or the post-incident review arrives, the question is never only what the model saw. It is what your enterprise did about it — and this is precisely where most vision stacks fall silent.
The read has genuinely gotten good
Let's concede the real progress first, because pretending otherwise is the fastest way to lose a technical reader. The modern computer-vision platforms have democratized a discipline that used to require a research team. Expert-guided labeling and rapid model training put a working detector in the hands of a domain expert who has never written a training loop. That is a legitimate shift, and it is not marketing.
The read has become more trustworthy too. Pixel-level grounding and source citations — a bounding box tying an extracted value to its exact coordinates on the page, a highlighted region proving where a classification came from — make the model's output verifiable rather than asking you to trust a number in isolation. Done well, this measurably cuts the hallucinated field, the confidently-wrong extraction. Pair it with confidence scores and human-in-the-loop review, and you have a pragmatic middle path: route the low-confidence read to a person, auto-accept the rest. Add edge deployment and enterprise security posture — SOC 2, HIPAA, zero-data-retention — and you have real assurance that the pixels never left the building and the model behaved.
All of this is real. And all of it governs the same two things: the model (its accuracy, its drift, the quality of its labels) and the read (the provenance of the pixel). It is a complete story about perception. It is not a story about consequence.
The audit that stops at the pixel
Here is the uncomfortable part. Two of the loudest claims in this category are both true and both incomplete — in the same way.
The document-AI platforms market field-level citations and pixel-level grounding as, in their framing, an "auditable trail for compliance." And it is a trail — of the read. It proves that the total on line 14 was extracted from these exact coordinates on this exact page. Separately, the managed cloud detection and extraction endpoints offer their own "audit": an API-call log. It proves that a request went in and a JSON response came back with a bounding box and a confidence score.
Notice what neither one records. The citation proves where the value sat. The API log proves the model was called. Neither says a word about what happened next — whether a payment posted, a batch was quarantined, a badge was flagged, a claim was approved. Both prove the read with real rigor and are structurally silent on the act. An audit that ends at the pixel is an audit of perception dressed up as an audit of decisions. When your controls, your regulator, and your liability all live downstream of the detection, that is the wrong ledger.
Where the liability actually lives
The compliance question is not "was the detection correct?" It is a chain: which action fired, under what policy, approved by whom, and with what result? That chain is the thing an auditor reconstructs, the thing a plaintiff subpoenas, the thing an incident review dissects. And in the standard architecture, that chain is authored somewhere the vision system cannot see.
The detection is returned — as an endpoint response or an edge event — and thrown over a wall into a separate system that takes the consequential action: a PLC that stops a line, an MES that dispositions a lot, an ERP that posts an entry, an RPA bot that files a claim. To be fair, the more sophisticated vendors do close this loop on paper: they write the detection into a PLC tag, or pipe the structured fields straight into an MES, ERP, or RPA workflow. That is genuinely more end-to-end than a bare detector that just emits a number, and it deserves credit.
But be precise about what "closing the loop" actually is here. The decision-and-action logic — the confidence threshold that gates the stop, the segregation-of-duties rule on the posting, the approval that must precede the disposition — is authored in a different tool, by different people, in a different governance regime. No policy travels across the wire; only a value does. And the record splits in two: the vision system logs what the camera saw, the second system logs what it did, and nothing binds them into one accountable transaction. That split-brain audit — two half-records that have to be reconciled after the fact, if anyone can — is not a tooling gap you patch. It is the shape of the architecture. Detection is governed as a model on one side of the wall; the action executes ungoverned on the other.
The record that binds saw to did
Entroid's wedge here is not a better detector. Assume the read is excellent — grounding, citations, confidence, and all. The difference is where the model runs. On the ES Composable Process Fabric, a vision model runs as a governed Function inside a Deterministic Workflow, on one runtime, over a Semantic Ontology. The detection is not returned to a caller and abandoned; it is consumed by the next step of the same executing process — a step that carries the policy inline.
Concretely, the confidence-gated logic, the human-in-the-loop pause (a first-class primitive, not a bolt-on), the permission on the actor, and the action itself — reaching an external PLC, MES, or ERP through a governed Connector, the only primitive allowed to touch an outside system — all execute as one accountable transaction. And that transaction emits a single immutable record binding the whole chain:
- Detection → confidence → policy → action → actor → result. What the model saw, how sure it was, which rule evaluated, what action that rule authorized, who (or which agent, under whose authority) it ran as, and what actually happened downstream — one row, one runtime, one audit, by construction rather than by later reconciliation.
This is what reframes "auditability" from provenance-of-the-read to accountability-of-the-act. Grounding and citations, however good, prove the first link and stop. The per-action record is the whole chain. That is explainability by design: not "the model was 0.94 confident," but "at 0.94 confidence the quarantine policy fired, no human override was required, the Connector dispositioned lot 8817, and here is the actor and the result" — one query, not a forensic reconstruction across two systems.
Why this is the universal wedge
This is not a manufacturing point or a document point. It is the same point everywhere a detection triggers a consequence:
- Manufacturing: a defect detection triggers a governed quarantine-or-rework workflow, HITL-gated below a confidence floor, instead of a signal firing a PLC that stops a line with no policy attached and no record of why.
- Document / IDP: an extracted invoice field feeds a governed posting with segregation of duties enforced inline — the read and the financial control live in the same transaction, not in a citation on one side and an ERP entry on the other.
- Safety and security: a monitoring detection escalates through a permissioned, HITL-first workflow with the escalation and its outcome captured — not an alert thrown at an external system that acts, or doesn't, off-ledger.
One honest caveat, because over-claiming would make this whole argument suspect: ES is not a rip-and-replace. It does not pretend your PLCs, MES, and ERP vanish. It runs over the existing estate through governed Connectors — the integration is real, and the fabric's value is precisely that the governance travels with the action across that Connector, rather than stopping at the wall. The detect-decide-act-audit loop is one transaction on one runtime; the estate it reaches into stays exactly where it is.
Proving what the model saw is provenance. Proving what your enterprise did about it is governance — and only one of them shows up in the audit.
See what this looks like for your enterprise.
Not a demo. A strategic conversation about how your enterprise could operate
when every process runs on one governed fabric.
