Every data security platform makes one foundational choice that can't be undone later by adding features. It isn't pricing tiers or how nice the dashboard looks in a demo. It's architecture. Does data ever leave the customer's environment to get scanned, or does the scanning happen where the data already lives. That single choice determines scan performance, classification accuracy, operational cost, and regulatory defensibility, and it can't be patched over later with a clever integration. It's poured into the foundation.
Key takeaways
- In-environment architecture keeps data inside the customer's own cloud account during analysis. Data-egress architecture extracts it to the vendor's cloud, introducing lag and a compliance question that eventually needs explaining.
- Agentic AI runs continuously, and governance built for periodic review cycles cannot keep pace with continuous agent activity or incident timelines now measured in hours.
- Compliant storage does not guarantee compliant AI outputs, since RAG and embeddings can recombine document fragments in ways original access controls never anticipated.
- At petabyte scale, in-environment architectures have scanned 9 petabytes in under 72 hours at roughly $40,000 a year, versus roughly $400,000 for data-egress alternatives.
- The one question to ask any vendor is where the data's actual content sits during classification, not just where results are stored afterward.
Two models, one fork in the road
Model A keeps everything in-environment. Discovery, classification, and analysis all happen inside the customer's own cloud account, and only metadata, not the actual sensitive content, travels back to the vendor's platform. The data never crosses a boundary it didn't already live inside.
Model B is the data-egress approach. Data gets extracted, copied to the vendor's cloud, processed there, and results get sent back. It's a workable model for plenty of software categories. For sensitive data specifically, it introduces a structural lag, since data has to travel before it can be analyzed, and a compliance question that doesn't go away just because the vendor has a good SOC 2 report. Regulated data is now sitting somewhere outside the customer's walls, and someone eventually has to explain that to an auditor.
Why this stopped being a minor implementation detail
Five years ago this distinction mattered less than it does now. Point-in-time scans, run occasionally, tolerated the lag that comes with an egress model reasonably well. But McKinsey's analysis of AI data readiness makes an argument worth sitting with. Governance mechanisms designed for periodic review fail in environments where interactions occur continuously. Agentic AI runs continuously. It doesn't wait for a quarterly scan window, and neither does the risk.
McKinsey also makes a point that undercuts a comfortable assumption a lot of security teams still lean on. Compliant storage does not guarantee compliant outputs. Document-level access controls were designed for a world where a human opens a file and reads it whole. They don't map cleanly onto a world where an AI system retrieves scattered fragments of that same file through embeddings and RAG pipelines, recombining pieces in ways the original access control never anticipated. OWASP's framing on incident response makes the timeline problem explicit too. Governance practices designed for quarterly cycles cannot meet incident reporting clocks now measured in hours, not weeks.
Put plainly, the lag baked into a data-egress architecture, tolerable when review cycles were slow, becomes a structural liability the moment a governance model needs to operate on the same timescale as the agents it's supposed to be watching.
What the choice actually costs, in real numbers
The performance gap between these two models isn't subtle at scale. An in-environment, agentless architecture has demonstrated scanning 9 petabytes of data in under 72 hours, at roughly $40,000 in annual cost at 100 petabyte scale, with better than 98% classification accuracy. A comparable data-egress competitor failed to complete a scan of just 0.9 petabytes, one tenth the volume, in that same window, and egress-model total cost of ownership at that scale runs closer to $400,000 or more. That's not a rounding error between two similar options. That's a full order of magnitude, in both directions, speed and cost, stemming from one architectural decision made before either vendor's sales team ever picked up the phone.
Those numbers don't even include the softer costs. No data retention windows to manage, no deletion attestation paperwork to produce for every audit cycle, and no explaining to a regulator why customer financial records spent months sitting in someone else's cloud account.
The question to actually ask a vendor
Skip the feature checklist for a minute and ask this instead. When a platform classifies your data, where does the actual content sit while that happens. If the honest answer involves data leaving the environment at any point, that's a Model B platform with better marketing, and no amount of encryption in transit changes the fact that a copy of sensitive data now exists somewhere outside direct control.
This is the same question that sits underneath Sentra's approach to AI data readiness more broadly. What can AI actually see, what can it do with that access, and how do you keep governing it as things change. Architecture is what determines whether those three questions can be answered continuously or only as of the last scan.
See the architecture comparison. Book a demo to learn more.
