Layered detection, evaluated per entity.
Detection across 40+ entity types through 80+ pattern detectors and 40 ML/NER entity recognizers, confirmed in context by a 3-tier pipeline.
Capabilities
Layered Detection
Every detection passes through more than one check. Pattern rules, ML entity recognizers, and a contextual validator each inspect the same content, so a match surfaced by one stage can be confirmed or rejected by another. That layering — not any single detector — is what keeps precision usable across 80+ pattern detectors and 40 ML/NER entity recognizers.
Detection Coverage
Coverage spans five categories of sensitive data: personal identifiers, financial data, healthcare identifiers, credentials and secrets, and infrastructure identifiers. Gap analysis identifies entity types where detection coverage needs improvement, so security teams can tune thresholds before a gap becomes an incident.
Confidence Calibration
DLP confidence scores are calibrated and validated — detections are only surfaced when contextual evidence supports the classification. Threshold tuning lets security teams set per-entity-type confidence thresholds: tighter for financial data, more permissive for general PII, without manual rule writing or trial-and-error in production.
Evaluated Against a Labeled Corpus
Detection quality is evaluated per entity type against a labeled corpus covering real-world prompt patterns and adversarial edge cases. That evaluation is an internal engineering practice today. Arbitex does not currently publish per-entity accuracy figures — the corpus is not yet large enough for those numbers to be meaningful, and we would rather publish nothing than publish a figure we cannot stand behind.
Detection Categories
The inspection pipeline covers five categories of sensitive data. Each is detected by a combination of pattern matching, ML entity recognition, and contextual validation.
Detection quality is evaluated per category against a labeled corpus as part of internal testing. Arbitex does not currently publish per-category accuracy figures — the evaluation corpus is not yet large enough for those numbers to be meaningful.
Three tiers. One inspection pass.
Low-ambiguity patterns resolve at Tier 1. Unstructured content escalates through AI entity recognition and contextual validation. Each tier narrows what the next one has to consider.
DLP pipeline: Input → Tier 1 Regex and Validators → Tier 2 AI entity recognition → Tier 3 Contextual validation → Action
How it works
Labeled validation corpus
A curated, labeled dataset covers supported entity types — real-world prompt patterns, multi-language coverage, and adversarial edge cases designed to surface detection gaps. The dataset grows as new entity types and evasion patterns emerge. Every labeled example carries provenance: where it came from, when it was added, and which pipeline tiers it exercises.
Per-entity evaluation
Detection quality is evaluated independently for each category rather than as a single blended score, because an aggregate hides the weak spots that matter most. Where a category underperforms, that is what drives the next round of detector and corpus work.
Corpus expansion before publication
Arbitex is expanding the evaluation corpus to the point where per-entity figures carry a meaningful confidence interval. Until it reaches that size, no per-entity accuracy numbers are published. When they are, they will be published with their corpus size and measurement date attached.
Related Resources
See the pipeline on your own data.
Three inspection tiers, checksum-validated formats, and contextual confirmation. Run it against your own traffic and see what it catches.