Skip to main content
Protect

Layered detection, evaluated per entity.

Detection across 40+ entity types through 80+ pattern detectors and 40 ML/NER entity recognizers, confirmed in context by a 3-tier pipeline.

How detection quality is built

Three tiers, validated formats, confirmed in context

Multi-tier

Three independent detection tiers

Pattern matching, ML entity recognition, and contextual validation each inspect the same content. A detection that one tier surfaces, another can confirm or reject.

Validated

Checksum and format validation

Structured identifiers — payment card numbers, IBANs, and other formats carrying a check digit — are arithmetically validated, not just pattern-matched.

In context

Contextual confirmation

An AI-powered contextual validator evaluates each match against the surrounding text, so an identifier in a code fixture is treated differently from one in a patient record.

Detection quality is evaluated internally per entity type against a labeled corpus. Arbitex does not currently publish per-entity accuracy figures — the evaluation corpus is not yet large enough for those numbers to be meaningful, and publishing a figure we cannot stand behind would be worse than publishing none.

Capabilities

Layered Detection

Every detection passes through more than one check. Pattern rules, ML entity recognizers, and a contextual validator each inspect the same content, so a match surfaced by one stage can be confirmed or rejected by another. That layering — not any single detector — is what keeps precision usable across 80+ pattern detectors and 40 ML/NER entity recognizers.

Detection Coverage

Coverage spans five categories of sensitive data: personal identifiers, financial data, healthcare identifiers, credentials and secrets, and infrastructure identifiers. Gap analysis identifies entity types where detection coverage needs improvement, so security teams can tune thresholds before a gap becomes an incident.

Confidence Calibration

DLP confidence scores are calibrated and validated — detections are only surfaced when contextual evidence supports the classification. Threshold tuning lets security teams set per-entity-type confidence thresholds: tighter for financial data, more permissive for general PII, without manual rule writing or trial-and-error in production.

Evaluated Against a Labeled Corpus

Detection quality is evaluated per entity type against a labeled corpus covering real-world prompt patterns and adversarial edge cases. That evaluation is an internal engineering practice today. Arbitex does not currently publish per-entity accuracy figures — the corpus is not yet large enough for those numbers to be meaningful, and we would rather publish nothing than publish a figure we cannot stand behind.

Coverage Breadth

Detection Categories

The inspection pipeline covers five categories of sensitive data. Each is detected by a combination of pattern matching, ML entity recognition, and contextual validation.

Names, addresses, national identifiers, dates of birth
Payment card numbers, IBANs, and account identifiers
Medical record numbers and other healthcare identifiers
API keys, tokens, private keys, and connection strings
Network addresses and infrastructure identifiers

Detection quality is evaluated per category against a labeled corpus as part of internal testing. Arbitex does not currently publish per-category accuracy figures — the evaluation corpus is not yet large enough for those numbers to be meaningful.

Detection Architecture

Three tiers. One inspection pass.

Low-ambiguity patterns resolve at Tier 1. Unstructured content escalates through AI entity recognition and contextual validation. Each tier narrows what the next one has to consider.

DLP pipeline: Input → Tier 1 Regex and Validators → Tier 2 AI entity recognition → Tier 3 Contextual validation → Action

How it works

01

Labeled validation corpus

A curated, labeled dataset covers supported entity types — real-world prompt patterns, multi-language coverage, and adversarial edge cases designed to surface detection gaps. The dataset grows as new entity types and evasion patterns emerge. Every labeled example carries provenance: where it came from, when it was added, and which pipeline tiers it exercises.

02

Per-entity evaluation

Detection quality is evaluated independently for each category rather than as a single blended score, because an aggregate hides the weak spots that matter most. Where a category underperforms, that is what drives the next round of detector and corpus work.

03

Corpus expansion before publication

Arbitex is expanding the evaluation corpus to the point where per-entity figures carry a meaningful confidence interval. Until it reaches that size, no per-entity accuracy numbers are published. When they are, they will be published with their corpus size and measurement date attached.

Related Resources

DLP Protection

Inspect every AI prompt for sensitive data

DLP Pipeline

3-tier content inspection pipeline

Compliance Frameworks

Pre-built policy packs for regulatory requirements

Policy Engine

Rules-based governance for every AI request

See the pipeline on your own data.

Three inspection tiers, checksum-validated formats, and contextual confirmation. Run it against your own traffic and see what it catches.