Skip to main content

Know what leaves. Block what shouldn't.

DLP detects entities — PII, credentials, and regulated data — across every channel through a 3-tier DLP pipeline at under 2ms p99 latency. Content Categories classify topics — classifying prompts by topic across 8 domains and 26 sub-categories. Together, they form two complementary detection surfaces in a single inspection pipeline, enabling governance rules that fire on both what data is in the request and what the request is about.

Every request and response is inspected in real time — flagged events appear in the detection log with severity, matched pattern, and the enforcement action taken.

DLP detection table showing flagged content with severity levels and enforcement actions

Capabilities

80+ Pattern Detectors, with checksum validation where applicable

Detect credit card numbers, IBANs, Social Security numbers, and government IDs across the regex layer. Where the entity type carries a check digit or structural format spec — IBAN (MOD-97), ABA routing, NPI (Luhn), DEA, ITIN, EIN, Canadian SIN, IMEI, SWIFT/BIC, and similar regulated identifiers — a validator runs after the regex match to confirm the data is structurally real, eliminating false positives on invoice numbers and arbitrary digit strings. Patterns that lack a published checksum spec rely on context keywords and downstream tier-2/tier-3 confirmation. The credential detection layer covers API keys, database connection strings, OAuth tokens, and cloud provider secrets across every major platform.

ML-Driven Named Entity Recognition with Contextual Validation

ML-driven named entity recognition runs inference to catch names, addresses, medical record numbers, financial account identifiers, and more. An ML-driven contextual validation model (under 2ms p99 latency) evaluates each match against surrounding text to reduce false positives. The result: fewer false alarms and faster review cycles for security teams.

12 Compliance Frameworks, Pre-Mapped

Enforce PCI-DSS, HIPAA, GDPR, GLBA, SOX, CCPA, BSA/AML, SEC Reg FD, FERPA, EU AI Act, NIST AI RMF, ISO/IEC 42001 with pre-built compliance bundles. Each bundle activates the correct detectors across all 3 DLP layers and maps enforcement actions — block, redact, or log — to the specific data types each regulation covers. Activate a framework and Arbitex handles the mapping.

Content Category Detection

Topic-based classification as a policy engine condition. Content Categories classify prompts across 8 content domains and 26 sub-categories — legal, financial, medical, security, competitive, and more. The detection surface extends beyond named topics to include code detection, language identification, and prompt safety classifiers (jailbreak, prompt leak, off-topic) — all keyword and rule-based. All classifications fire as content_category conditions in the same policy rule chain as DLP entity findings.

How Detection Quality Is Built

Detection quality does not come from any single detector. It comes from layering — pattern rules, ML entity recognition, and contextual validation each inspecting the same content, so that a match one stage surfaces, another can confirm or reject.

Validated
Structured Formats

Payment card numbers, IBANs, and other identifiers carrying a check digit are arithmetically validated, not just pattern-matched. When the format is unambiguous, a validated match is close to unambiguous too.

In Context
Contextual Confirmation

ML-driven contextual validation analyzes surrounding text to distinguish genuine sensitive data from false positives. An identifier in a code fixture is treated differently from the same identifier in a patient record.

Per Entity
Evaluated by Category

Detection quality is evaluated per entity type against a labeled corpus rather than as a single blended score, because an aggregate hides the weak spots that matter most when tuning detectors.

Per-entity evaluation is an internal engineering practice. Arbitex does not currently publish per-entity accuracy figures — the evaluation corpus is not yet large enough for those numbers to be meaningful.

How It Works

01

Define your compliance scope

Select which regulatory frameworks, data types, and content categories to enforce. Compliance bundles activate the correct detectors across regex, entity recognition, and contextual validation layers. Content Categories add topic classification — 8 domains and 26 sub-categories — as a separate configuration surface alongside entity detection. Custom rules extend coverage for organization-specific patterns.

02

Inspect on input and output

Requests pass through all 3 detection layers. Response inspection available on select plans. Regex detectors validate format and checksum. ML-driven entity recognition extracts entities from free text. A context-aware language model confirms or rejects each detection. Compliance bundles enforce framework-specific rules. Each match triggers a configurable action: block the request, redact the sensitive content, or log the detection and allow it through.

03

Manage from the DLP console

Build custom detection rules in the rule builder. Test patterns against sample payloads with the preview endpoint. Monitor detection trends and false positive rates from the analytics dashboard. Bulk import and export rules across environments.

Protection Add-on

Credential Intelligence

Detect known compromised credentials before they reach any AI model.

Pattern detection catches what looks like a credential. Credential Intelligence catches what is one — and tells you how dangerous it is.

When a user sends text through the AI gateway, Credential Intelligence checks it against a compromised credential dataset using secure comparison. No cleartext is ever stored or transmitted. The check runs in parallel with the DLP pipeline, adding negligible latency to each request.

What distinguishes this from generic credential detection is the frequency signal. A credential seen in millions of breaches carries a different risk profile than one seen once. Credential Intelligence assigns each detection to a risk level — Critical, High, Medium, or Low — so security teams can act on the highest-risk exposures immediately.

Standard

Monthly dataset refresh

Secure comparison against a curated compromised credential dataset. Frequency-weighted risk levels: Critical, High, Medium, Low.

Enterprise

Credential Intelligence

Checks every request against a compromised credential dataset. Frequency-weighted risk levels: Critical, High, Medium, Low.

Related Resources

Policy Engine

Rules-based governance for every AI request

Content Categories

Topic classification as a policy engine condition

Compliance Frameworks

Pre-built policy packs for regulatory requirements

Audit Log

Tamper-proof activity trail

Healthcare

HIPAA compliance and PHI detection

Financial Services

PCI-DSS and SOX compliance

Read the DLP admin guide

Stop data leaks at the gateway.