DevvProxy Integration Guide
Learn how to route your LLM traffic through DevvProxy to enforce edge PII sanitization and 0ms deterministic caching with zero application code changes.
1. Zero-Code-Refactor Quickstart
DevvProxy implements the exact HTTP wire protocol of the OpenAI Chat Completions API. Point base_url to DevvProxy to activate edge PII redaction and deterministic caching:
from openai import OpenAI
client = OpenAI(
base_url="https://devvproxy.vercel.app/api/v1", # Point to DevvProxy
api_key="devv_live_demo_9481b37c" # Or your virtual key
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Draft invoice for user@corp.com with card 4532-1188-9922-3344"}]
)
print(response.choices[0].message.content)
2. ReDoS-Resistant PII Redaction Rules
DevvProxy runs a linear-time, non-backtracking redaction engine at the edge before packets are forwarded to upstream LLM providers:
Validates 13–19 digit sequences using the mathematical Luhn checksum algorithm. Non-card digit sequences are preserved safely.
Placeholder: [REDACTED_CREDIT_CARD_N]Linear RFC 5322 compliant regular expression with zero nested quantifiers, resistant to ReDoS attacks.
Placeholder: [REDACTED_EMAIL_N]Matches formatted standard XXX-XX-XXXX patterns with boundary protections and reserved area code filtering.
Placeholder: [REDACTED_SSN_N]Detects high-entropy keys: OpenAI (sk-...), GitHub tokens (ghp_...), and AWS credentials.
3. Deterministic Caching & Custom Response Headers
DevvProxy inspects response headers on every query. If a request has been processed before by your team, it resolves in under 5ms:
| Header | Values | Description |
|---|---|---|
| X-Devv-Cache | HIT | MISS | Indicates whether response was served from 0ms cache. |
| X-Devv-Provider | cache | openai | groq | simulator | Identifies which upstream engine fulfilled the request. |
| X-Devv-PII-Scrubbed | Integer (e.g. 2) | Number of sensitive entities masked at the edge. |
Server-Sent Events (SSE) Streaming Compatibility
DevvProxy natively supports wire-compatible Server-Sent Events (stream: true). Incoming streams pass raw tokens directly to the client with zero buffering delay (<5ms TTFT overhead), while background collectors asynchronously assemble the full response to populate the L1 Hot Cache and record non-blocking telemetry.
X-Devv-Cache: HIT) in under 15ms.