All sectors
Input guardrail for an internal general-audience AI assistant
Every message is filtered before it reaches the model: injections, jailbreaks, harmful content and personal data. The filter fits on a single GPU.

Context
An AI assistant open to several thousand internal users, including a young audience, with two sensitivity profiles.
The problem
Blocking dangerous uses without blocking legitimate questions, at a sustainable latency and cost.
What we built
A multi-layer guardrail: anti-injection rules, detection and masking of personal data, a compact classifier trained in-house, combined with an open safety classifier. A benchmark of 500 hand-annotated messages in French and English, a comparison of 13 classifiers, and an internal tool for testing versions side by side.
Steps
- 01Building a set of 500 annotated messages
- 02Comparative benchmark of commercially available classifiers under the same conditions
- 03Choice of the combination and deployment on a T4 GPU with scale-to-zero
- 04Retraining and version comparison tool
Hosting and models
Azure Container Apps with T4 GPU (West Europe); mDeBERTa classifier served by vLLM and quantised Granite Guardian 3.1 2B (llama.cpp). No calls to an external model for filtering.
Services involved
AI in your applications · AI & Cloud Strategy
We name a client only with their written agreement. The budgets and detailed results of our engagements remain confidential.