Skip to content
Back to use cases

All sectors

Input guardrail for an internal general-audience AI assistant

Every message is filtered before it reaches the model: injections, jailbreaks, harmful content and personal data. The filter fits on a single GPU.

Illustrative mock-up of the tool : guardrail · request log
Illustrative mock-up: the names, values and documents shown are fictitious.

Context

An AI assistant open to several thousand internal users, including a young audience, with two sensitivity profiles.

The problem

Blocking dangerous uses without blocking legitimate questions, at a sustainable latency and cost.

What we built

A multi-layer guardrail: anti-injection rules, detection and masking of personal data, a compact classifier trained in-house, combined with an open safety classifier. A benchmark of 500 hand-annotated messages in French and English, a comparison of 13 classifiers, and an internal tool for testing versions side by side.

Steps

  1. 01Building a set of 500 annotated messages
  2. 02Comparative benchmark of commercially available classifiers under the same conditions
  3. 03Choice of the combination and deployment on a T4 GPU with scale-to-zero
  4. 04Retraining and version comparison tool

Hosting and models

Azure Container Apps with T4 GPU (West Europe); mDeBERTa classifier served by vLLM and quantised Granite Guardian 3.1 2B (llama.cpp). No calls to an external model for filtering.

Services involved

AI in your applications · AI & Cloud Strategy

We name a client only with their written agreement. The budgets and detailed results of our engagements remain confidential.

Tell us about your use case.

An engineer replies within one business day with the scope of a first use case and a schedule.