Skip to content
Back to use cases

All sectors

Multi-model gateway with complexity-based routing

Each request is classified as simple, medium or complex, then sent to the leanest model capable of answering it. The organisation can change models without redoing the integration.

Illustrative mock-up of the tool : gateway · request routing
Illustrative mock-up: the names, values and documents shown are fictitious.

Context

An organisation that wants to use several models (commercial and open) without multiplying integrations or costs.

The problem

A top-tier model for every message is expensive and slow; a single model creates dependency.

What we built

A single LiteLLM proxy for all models, a complexity judge ahead of each request, configuration variants per user profile, and benchmarks for choosing the models at each level on programmatically verifiable criteria.

Steps

  1. 01Single gateway and per-application keys
  2. 02Three-level complexity judge
  3. 03Model selection benchmarks (accuracy, latency, cost)
  4. 04Documented choice of models per level

Hosting and models

LiteLLM proxy on the client's cloud; Gemini and Claude models via Vertex AI; benchmarks run on an open-model API platform hosted in France (Scaleway Generative APIs).

Services involved

Agentic platforms · AI & Cloud Strategy

We name a client only with their written agreement. The budgets and detailed results of our engagements remain confidential.

Tell us about your use case.

An engineer replies within one business day with the scope of a first use case and a schedule.