All sectors
Multi-model gateway with complexity-based routing
Each request is classified as simple, medium or complex, then sent to the leanest model capable of answering it. The organisation can change models without redoing the integration.

Context
An organisation that wants to use several models (commercial and open) without multiplying integrations or costs.
The problem
A top-tier model for every message is expensive and slow; a single model creates dependency.
What we built
A single LiteLLM proxy for all models, a complexity judge ahead of each request, configuration variants per user profile, and benchmarks for choosing the models at each level on programmatically verifiable criteria.
Steps
- 01Single gateway and per-application keys
- 02Three-level complexity judge
- 03Model selection benchmarks (accuracy, latency, cost)
- 04Documented choice of models per level
Hosting and models
LiteLLM proxy on the client's cloud; Gemini and Claude models via Vertex AI; benchmarks run on an open-model API platform hosted in France (Scaleway Generative APIs).
Services involved
Agentic platforms · AI & Cloud Strategy
We name a client only with their written agreement. The budgets and detailed results of our engagements remain confidential.