Class 1 Decider: A Fast Router for LLM Requests
A small, very fast decision layer in front of LLM calls that sends each request to the right model, tool, or agent without paying a large model to decide.
- Industries
- Legal, Healthcare
- Author
- Solvren AI Engineering
- Published
- Last updated
Solvren AI is an AI forward deployment company: our engineers embed with the client’s team and build inside the client’s own cloud.
Problem
Many AI systems ask a large LLM to decide what to do with each request before any real work starts: which model, which tool, which agent.
That first decision adds latency and cost to every request, and the large model is far more capability than a routing choice needs.
What we deployed
- A router/classifier built on a joint-embedding, contrastive-style approach, deployed in the client’s environment.
- Requests and possible destinations are embedded into the same vector space. The router picks the closest destination in a single fast pass instead of asking an LLM to generate an answer.
How it works
-
1Request arrives
A user or system request enters the AI stack.
-
2Class 1 Decider
Embeds the request and picks a route with a confidence score.
-
3Route
Small model, large model, tool, skill from the registry, or the payments agent.
-
4Fallback
Low-confidence requests go to the larger model.
One decision, before the LLM
Every request passes through the Decider first. It returns a route (a model, a tool, an agent, or “no LLM needed”) and a confidence score. Low-confidence requests fall back to the larger model, so accuracy is protected.
Where it fits
The Decider is designed to sit in front of the other two systems: it chooses which skill to load from the Skills Registry, and decides whether a request needs a paid call before the payments agent runs. Together, the three case studies describe one stack.
How we measure it
The benchmark covers p50 and p95 latency, accuracy against a larger-model baseline, cost per million decisions, and the hardware used. The dataset and test method are stated alongside the numbers so anyone can repeat the test.
Controls
- Runs inside the client’s environment; requests are not sent to an external routing API.
- Confidence threshold with fallback to the larger model.
- Every routing decision logged for audit and re-evaluation.
Results
Note: The full benchmark (p50/p95 latency, accuracy, cost per million decisions, hardware, dataset, and method) will be published here together with the evaluation harness.
| Measure | Result |
|---|---|
| p95 routing latency | Being measured |
| Accuracy vs. larger-model baseline | Being measured |
| Cost per decision vs. LLM routing | Being measured |
We publish only numbers we measured.
Why it fits regulated teams
A fast, local decision layer keeps sensitive requests inside the client’s environment and makes every routing choice auditable, which matters for legal and healthcare workloads.
Related industries: Legal, Healthcare
FAQ
What is an LLM router?
An LLM router is a decision layer that looks at each incoming request and sends it to the right model, tool, or agent. A fast router avoids paying a large model just to make that choice.
Why not let the large model route requests itself?
Routing with a large LLM adds a full model call of latency and cost to every request before real work starts. A small embedding-based router makes the same choice in a single fast pass.
What happens when the router is unsure?
Requests below a confidence threshold fall back to the larger model, so accuracy on hard cases is protected.