All case studies

Class 1 Decider: A Fast Router for LLM Requests

A small, very fast decision layer in front of LLM calls that sends each request to the right model, tool, or agent without paying a large model to decide.

Industries
Legal, Healthcare
Published
Last updated

Solvren AI is an AI forward deployment company: our engineers embed with the client’s team and build inside the client’s own cloud.

Problem

Many AI systems ask a large LLM to decide what to do with each request before any real work starts: which model, which tool, which agent.

That first decision adds latency and cost to every request, and the large model is far more capability than a routing choice needs.

What we deployed

  • A router/classifier built on a joint-embedding, contrastive-style approach, deployed in the client’s environment.
  • Requests and possible destinations are embedded into the same vector space. The router picks the closest destination in a single fast pass instead of asking an LLM to generate an answer.

How it works

Where the Decider sits
  1. 1Request arrives

    A user or system request enters the AI stack.

  2. 2Class 1 Decider

    Embeds the request and picks a route with a confidence score.

  3. 3Route

    Small model, large model, tool, skill from the registry, or the payments agent.

  4. 4Fallback

    Low-confidence requests go to the larger model.

One decision, before the LLM

Every request passes through the Decider first. It returns a route (a model, a tool, an agent, or “no LLM needed”) and a confidence score. Low-confidence requests fall back to the larger model, so accuracy is protected.

Where it fits

The Decider is designed to sit in front of the other two systems: it chooses which skill to load from the Skills Registry, and decides whether a request needs a paid call before the payments agent runs. Together, the three case studies describe one stack.

How we measure it

The benchmark covers p50 and p95 latency, accuracy against a larger-model baseline, cost per million decisions, and the hardware used. The dataset and test method are stated alongside the numbers so anyone can repeat the test.

Controls

  • Runs inside the client’s environment; requests are not sent to an external routing API.
  • Confidence threshold with fallback to the larger model.
  • Every routing decision logged for audit and re-evaluation.

Results

Note: The full benchmark (p50/p95 latency, accuracy, cost per million decisions, hardware, dataset, and method) will be published here together with the evaluation harness.

Results for Class 1 Decider: A Fast Router for LLM Requests
Measure Result
p95 routing latency Being measured
Accuracy vs. larger-model baseline Being measured
Cost per decision vs. LLM routing Being measured

We publish only numbers we measured.

Why it fits regulated teams

A fast, local decision layer keeps sensitive requests inside the client’s environment and makes every routing choice auditable, which matters for legal and healthcare workloads.

Related industries: Legal, Healthcare

See the fixed-scope offer

FAQ

What is an LLM router?

An LLM router is a decision layer that looks at each incoming request and sends it to the right model, tool, or agent. A fast router avoids paying a large model just to make that choice.

Why not let the large model route requests itself?

Routing with a large LLM adds a full model call of latency and cost to every request before real work starts. A small embedding-based router makes the same choice in a single fast pass.

What happens when the router is unsure?

Requests below a confidence threshold fall back to the larger model, so accuracy on hard cases is protected.

More case studies

Start with one workflow

Book a free 30-minute AI audit. We look at one workflow with you and tell you whether it can reach production in 6 to 8 weeks. No commitment, just 30 minutes with an engineer.

Book a free AI audit
Book a free AI audit