In the multi-model era, your application already talks to more than one LLM: a cheap fast model for casual chat, an expensive smart one for reasoning, plus dedicated code models, vision models, and local small models. That raises a question — who gets each request?
The solon-ai-router module in Solon AI (v4.1.0+) answers this the "Solon way": a router so small it almost doesn't exist. Only 6 public classes in the whole module, yet it makes model selection clean.
Repo: https://github.com/opensolon/solon-ai (module: solon-ai-router)
1. It Does Exactly One Thing: Pick a ChatModel for You
The scope of ChatModelRouter is deliberately narrow:
- When a chat request is created, select one physical model from a set of already-built
ChatModelinstances; - Return the
ChatRequestDescproduced by that model, untouched.
Note the second half — the Router never hijacks request logic. After routing, session(), role(), instruction(), systemPrompt(), options(), call(), and stream() are all handled by the target model's own implementation. No Agent, no Flow, no Harness — just a pure model selector.
That differs from the common "AI gateway" approach where routing, retries, fallback, caching, and rate limiting all fuse into one big blob. Solon AI draws the line clearly: routing is routing, execution is execution.
<dependency>
<groupId>org.noear</groupId>
<artifactId>solon-ai-router</artifactId>
<version>4.1.0</version>
</dependency>
2. Basic Usage
Register candidate models, wrap them with a strategy, then send prompts like a normal ChatModel:
ChatModelRouter router = new ChatModelRouter(Arrays.asList(
new ChatModelRoute("fast", "Fast, cheap model for routine Q&A", 1, fastModel),
new ChatModelRoute("reasoning", "Model for complex analysis and multi-step reasoning", 1, reasoningModel)
), new RoundRobinRoutingStrategy());
ChatResponse response = router.prompt("Analyze this concurrency issue")
.options(options -> options.temperature(0.2F))
.call();
A ChatModelRoute takes four arguments: id, description (used later by the LLM classifier), weight (must be > 0), and the chatModel instance.
ChatModelRouter mirrors the four ChatModel entry points, so migration cost is near zero:
router.prompt(prompt);
router.prompt(messages);
router.prompt(systemMessage, userMessage);
router.prompt("user message");
One subtle detail: routing happens at prompt(...) time. Even if you only create the request and never call call() or stream(), one round-robin slot has already been consumed. Don't hoard prompt() outside a loop.
3. Four Built-in Strategies, From Dumb to Smart
1) RoundRobinRoutingStrategy
new RoundRobinRoutingStrategy()
Cycles through candidates in registration order. Backed by an AtomicLong cursor — lock-free, thread-safe, shareable across threads. Perfect for a pool of equivalent models (e.g., same spec, multiple accounts) to spread quota evenly.
2) WeightedRoundRobinRoutingStrategy
new WeightedRoundRobinRoutingStrategy()
Smooth weighted round-robin over ChatModelRoute.getWeight() — the same algorithm nginx uses. With a 5:1 weight ratio, requests spread out evenly instead of "5 hits on A, then 1 on B", which is friendlier to rate-limit windows.
A design decision worth noting: once used, a strategy instance pins the candidate IDs, order, and weights. If the candidate list topology changes afterwards (added/removed candidates, changed weights), it throws RoutingException instead of quietly running with the new config. Configuration drift should be visible, not swallowed.
3) RuleBasedRoutingStrategy
RuleBasedRoutingStrategy strategy = new RuleBasedRoutingStrategy(Arrays.asList(
new RoutingRule("reasoning", context -> {
String content = context.getPrompt().getUserContent();
return content != null && content.contains("analyze");
}),
new RoutingRule("fast", context -> true) // catch-all: keep it last
));
Rules execute in registration order; the first match wins. RoutingContext exposes the Prompt, so you can branch on message content, attributes, or anything else — tenant tier, task type, message length.
Two semantics you must know:
- If no rule matches, the strategy throws
RoutingExceptionimmediately. It won't guess a default model for you — add an explicit catch-all rule if every request needs a home; - A predicate that throws a
RuntimeExceptiongets wrapped in aRoutingExceptionand propagated, never silently skipped.
4) SmartRoutingStrategy — an LLM as the Dispatcher
ChatModelRouter router = new ChatModelRouter(Arrays.asList(
new ChatModelRoute("fast", "Routine Q&A", 3, fastModel),
new ChatModelRoute("reasoning", "Complex reasoning", 1, reasoningModel)
), new SmartRoutingStrategy(classifierModel));
ChatResponse response = router.prompt("Analyze this concurrency issue")
.options(options -> options.temperature(0.2F))
.call();
The most interesting one: a dedicated, ordinary ChatModel acts as the classifier. It reads the candidates' description fields plus the original messages, then emits a structured RoutingDecision (with routeId + reasoning).
Look inside SmartRoutingStrategy and the classifier prompt is refreshingly plain:
You are a routing classifier. Select only one candidate route id
and explain the reason.
Candidate routes:
- fast: Routine Q&A
- reasoning: Complex reasoning
...
The implementation leans on Solon AI's structured output: options.outputSchema(RoutingDecision.class) constrains the classifier to a valid shape, then message.toBean(RoutingDecision.class) deserializes it, then it validates that routeId is non-empty, reasoning is non-empty, and the id is a registered candidate. Any failed step fails the request immediately — it never proceeds to a business model with uncertainty.
This is why ChatModelRoute.description is required — it's the model's "business card" shown to the classifier. How well you write it directly determines routing accuracy.
The cost is equally clear: one extra model call per request (latency included). Use a small, fast classifier — e.g., a local 1.5B model via Ollama — not your flagship reasoning model.
4. Explicit Routing: The Escape Hatch
A Prompt can bypass the strategy via a fixed attribute:
Prompt prompt = Prompt.of("Summarize this")
.attrPut(ChatModelRouter.ATTR_ROUTE_ID, "fast");
router.prompt(prompt).call();
Semantics stay deliberate: the explicit id must be a registered, non-empty string; an unknown or invalid id throws RoutingException and never falls back to the configured strategy. This makes "let advanced users pick the model" a clean feature — no more silent model swaps the user never asked for.
5. Error Semantics: No Magic, Only Determinism
List every failure path of ChatModelRouter.prompt() and you see the module's full personality:
| Scenario | Behavior |
|---|---|
| Empty decision / empty routeId / unknown candidate | Throws RoutingException immediately |
| No rule matched (rule strategy) | Throws RoutingException immediately |
| Classifier failed / invalid output / unknown candidate | Throws RoutingException; business models are never called
|
| Invalid explicit route id | Throws RoutingException; no fallback to the strategy
|
Target model call()/stream() throws |
Original type and propagation preserved |
One sentence: the Router provides no default candidate, no failover, no silent fallback.
In production, that's a life-saver. The most common AI disaster isn't "the request failed" — it's "the request quietly succeeded some other way": degraded to a model the user never approved, silently switched to one 10× more expensive, or misclassified with nobody knowing. Solon AI makes every routing failure loud and hands the degradation decision back to business code. Need a fallback? Catch RoutingException and write your own — three lines of code, but the control is yours.
6. Choosing a Strategy
| Scenario | Strategy |
|---|---|
| Spread load/quota across equivalent models | RoundRobinRoutingStrategy |
| Split traffic by ratio (e.g., 3:1) | WeightedRoundRobinRoutingStrategy |
| Deterministic business rules (VIP routing, long-input routing) | RuleBasedRoutingStrategy |
| Rules can't enumerate everything; let the model decide |
SmartRoutingStrategy (with a small classifier) |
| User manually picks the model |
ATTR_ROUTE_ID explicit routing |
They compose too: put deterministic rules first, then nest smart routing as the catch-all for the long tail.
Wrapping Up
solon-ai-router is 6 public classes that decompose a real engineering problem — multi-model selection — cleanly: single responsibility, explicit failure, pluggable strategies. It doesn't aspire to be an AI gateway; it does model selection, and does it without surprises.
That's the taste the Solon ecosystem keeps showing — restrained, but never compromising.
Links
- Solon AI repo: https://github.com/opensolon/solon-ai
- Solon website: https://solon.noear.org
- Getting started with Solon AI: https://solon.noear.org/article/learn-solon-ai
All APIs and behaviors in this post were verified against the solon-ai repo's solon-ai-router module source (since 4.1).
Top comments (0)