DEV Community

Solon Framework
Solon Framework

Posted on Originally published at dev.to

One Router, Four Strategies: How Solon AI Picks the Right ChatModel

In the multi-model era, your application already talks to more than one LLM: a cheap fast model for casual chat, an expensive smart one for reasoning, plus dedicated code models, vision models, and local small models. That raises a question — who gets each request?

The solon-ai-router module in Solon AI (v4.1.0+) answers this the "Solon way": a router so small it almost doesn't exist. Only 6 public classes in the whole module, yet it makes model selection clean.

Repo: https://github.com/opensolon/solon-ai (module: solon-ai-router)


1. It Does Exactly One Thing: Pick a ChatModel for You

The scope of ChatModelRouter is deliberately narrow:

  • When a chat request is created, select one physical model from a set of already-built ChatModel instances;
  • Return the ChatRequestDesc produced by that model, untouched.

Note the second half — the Router never hijacks request logic. After routing, session(), role(), instruction(), systemPrompt(), options(), call(), and stream() are all handled by the target model's own implementation. No Agent, no Flow, no Harness — just a pure model selector.

That differs from the common "AI gateway" approach where routing, retries, fallback, caching, and rate limiting all fuse into one big blob. Solon AI draws the line clearly: routing is routing, execution is execution.

<dependency>
    <groupId>org.noear</groupId>
    <artifactId>solon-ai-router</artifactId>
    <version>4.1.0</version>
</dependency>
Enter fullscreen mode Exit fullscreen mode

2. Basic Usage

Register candidate models, wrap them with a strategy, then send prompts like a normal ChatModel:

ChatModelRouter router = new ChatModelRouter(Arrays.asList(
        new ChatModelRoute("fast", "Fast, cheap model for routine Q&A", 1, fastModel),
        new ChatModelRoute("reasoning", "Model for complex analysis and multi-step reasoning", 1, reasoningModel)
), new RoundRobinRoutingStrategy());

ChatResponse response = router.prompt("Analyze this concurrency issue")
        .options(options -> options.temperature(0.2F))
        .call();
Enter fullscreen mode Exit fullscreen mode

A ChatModelRoute takes four arguments: id, description (used later by the LLM classifier), weight (must be > 0), and the chatModel instance.

ChatModelRouter mirrors the four ChatModel entry points, so migration cost is near zero:

router.prompt(prompt);
router.prompt(messages);
router.prompt(systemMessage, userMessage);
router.prompt("user message");
Enter fullscreen mode Exit fullscreen mode

One subtle detail: routing happens at prompt(...) time. Even if you only create the request and never call call() or stream(), one round-robin slot has already been consumed. Don't hoard prompt() outside a loop.


3. Four Built-in Strategies, From Dumb to Smart

1) RoundRobinRoutingStrategy

new RoundRobinRoutingStrategy()
Enter fullscreen mode Exit fullscreen mode

Cycles through candidates in registration order. Backed by an AtomicLong cursor — lock-free, thread-safe, shareable across threads. Perfect for a pool of equivalent models (e.g., same spec, multiple accounts) to spread quota evenly.

2) WeightedRoundRobinRoutingStrategy

new WeightedRoundRobinRoutingStrategy()
Enter fullscreen mode Exit fullscreen mode

Smooth weighted round-robin over ChatModelRoute.getWeight() — the same algorithm nginx uses. With a 5:1 weight ratio, requests spread out evenly instead of "5 hits on A, then 1 on B", which is friendlier to rate-limit windows.

A design decision worth noting: once used, a strategy instance pins the candidate IDs, order, and weights. If the candidate list topology changes afterwards (added/removed candidates, changed weights), it throws RoutingException instead of quietly running with the new config. Configuration drift should be visible, not swallowed.

3) RuleBasedRoutingStrategy

RuleBasedRoutingStrategy strategy = new RuleBasedRoutingStrategy(Arrays.asList(
        new RoutingRule("reasoning", context -> {
            String content = context.getPrompt().getUserContent();
            return content != null && content.contains("analyze");
        }),
        new RoutingRule("fast", context -> true)   // catch-all: keep it last
));
Enter fullscreen mode Exit fullscreen mode

Rules execute in registration order; the first match wins. RoutingContext exposes the Prompt, so you can branch on message content, attributes, or anything else — tenant tier, task type, message length.

Two semantics you must know:

  • If no rule matches, the strategy throws RoutingException immediately. It won't guess a default model for you — add an explicit catch-all rule if every request needs a home;
  • A predicate that throws a RuntimeException gets wrapped in a RoutingException and propagated, never silently skipped.

4) SmartRoutingStrategy — an LLM as the Dispatcher

ChatModelRouter router = new ChatModelRouter(Arrays.asList(
        new ChatModelRoute("fast", "Routine Q&A", 3, fastModel),
        new ChatModelRoute("reasoning", "Complex reasoning", 1, reasoningModel)
), new SmartRoutingStrategy(classifierModel));

ChatResponse response = router.prompt("Analyze this concurrency issue")
        .options(options -> options.temperature(0.2F))
        .call();
Enter fullscreen mode Exit fullscreen mode

The most interesting one: a dedicated, ordinary ChatModel acts as the classifier. It reads the candidates' description fields plus the original messages, then emits a structured RoutingDecision (with routeId + reasoning).

Look inside SmartRoutingStrategy and the classifier prompt is refreshingly plain:

You are a routing classifier. Select only one candidate route id
and explain the reason.

Candidate routes:
- fast: Routine Q&A
- reasoning: Complex reasoning
...
Enter fullscreen mode Exit fullscreen mode

The implementation leans on Solon AI's structured output: options.outputSchema(RoutingDecision.class) constrains the classifier to a valid shape, then message.toBean(RoutingDecision.class) deserializes it, then it validates that routeId is non-empty, reasoning is non-empty, and the id is a registered candidate. Any failed step fails the request immediately — it never proceeds to a business model with uncertainty.

This is why ChatModelRoute.description is required — it's the model's "business card" shown to the classifier. How well you write it directly determines routing accuracy.

The cost is equally clear: one extra model call per request (latency included). Use a small, fast classifier — e.g., a local 1.5B model via Ollama — not your flagship reasoning model.


4. Explicit Routing: The Escape Hatch

A Prompt can bypass the strategy via a fixed attribute:

Prompt prompt = Prompt.of("Summarize this")
        .attrPut(ChatModelRouter.ATTR_ROUTE_ID, "fast");

router.prompt(prompt).call();
Enter fullscreen mode Exit fullscreen mode

Semantics stay deliberate: the explicit id must be a registered, non-empty string; an unknown or invalid id throws RoutingException and never falls back to the configured strategy. This makes "let advanced users pick the model" a clean feature — no more silent model swaps the user never asked for.


5. Error Semantics: No Magic, Only Determinism

List every failure path of ChatModelRouter.prompt() and you see the module's full personality:

Scenario Behavior
Empty decision / empty routeId / unknown candidate Throws RoutingException immediately
No rule matched (rule strategy) Throws RoutingException immediately
Classifier failed / invalid output / unknown candidate Throws RoutingException; business models are never called
Invalid explicit route id Throws RoutingException; no fallback to the strategy
Target model call()/stream() throws Original type and propagation preserved

One sentence: the Router provides no default candidate, no failover, no silent fallback.

In production, that's a life-saver. The most common AI disaster isn't "the request failed" — it's "the request quietly succeeded some other way": degraded to a model the user never approved, silently switched to one 10× more expensive, or misclassified with nobody knowing. Solon AI makes every routing failure loud and hands the degradation decision back to business code. Need a fallback? Catch RoutingException and write your own — three lines of code, but the control is yours.


6. Choosing a Strategy

Scenario Strategy
Spread load/quota across equivalent models RoundRobinRoutingStrategy
Split traffic by ratio (e.g., 3:1) WeightedRoundRobinRoutingStrategy
Deterministic business rules (VIP routing, long-input routing) RuleBasedRoutingStrategy
Rules can't enumerate everything; let the model decide SmartRoutingStrategy (with a small classifier)
User manually picks the model ATTR_ROUTE_ID explicit routing

They compose too: put deterministic rules first, then nest smart routing as the catch-all for the long tail.


Wrapping Up

solon-ai-router is 6 public classes that decompose a real engineering problem — multi-model selection — cleanly: single responsibility, explicit failure, pluggable strategies. It doesn't aspire to be an AI gateway; it does model selection, and does it without surprises.

That's the taste the Solon ecosystem keeps showing — restrained, but never compromising.


Links

All APIs and behaviors in this post were verified against the solon-ai repo's solon-ai-router module source (since 4.1).

Top comments (0)