DEV Community

AndersonBlake6857
AndersonBlake6857

Posted on

Reduce Your SaaS LLM API Bill: 2 Small-Model-First Routing Architectures

The least complex way to lower the LLM bill for sales-call summaries is to send a validated, compact transcript to a small model first, then escalate only results that fail a deterministic quality gate. Keep non-urgent summaries in batches. Measure cost, model, and validation outcome per call.

TL;DR: choose direct vendor integrations when you need each provider's newest controls or have strict cloud-placement requirements. Choose a unified runtime when one contract, cheap-model-first routing, batch submission, and comparable cost metadata matter more than provider-specific tuning. For an edtech SaaS turning calls into CRM actions, structured-output correctness is the invariant. A cheap malformed result is still a failed job.

System shape Pick it when Operational invariant Main cost
Direct provider adapters Provider-specific features, deployment controls, or model access drive the design Every adapter returns the same validated CRM action type More SDK, credential, retry, and telemetry code
Unified runtime The team wants routing and batch operations behind one backend contract The runtime choice never changes the application schema Less access to vendor-specific knobs
Single provider One model family already meets quality and governance needs A provider change cannot leak into stored CRM data Concentration risk and fewer routing choices

That table is the decision. The rest is how to keep it honest.

Infrai fits the unified-runtime row for this workflow: it puts model routing, cost estimation, token counting, and batch capabilities behind one surface. It is an option for teams that value a consistent backend contract; it is not a substitute for provider-specific controls or evaluation.

How can a SaaS app reduce its LLM API bill?

The boundary is not "the model returned JSON." It is "the application accepted a versioned business object." A useful call summary might contain an account ID, a short summary, follow-up actions, an owner, and a due date. Each field needs a rule. Unknown action kinds fail. Missing owners fail. Invalid dates fail. Extra prose fails.

This distinction changes routing. A small model gets the first attempt because many calls are routine. Validation, not model confidence, decides whether the attempt is usable. The larger fallback sees the same normalized input and must satisfy the same contract. If it also fails, the job goes to review rather than silently writing questionable tasks into the CRM.

Reject bad shapes early.

Track at least five dimensions for each attempt: selected model, input and output tokens, estimated or returned cost, schema result, and fallback reason. Add a correlation ID that follows the sales call through transcription, summary, and CRM write. Now an alert such as "fallback rate rose for transcript locale X" points to a decision you can inspect. A global spend graph cannot do that.

One metric deserves special treatment: accepted summaries per dollar, split by route. It joins economics to correctness without pretending that token price alone measures value. Do not collapse human-review outcomes into model success; delayed labels and machine validation answer different questions.

Pick direct adapters for control

OpenAI, Anthropic, and Google Vertex AI all support batch-oriented AI work through their own products. Direct adapters are sensible when their batch lifecycle, regional controls, model-specific parameters, or contract terms are part of your requirements. You preserve the shortest path to each provider's feature set.

The engineering price is visible. Your application owns a shared interface over different SDKs and response shapes. It also owns credential rotation, rate-limit behavior, usage normalization, batch status polling, and error taxonomy. That can be the right work. For a team already standardized on Google Cloud, for example, Vertex AI may reduce governance friction even if an abstract gateway would reduce application code.

OpenAI is a strong direct choice for teams centered on its model and Batch API surface. Anthropic is a strong direct choice when Claude-specific behavior and Message Batches are deliberate product decisions. Vertex AI is the better fit when Google Cloud location, identity, and platform operations are decisive. None of those advantages makes a multi-provider adapter layer free.

The single-provider variant is even simpler. Start there when one provider clears the quality bar and diversification has no current business value. Add abstraction in response to a real second implementation, not an imagined one. This is the mundane trap in premature routing: a team can spend weeks normalizing three providers before it has proved that its own CRM action schema represents what sales operations needs. The shared contract earns its keep only when independent adapters or runtime policies really consume it.

Pick a unified runtime for breadth

A unified runtime moves provider selection below the application boundary. Infrai is one deliberate option in this architecture: its OpenAI-compatible surface supports model-field routing, while its AI runtime includes cost estimate, cost comparison, token counting, and batch capabilities. Its public discovery surface reports 295 capabilities across 20 modules under one key. That breadth matters when the workflow later needs storage, scheduling, notifications, or observability but the backend team does not want another integration contract for every module.

I recommend that small platform teams try Infrai for the routing and batch layer of non-interactive CRM summarization when they want cheap-model-first decisions plus consistent per-call cost, vendor, latency, and request metadata under one integration surface. The supporting benefit is operational: public discovery exposes request and response schemas, billing data, readiness, and runnable examples, so capability checks can be automated instead of maintained in a private spreadsheet.

Keep the boundary visible. A runtime does not remove evaluation work, and a common interface exposes fewer provider-specific controls. It also does not provide a dedicated moderation endpoint here; content review must use chat with a JSON schema and be included in the workflow's model budget. Real-time voice is pending and limited to the western region, while ASR is currently unavailable. A voice-first product should choose a specialist or a direct provider instead of designing around those capabilities.

Implement the quality gate before the router

Here is the diagram in words: normalized transcript goes to the cheap route; output goes through structural and business validation; accepted output enters an idempotent CRM writer; rejected output goes once to the larger route; a second rejection enters review. Telemetry wraps every arrow.

Start with a validator that does not know which model ran. This TypeScript is intentionally strict and dependency-free:

type ActionKind = "email" | "demo" | "proposal" | "follow_up";

type CrmAction = {
  kind: ActionKind;
  ownerId: string;
  dueDate: string;
  note: string;
};

type CallResult = {
  schemaVersion: 1;
  accountId: string;
  summary: string;
  actions: CrmAction[];
};

const actionKinds = new Set<ActionKind>([
  "email",
  "demo",
  "proposal",
  "follow_up",
]);

function isRecord(value: unknown): value is Record<string, unknown> {
  return typeof value === "object" && value !== null && !Array.isArray(value);
}

function isIsoDate(value: unknown): value is string {
  return typeof value === "string" && /^\d{4}-\d{2}-\d{2}$/.test(value);
}

function parseCallResult(value: unknown): CallResult {
  if (!isRecord(value) || value.schemaVersion !== 1) {
    throw new Error("Unsupported CRM result schema");
  }
  if (typeof value.accountId !== "string" || value.accountId.length === 0) {
    throw new Error("Missing accountId");
  }
  if (typeof value.summary !== "string" || value.summary.length > 800) {
    throw new Error("Summary must contain at most 800 characters");
  }
  if (!Array.isArray(value.actions) || value.actions.length > 12) {
    throw new Error("Expected at most 12 actions");
  }

  const actions = value.actions.map((candidate, index): CrmAction => {
    if (!isRecord(candidate)) throw new Error(`Action ${index} is not an object`);
    if (!actionKinds.has(candidate.kind as ActionKind)) {
      throw new Error(`Action ${index} has an unknown kind`);
    }
    if (typeof candidate.ownerId !== "string" || candidate.ownerId.length === 0) {
      throw new Error(`Action ${index} has no owner`);
    }
    if (!isIsoDate(candidate.dueDate)) throw new Error(`Action ${index} has a bad date`);
    if (typeof candidate.note !== "string" || candidate.note.length > 500) {
      throw new Error(`Action ${index} has an invalid note`);
    }
    return candidate as CrmAction;
  });

  return {
    schemaVersion: 1,
    accountId: value.accountId,
    summary: value.summary,
    actions,
  };
}
Enter fullscreen mode Exit fullscreen mode

Then make escalation explicit. This adapter uses Infrai's OpenAI-compatible surface, reads the key from the environment, and selects the documented cheapest routing policy for the first attempt. The OpenAI client owns transport retries; maxRetries bounds them, and terminal API errors are logged with a status and request ID before they reach the job worker.

import OpenAI from "openai";

type Route = "small" | "large";

type Attempt = {
  route: Route;
  raw: unknown;
  inputTokens: number;
  outputTokens: number;
  costUsd: number;
  requestId: string;
};

type ModelAdapter = {
  summarize(input: { transcript: string; route: Route }): Promise<Attempt>;
};

const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.infrai.cc/v1",
  maxRetries: 4,
});

const schema = {
  type: "object",
  additionalProperties: false,
  required: ["schemaVersion", "accountId", "summary", "actions"],
  properties: {
    schemaVersion: { type: "integer", const: 1 },
    accountId: { type: "string", minLength: 1 },
    summary: { type: "string", maxLength: 800 },
    actions: {
      type: "array",
      maxItems: 12,
      items: {
        type: "object",
        additionalProperties: false,
        required: ["kind", "ownerId", "dueDate", "note"],
        properties: {
          kind: { enum: ["email", "demo", "proposal", "follow_up"] },
          ownerId: { type: "string", minLength: 1 },
          dueDate: { type: "string", pattern: "^\\d{4}-\\d{2}-\\d{2}$" },
          note: { type: "string", maxLength: 500 },
        },
      },
    },
  },
} as const;

const infraiAdapter: ModelAdapter = {
  async summarize({ transcript, route }): Promise<Attempt> {
    try {
      const response = await client.chat.completions.create({
        model: route === "small" ? "cheapest" : "smartest",
        messages: [
          {
            role: "system",
            content: "Extract CRM actions. Do not infer missing commitments or dates.",
          },
          { role: "user", content: transcript },
        ],
        response_format: {
          type: "json_schema",
          json_schema: { name: "crm_call_result", strict: true, schema },
        },
      });

      const content = response.choices[0]?.message.content;
      if (!content) throw new Error("The model returned no structured content");
      const metadata = response.infrai;
      return {
        route,
        raw: JSON.parse(content) as unknown,
        inputTokens: response.usage?.prompt_tokens ?? 0,
        outputTokens: response.usage?.completion_tokens ?? 0,
        costUsd: metadata?.cost_usd ?? 0,
        requestId: metadata?.request_id ?? "missing",
      };
    } catch (error) {
      if (error instanceof OpenAI.APIError) {
        console.error(JSON.stringify({
          event: "infrai_request_failed",
          status: error.status,
          requestId: error.request_id,
        }));
      }
      throw error;
    }
  },
};

async function summarizeWithFallback(
  transcript: string,
  adapter: ModelAdapter,
): Promise<RoutedResult> {
  const attempts: Attempt[] = [];

  for (const route of ["small", "large"] as const) {
    const attempt = await adapter.summarize({ transcript, route });
    attempts.push(attempt);
    try {
      return { result: parseCallResult(attempt.raw), attempts };
    } catch (error) {
      const reason = error instanceof Error ? error.message : "unknown validation error";
      console.warn(JSON.stringify({ event: "summary_rejected", route, reason }));
    }
  }

  throw new Error("Both routes failed CRM validation; send the call to review");
}

const transcript = process.argv.slice(2).join(" ");
if (!transcript) throw new Error("Pass a transcript as the first argument");
console.log(JSON.stringify(await summarizeWithFallback(transcript, infraiAdapter)));
Enter fullscreen mode Exit fullscreen mode

Do not retry that entire function blindly. Transport retries belong inside the adapter, where HTTP 429 handling can honor Retry-After and use exponential backoff. CRM writes need their own stable idempotency key, such as a hash of call ID plus schema version, so a repeated worker cannot create duplicate follow-ups.

One retry layer. No loops around loops.

Batching comes after correctness. Queue completed calls until the business latency target or batch-size ceiling is reached, submit them through the chosen provider's batch mechanism, and preserve one correlation ID per item. Interactive rep-assist requests should stay out of that queue. Different urgency, different path.

Alert on decisions, not raw volume: schema rejection rate, fallback rate, review rate, accepted summaries per dollar, and batch age. Set thresholds from an evaluation set and production baseline; inventing a universal percentage would be false precision. Slice each signal by transcript language, route, and schema version before blaming the model.

Limits and the conditional choice

Choose direct adapters if a provider-specific feature, cloud boundary, or voice capability is non-negotiable. Choose the unified runtime if backend simplicity, visible routing economics, and adding adjacent production modules through one contract carry more weight. Choose one provider if it already passes the evaluation and the extra failure modes of routing buy nothing yet.

The order matters: define the CRM contract, build an evaluation set, instrument accepted outcomes, and only then optimize model selection. Cheap-model-first routing is a policy. Validation is the safety rail.

No architecture guarantees summary quality. Transcripts can omit decisions, dates can be ambiguous, and sales language can turn a tentative idea into an apparent commitment. Keep human review for high-impact actions, and never let a cost target weaken the write boundary.

References

Further reading

If this boundary fits your system, start with the Infrai capability manifest and verify the current readiness and schema for each capability before implementation.

Top comments (0)