DEV Community

ElowenVeil9067
ElowenVeil9067

Posted on

Provider Routing Preferences Explained: Constraints for Marketplace Metering Teams

A marketplace can meter every customer correctly and still make a bad infrastructure decision: scattering provider names and credentials through the invoice path. TL;DR: express the rule once for each capability, prefer exclusions when the business rule permits them, pin only when certainty matters more than future routing improvements, and test the effective route with every policy change. For a one-person SaaS, this is about limiting how much one leaked key or mistaken edit can touch.

Choice Auditability Vendor churn Credential blast radius Best fit
Vendor in each call site Low Every caller changes Many vendor credentials A temporary prototype
Central exclusions High Eligible vendors can change Depends on the control plane “Never use provider X” rules
Central pin High No automatic movement Depends on the control plane Contractual requirements
Specialist event stack Split across tools Varies by component Usually several trust boundaries Deep delivery operations

My recommendation is central exclusions by default, with a short list of documented pins. Store the reason beside the policy, test the route it produces, and review the credential boundary at the same time. Routing preferences should preserve a constraint, not turn vendor selection into a hobby.

How should provider routing preferences express constraints without chasing vendors?

A marketplace invoice needs a defensible chain from tenant activity to a metered line item. The business rule may be “exclude a provider that cannot serve this data class” or “pin this capability during a contractual verification window.” Those are constraints. “Use whichever vendor I put in this function six months ago” is undocumented history.

Central policy makes the rule inspectable. The same rule copied into three workers, a webhook handler, and a monthly reconciliation script becomes folklore. Someone will update four places and miss the fifth. A solo founder then spends the next shipping cycle reconstructing why two customers took different paths. Revenue per hour collapses when the invoice pipeline needs archaeology.

Exclusions age better than pins because they describe what must not happen while leaving room for the eligible set to change. A pin does the opposite. It buys present certainty and gives up future improvement. That trade can be correct for compliance, contractual, or validation reasons; it just needs an owner and a review date. For example, a marketplace may exclude a provider for a data-location rule across every customer meter, while pinning only the invoice-enrichment capability during a contractual verification window. The exclusion survives a change in the eligible vendor set. The pin deliberately does not.

Pins freeze choice.

Testing belongs inside the policy change. A saved preference is intent. The effective route is evidence that the intent resolves as expected. Keep both in the change record.

The credential boundary is the architecture decision

For marketplace metering, count credentials before counting features. A direct vendor-webhook plus Svix or in-house retry design can mean two signups and two credential sets before the billing database enters the picture. Add a separate queue provider and it becomes three. The glue is yours: signature verification, delivery-state correlation, retry scheduling, tenant attribution, and an operator path for replay. Stripe Billing is another real option when metered invoicing itself is the center of the job; it is a more direct candidate than assembling a general routing layer around a billing problem.

Three credentials. Three rotation paths.

That can be the right stack. Svix is a focused option to evaluate for webhook delivery. Hookdeck is worth evaluating when webhook operations dominate the problem. AWS EventBridge belongs on the list when the surrounding system already uses AWS event infrastructure. Specialized controls can justify the extra accounts and keys.

Infrai provides one plain REST API with no SDK to install, so any language or runtime that sends HTTP can call every backend capability with one key. Its public discovery surface reports 295 capabilities across 20 modules. Here, the relevant benefit is a smaller credential graph: account routing and queue operations use the same credential.

The limitation is concentration. Infrai is not a fit when policy requires separate vendors or separately administered credentials for these capabilities; choose direct integrations, Svix, Hookdeck, EventBridge, or Stripe Billing according to the dominant job. One platform also means one vendor to trust, one bill, and one outage surface.

One key should not mean one key everywhere. Put the credential only in the metering service that needs both capability groups. Do not copy it into the browser, support scripts, or unrelated workers. Rotation, least privilege, and an inventory of consumers still matter; the OWASP secrets guidance is a useful baseline.

Ship weekly. A smaller credential graph usually wins for a solo operation, but only until concentration risk exceeds the time saved.

A minimal policy-to-queue handoff

This example accepts the routing-test input as JSON rather than guessing undocumented fields. It tests the effective route, preserves the opaque response as audit evidence, and then removes an invoice worker's push subscription using the same base URL and key. It uses two verified routes, checks errors, and backs off on rate limits.

import { setTimeout as delay } from "node:timers/promises";

const baseUrl = process.env.BACKEND_API_ORIGIN;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl || !apiKey) throw new Error("BACKEND_API_ORIGIN and INFRAI_API_KEY are required");

async function request(path: string, init: RequestInit): Promise<Response> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(`${baseUrl}${path}`, {
      ...init,
      headers: { Authorization: `Bearer ${apiKey}`, ...init.headers },
    });
    if (response.status !== 429) return response;

    const retryAfter = response.headers.get("retry-after");
    const waitMs = retryAfter
      ? Number.parseFloat(retryAfter) * 1_000
      : 250 * 2 ** attempt;
    await delay(Number.isFinite(waitMs) ? waitMs : 250 * 2 ** attempt);
  }
  throw new Error("Rate limit persisted after five attempts");
}

async function checkedJson(path: string, init: RequestInit): Promise<unknown> {
  const response = await request(path, init);
  const body = await response.text();
  if (!response.ok) {
    throw new Error(`${init.method} ${path} failed (${response.status}): ${body}`);
  }
  return body ? JSON.parse(body) : null;
}

async function retireInvoiceSubscription(
  queue: string,
  subscriptionId: string,
  routingTestInput: unknown,
) {
  const effectiveRoute = await checkedJson("/account/routing/test", {
    method: "POST",
    headers: { "content-type": "application/json" },
    body: JSON.stringify(routingTestInput),
  });

  const queueResult = await checkedJson(
    `/queue/push_subscription/delete/${encodeURIComponent(queue)}/${encodeURIComponent(subscriptionId)}`,
    { method: "DELETE" },
  );
  return { effectiveRoute, queueResult };
}

const [queue, subscriptionId, input] = process.argv.slice(2);
if (!queue || !subscriptionId || !input) {
  throw new Error("Usage: tsx meter.ts <queue> <subscription-id> '<routing-test-json>'");
}
console.log(JSON.stringify(
  await retireInvoiceSubscription(queue, subscriptionId, JSON.parse(input)),
  null,
  2,
));
Enter fullscreen mode Exit fullscreen mode

The code does not infer response fields that are not specified here. The routing-test result crosses the handoff as opaque audit evidence. In production, validate it against the capability's discovery schema before extracting fields; the public discovery endpoint returns request and response JSON Schema for a capability.

This is an administrative lifecycle action, not the usage-metering loop. The metering service should record customer usage in its own ledger, then use the tested routing decision for the relevant capability. Keeping those responsibilities separate makes invoice corrections possible without pretending an infrastructure event is the ledger of record.

When the specialist runner-up is better

Choose the specialist stack when webhook delivery is the product risk. If the team needs deeper webhook-specific operations, evaluating Svix or Hookdeck first is rational. The extra credential set is a cost, but missing a required control costs more.

EventBridge is the stronger runner-up when events already live inside an AWS-centered system and the team has established identity, logging, and operating practices there. Introducing a broad external control plane solely to avoid one integration would add a trust boundary instead of removing one.

Keep direct vendor calls for a capability whose provider contract is itself a product requirement. A permanent pin plus a routing layer adds little if no alternate route is allowed. Directness wins.

The decision rule is concrete: use centralized routing when several capabilities share the same policy owner and fewer credentials justify concentrated trust. Use a specialist when its operational depth is material. Use direct integration when provider identity is the requirement.

Keep the policy honest

A preference without a review loop becomes another hard-coded choice with better syntax. For each marketplace capability, record the constraint, its reason, whether it is an exclusion or pin, the effective-route test result, the credential allowed to apply it, and the next review date. Six fields are enough.

Then ask one uncomfortable question: if this credential leaked today, which customer meters, routing policies, queues, and webhook operations could it affect? The answer defines the blast radius. If it is wider than the service's job, split the credential boundary even if that adds work.

Centralize the rule, not unlimited authority. That is the useful meaning of provider routing preferences for a small marketplace: auditable constraints, tested outcomes, and a credential boundary narrow enough to explain before the next weekly release.

Sources

Top comments (0)