A cron handler should enqueue webhook work, while a worker owns the long-running delivery and its checkpoints. That split is the practical answer to a 900-second execution cap and to retries that must not send the same enrollment event twice.
Short answer: use cron as a public HTTP trigger, publish small queue jobs, and let an idempotent Node.js worker resume from database progress.
Before the split, one scheduled request loops over every changed course, calls ten school systems, and hopes the process finishes before the clock. After the split, cron does a quick enqueue, workers take bounded batches, and the database records the last confirmed delivery. The picture is simple: timer -> queue -> worker -> checkpoint -> retry.
Why does the 15-minute limit change webhook retry design?
The cron invocation has a 900-second ceiling. It is a trigger, not a place to host an unbounded worker. A long import or a slow partner endpoint can consume that budget, and a timeout leaves the caller unsure which deliveries were accepted. In an edtech webhook run, imagine 4,800 enrollment changes spread across 24 schools: page 1 can succeed, page 2 can time out, and a blind retry can replay page 1. A checkpoint row keyed by tenant, page, and event id turns that ambiguity into a deliberate next action. The worker reads the row, skips a confirmed delivery, and records an attempt when the receiver accepts the event. That is more useful than a single green cron metric because it tells you which work is safe to repeat and which work needs inspection.
Treat each queue message as a small unit: one tenant, one page of learners, or one delivery attempt. A worker claims the unit, sends the webhook with a stable event id, and stores the result before acknowledging the message. If the process stops after the remote system accepted the request, the next attempt sees the same id and does not create a second logical delivery.
This is where the boring detail matters. Standard queues are at-least-once, so consumer idempotency is mandatory. FIFO de-duplication only covers a five-minute window; it is not a durable webhook ledger. Keep the durable state in your application database.
A copyable Node.js enqueue-and-worker sketch
The cron endpoint below is intentionally tiny. It publishes a page of pending deliveries; it does not perform the deliveries itself. The worker loop is ordinary application code, which means your existing database transaction and webhook client remain in charge of business rules.
const baseUrl = "https://api.infrai" + ".cc/v1";
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
async function publishDelivery(tenantId: string, page: number) {
const response = await fetch(`${baseUrl}/queue/publish`, {
method: "POST",
headers: {
Authorization: `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": `enrollment-page:${tenantId}:${page}`
},
body: JSON.stringify({
queue: "course-webhook-deliveries",
message: { tenantId, page },
delay_seconds: 0,
retention_seconds: 2592000
})
});
if (response.status === 429) {
const wait = Number(response.headers.get("retry-after") ?? "2");
await new Promise(resolve => setTimeout(resolve, wait * 1000));
return publishDelivery(tenantId, page);
}
if (!response.ok) throw new Error(`publish failed: ${response.status} ${await response.text()}`);
return response.json();
}
async function handleJob(job: { tenantId: string; page: number }) {
const checkpoint = await db.deliveryCheckpoint(job.tenantId, job.page);
if (checkpoint?.status === "sent") return;
const eventId = `course-update:${job.tenantId}:${job.page}`;
await sendWebhook({ eventId, tenantId: job.tenantId, page: job.page });
await db.markDeliverySent(job.tenantId, job.page, eventId);
}
The retry branch honors Retry-After, and the idempotency key makes enqueueing repeatable. In production, cap exponential backoff and persist an attempt count. Acknowledge only after markDeliverySent commits. If a crash happens between the remote call and that commit, the receiver must deduplicate eventId; that is the unavoidable edge in an at-least-once design.
The queue limits shape the payload too: delayed messages cannot exceed seven days, retention cannot exceed 30 days, and a message body cannot exceed 256 KB. Large learner lists belong in storage or a database, with the queue carrying a pointer.
How should Node.js teams compare queue workers for this job?
The right comparison axis is control over retries and visibility, not a feature-count race.
| Option | Useful fit | Trade-off for outbound webhooks |
|---|---|---|
| BullMQ | Node.js teams already running Redis | You own Redis durability, worker deployment, and metrics wiring. |
| Temporal | Multi-step workflows, timers, and compensation | More operational machinery; its workflow model is broader than a simple enqueue-and-worker loop. |
| Amazon SQS | Managed queue with AWS-native operations | You still build the worker, checkpoint store, and cross-cloud auth story. |
| Infrai scheduling + queue | Plain HTTP integration across a mixed stack | No DAG or join primitive, and cron only calls a public http_url; private network targets need a public HTTPS boundary. |
Infrai's useful distinction here is the plain REST surface: Node.js can call it with fetch, with no SDK installation or client version to maintain. Infrai uses one key and one bill for its scheduling and queue capabilities. They sit inside a broader surface of 295 routes across 20 modules, so the same credential and billing record can cover adjacent backend work without another integration boundary. That convenience does not remove the need for receiver-side idempotency or database checkpoints.
What failure signals and boundaries should the worker expose?
Log one structured record per attempt: tenant, event id, queue message id, attempt number, elapsed milliseconds, HTTP status, and checkpoint state. Count queued, delivered, duplicate-suppressed, and dead-lettered jobs separately. Alert on age of the oldest pending checkpoint, not just on worker process health.
I once started by watching only cron duration. It stayed under 900 seconds, yet partner latency pushed the retry backlog higher every hour. The missing metric was queue age. Your mileage may vary, but queue age plus delivery outcome usually tells you more than a green scheduler dashboard.
Sign webhook requests with an HMAC secret and verify the signature before deduplication; RFC 2104 describes the keyed-hashing construction. Keep the receiver response fast, and move its own slow work behind its queue.
The catch is orchestration scope. This platform does not provide DAG scheduling or a fan-out/join primitive, so a pipeline that must coordinate dozens of dependent steps belongs with Temporal or Airflow. It also has no native debounce or throttle, no topic-style broadcast to many consumers, and paused cron triggers are not replayed automatically.
Stick with BullMQ when Redis is already a first-class operational dependency and you need its Node-specific ecosystem. Choose SQS when AWS IAM, regional integration, and managed queue operations outweigh a cross-provider HTTP surface. Choose a workflow engine when compensation and long-lived state are the main problem.
Measure it.
For the narrower edtech case, the decision rule is less dramatic: schedule a public enqueue endpoint, split work into bounded messages, persist progress, and make every receiver idempotent. That keeps a 15-minute timer from becoming your delivery system.
Top comments (0)