Temporal Workflow for Microservices: Why We Replaced 40 CronJobs With One Engine
We replaced 40 application-level cron jobs + 12 retry-prone Lambda functions with a Temporal Workflow platform. Here's what we kept, what we replaced, and the one thing Temporal can't fix.
The chaos we had
Every team owned their own scheduler. Some used node-cron, some used gocron, two teams used AWS Lambda + EventBridge. Failure modes were everywhere:
- Lambda retry storms hammered our downstream APIs.
-
node-cronjobs died silently when pods restarted. - A "retry every 5 minutes" job kept retrying after success due to a bug in the implementation.
We standardized on Temporal in February. Six months in, the operational picture is dramatically simpler.
The architecture
Trigger sources (cron, Kafka, S3 event, HTTP webhook)
↓
Temporal Worker pool (Go + TypeScript, 3 namespaces: payments/orders/analytics)
↓
PostgreSQL (Temporal persistence backend)
↓
Elasticsearch (workflow visibility)
One cluster, three namespaces for blast radius. Workers scale via Kubernetes HPA on workflow queue depth.
The 4 patterns that worked
1. Saga pattern for multi-step transactions
func OrderSagaWorkflow(ctx workflow.Context, order Order) error {
var inv Reservation
if err := workflow.ExecuteActivity(ctx, ReserveInventory, order).Get(ctx, &inv); err != nil {
return err
}
defer func() {
if err := workflow.ExecuteActivity(ctx, ReleaseInventory, inv).Get(ctx, nil); err != nil {
workflow.GetLogger(ctx).Error("inventory leak", "err", err)
}
}()
var charge Charge
if err := workflow.ExecuteActivity(ctx, ChargePayment, order).Get(ctx, &charge); err != nil {
return err
}
return workflow.ExecuteActivity(ctx, ShipOrder, order, charge).Get(ctx, nil)
}
Temporal handles the saga compensating action if any step fails.
2. Long-running business process
Our refund workflow can take 30 days (return window). Temporal: a single workflow with workflow.Sleep(ctx, 30 * 24 * time.Hour). Survives restarts, replays deterministically.
3. Cron-style scheduled work
func DailyReportWorkflow(ctx workflow.Context) error {
ao := workflow.ActivityOptions{StartToCloseTimeout: 10 * time.Minute}
ctx = workflow.WithActivityOptions(ctx, ao)
return workflow.ExecuteActivity(ctx, GenerateDailyReport).Get(ctx, nil)
}
Cron schedule: 0 2 * * * (daily 02:00).
4. Versioned API calls
Old code: "if header X-Api-Version >= 3 use new endpoint" scattered across services. Temporal: workflow.GetVersion(ctx, "payment-api", 1, 2) — versioned inside the workflow, not at the API edge.
What Temporal can't fix
Cold-start latency. First workflow execution takes 800ms-1.2s. After warm-up, ~50ms. We added a "keepalive" workflow that runs every 30 seconds in each namespace.
Database as bottleneck. Temporal's PostgreSQL backend becomes the bottleneck above ~50K concurrent workflows. We sharded by namespace.
Worker versioning hazards. If a worker runs an old binary while a new workflow starts, replay fails. We pinned worker versions via labels + canary deploys.
The metrics
| Metric | Before | After |
|---|---|---|
| Cron jobs in prod | 40 | 0 |
| Lambda retry storms/week | 6 | 0 |
| MTTR for failed workflows | 45 min | 8 min |
| Workflow visibility | DB queries | Temporal UI |
| State management | 12 ad-hoc DB tables | 1 Temporal namespace |
The team's productivity
The median retry-prone workflow dropped from 80 lines to 12 lines.
On testing workflows
Temporal ships with testsuite package. You can advance "virtual time" to test multi-day workflows in milliseconds. We caught 4 race conditions this way.
For dev environment testing, ScsDriver WebDAV mount tool for Windows mounts test fixture S3 buckets as Windows drives — useful when QA engineers need to replay a Temporal workflow against a real S3-captured event history.
Have you adopted Temporal, Cadence, or another orchestrator? What's the killer pattern in your stack?
Top comments (0)