DEV Community

Ed Legaspi
Ed Legaspi

Posted on Originally published at czetsuyatech.com

Building Long-Running Payment Workflows with Spring Boot and NERV Event

Building Long-Running Payment Workflows with Spring Boot and NERV Event

Some payment operations finish in milliseconds.

Others don't.

Imagine a payment service that sends a request to an external payment provider.

Normally, the provider responds quickly.

But sometimes that operation can take several minutes — perhaps even 30 minutes.

The obvious implementation might look something like this:

@Transactional
public Payment processPayment(PaymentRequest request) {

    Payment payment = createPayment(request);

    GatewayResult result = paymentGateway.validate(payment);

    payment.complete(result);

    return payment;
}
Enter fullscreen mode Exit fullscreen mode

It looks safe.

The method is transactional.

If something fails, Spring rolls everything back.

But there is a problem hiding inside that transaction.

What happens if paymentGateway.validate() takes 30 minutes?

Your business workflow is now long-running.

And potentially, so is your database transaction.

That is usually not what we want.

The solution isn't simply increasing the timeout.

The better question is:

Should the lifetime of a database transaction really be the same as the lifetime of the business workflow?

For distributed payment workflows, the answer is usually no.


A 30-Minute Workflow Shouldn't Mean a 30-Minute Transaction

Consider what our original implementation is effectively doing:

BEGIN TRANSACTION

    Create Payment

    Call External Payment Gateway

    Wait...

    Wait...

    Wait...

    30 minutes later...

    Update Payment

COMMIT
Enter fullscreen mode Exit fullscreen mode

During those 30 minutes, the application may be holding database resources while doing no database work at all.

Depending on the transaction, this can contribute to:

  • long-held database connections
  • connection-pool pressure
  • locks being held longer than necessary
  • long-running MVCC snapshots
  • increased contention
  • harder failure recovery

But there is an even more fundamental problem.

The external payment provider does not participate in your database transaction anyway.

Rolling back PostgreSQL cannot undo an external payment request that has already been accepted.

So what exactly are we protecting by keeping the transaction open?

This is where the architecture needs to change.


Separate the Workflow from the Transaction

A long-running business workflow does not need to be implemented as one long-running database transaction.

Instead, we can divide it into short, durable steps.

The first transaction creates the payment and records what needs to happen next.

┌───────────────────────────────┐
│       Transaction #1          │
│                               │
│   Create Payment              │
│          +                    │
│   Create Outbox Event         │
│                               │
└──────────────┬────────────────┘
               │
             COMMIT
               │
               ▼
         Payment exists
Enter fullscreen mode Exit fullscreen mode

The payment and the event are stored atomically.

If the transaction rolls back:

Payment:       NOT CREATED
Outbox Event:  NOT CREATED
Enter fullscreen mode Exit fullscreen mode

If it commits:

Payment:       CREATED
Outbox Event:  CREATED
Enter fullscreen mode Exit fullscreen mode

We now have something extremely useful:

a durable record of what needs to happen next.


Commit Before Crossing the Network

This leads to a useful rule for long-running workflows:

Commit local state before performing slow remote work whenever the business semantics allow it.

Instead of:

BEGIN TX
   │
   ├── Save Payment
   │
   ├── Call External System
   │       │
   │       └── wait 30 minutes
   │
   └── Update Payment
COMMIT
Enter fullscreen mode Exit fullscreen mode

we move toward:

TX #1
   │
   ├── Save Payment
   ├── Save Outbox Event
   │
COMMIT
   │
   ▼
External Work
   │
   │ possibly 30 minutes
   ▼
TX #2
   │
   ├── Update Payment
   ├── Save Next Event
   │
COMMIT
Enter fullscreen mode Exit fullscreen mode

The business workflow can still take 30 minutes.

The individual database transactions do not.

That distinction is critical.


Let the Event Move the Workflow Forward

After the first transaction commits, NERV Event can dispatch the Outbox event.

The architecture becomes:

Client
  │
  ▼
Payment Service
  │
  │ Transaction #1
  ▼
Payment + Outbox
  │
  │ COMMIT
  ▼
Outbox Dispatcher
  │
  ▼
Message Broker
  │
  ▼
Payment Gateway Service
  │
  ▼
External Payment Provider
Enter fullscreen mode Exit fullscreen mode

The payment might now have an explicit state such as:

PENDING_VALIDATION
Enter fullscreen mode Exit fullscreen mode

That state tells us:

The payment exists, but the workflow has not finished.

This is an important architectural shift.

We stop using an open transaction to represent progress.

Instead, we use durable business state.


Long-Running External Work Does Not Need a Database Transaction

Suppose the Payment Gateway Service now needs to call the external provider.

If this operation can take a long time, we don't want an unnecessary database transaction surrounding it.

In Spring, one option is to make that boundary explicit:

@Transactional(propagation = Propagation.NOT_SUPPORTED)
public GatewayResult validate(PaymentRequest payment) {
    return externalGateway.validate(payment);
}
Enter fullscreen mode Exit fullscreen mode

NOT_SUPPORTED means the method should execute without an active transaction.

If a transaction exists, Spring suspends it while this method executes.

The distinction is important:

Short Database Work
        ≠
Long Network Operation
Enter fullscreen mode Exit fullscreen mode

The application may still have an HTTP request waiting.

A thread may still be occupied depending on the implementation.

The external operation may still take 30 minutes.

But we do not need to keep a database transaction open simply because the workflow is still running.


When the Provider Responds, Start Another Short Transaction

Eventually, the external provider returns a result.

Now we need another local transaction.

For example:

┌───────────────────────────────┐
│       Transaction #2          │
│                               │
│   Update Gateway State        │
│          +                    │
│   Create Result Event         │
│                               │
└──────────────┬────────────────┘
               │
             COMMIT
Enter fullscreen mode Exit fullscreen mode

The complete flow now looks like this:

TX #1
Create Payment
Create Outbox Event
      │
      ▼
    COMMIT
      │
      ▼
External Gateway Call
NO DATABASE TRANSACTION
      │
      │ possibly 30 minutes
      ▼
Gateway Response
      │
      ▼
TX #2
Update State
Create Outbox Event
      │
      ▼
    COMMIT
Enter fullscreen mode Exit fullscreen mode

The transaction lifetime now reflects the database work.

The workflow lifetime reflects the business process.

They no longer need to be the same.


But Now We Can't Roll Everything Back

Correct.

And that's one of the most important realizations when designing distributed workflows.

Suppose this happens:

Payment created
      │
      ▼
Transaction committed
      │
      ▼
External provider unavailable
Enter fullscreen mode Exit fullscreen mode

We cannot roll back the original payment transaction.

It already happened.

But even if we had kept a local database transaction open, rolling it back would not necessarily undo something that had already happened in another system.

The idea of one giant rollback is often an illusion once a workflow crosses service boundaries.

Instead, we need to ask different questions:

  • What has already happened?
  • What is the current business state?
  • What still needs to happen?
  • Can the failed step be retried?
  • Is retrying safe?
  • Does compensation need to occur?

Those are workflow questions.

Not database transaction questions.


Make Partial Progress Explicit

A payment workflow might use states such as:

CREATED
   │
   ▼
PENDING_VALIDATION
   │
   ▼
VALIDATING
   │
   ├───────────────► COMPLETED
   │
   └───────────────► FAILED
Enter fullscreen mode Exit fullscreen mode

The exact state machine depends on the payment domain.

The important part is that partial progress becomes explicit and durable.

Instead of the workflow existing primarily inside a Java call stack:

methodA()
   │
   ▼
methodB()
   │
   ▼
methodC()
Enter fullscreen mode Exit fullscreen mode

it exists in persisted state:

Payment State
     +
Outbox State
     +
Inbox State
Enter fullscreen mode Exit fullscreen mode

A JVM restart no longer erases our understanding of where the workflow is.


The Inbox Protects the Receiving Side

So far, we've concentrated mostly on publishing.

But reliable delivery is only half the problem.

Suppose the Payment Gateway Service receives:

PaymentValidationRequested
Enter fullscreen mode Exit fullscreen mode

The broker may deliver that message more than once.

This is normal in an at-least-once delivery architecture.

Without protection, the consumer might execute the same operation multiple times.

NERV Event uses the Inbox Pattern on the receiving side.

Conceptually:

Broker
   │
   ▼
Inbox
   │
   ▼
Handler
   │
   ▼
Business Operation
Enter fullscreen mode Exit fullscreen mode

The Inbox gives us a durable record of received messages.

It can support:

  • deduplication
  • idempotency
  • processing status
  • retry tracking
  • failure inspection

But there is an important detail.

The Inbox is not merely a table of message IDs.

Its transaction boundary matters.


Inbox + Business State + Outbox

Consider a consumer that receives one event and produces another.

For example:

PaymentValidationRequested
           │
           ▼
   Validate Payment
           │
           ▼
PaymentValidationCompleted
Enter fullscreen mode Exit fullscreen mode

The local transaction may need to include three things:

┌──────────────────────────────────┐
│        Database Transaction      │
│                                  │
│  Inbox Processing State          │
│           +                      │
│  Business State                  │
│           +                      │
│  Outbox Event                    │
│                                  │
└────────────────┬─────────────────┘
                 │
               COMMIT
Enter fullscreen mode Exit fullscreen mode

This creates a powerful reliability boundary.

The Inbox protects the incoming side.

The business transaction protects local state.

The Outbox protects the outgoing side.

Or more simply:

Inbox
  │
  ▼
Business State
  │
  ▼
Outbox
Enter fullscreen mode Exit fullscreen mode

All inside one local database transaction.

This is one of the cleanest patterns I've found for building durable event-driven workflows.


Why the Inbox Transaction Boundary Matters

Imagine this sequence:

Receive Event
    │
    ▼
Perform Business Operation
    │
    ▼
Mark Inbox PROCESSED
Enter fullscreen mode Exit fullscreen mode

Now suppose the business operation commits but updating the Inbox fails.

The broker delivers the message again.

The system may execute the business operation twice.

Now reverse the failure:

Inbox marked PROCESSED
        │
        ▼
Business transaction fails
Enter fullscreen mode Exit fullscreen mode

The infrastructure thinks the message has been processed.

But the business state says otherwise.

Both situations are dangerous.

Where possible, the business state and Inbox completion should participate in the same local transaction.

And if processing creates another event, the Outbox write belongs there too.

BEGIN

    Update Business State

    Mark Inbox Completed

    Insert Outbox Event

COMMIT
Enter fullscreen mode Exit fullscreen mode

Now either the whole processing step commits or it does not.


Durable Handoffs Are the Real Workflow

This led me to a useful way of thinking about long-running event-driven workflows.

Each service performs three responsibilities.

1. Receive responsibility

Usually through an incoming event recorded in an Inbox.

2. Perform a local unit of work

Using a short local transaction where appropriate.

3. Hand responsibility to the next step

By creating an Outbox event.

Conceptually:

Inbox
  │
  ▼
Local Work
  │
  ▼
Outbox
  │
  ▼
Next Service
Enter fullscreen mode Exit fullscreen mode

The event isn't merely a notification.

It represents a durable handoff of responsibility.

This model becomes particularly useful when a workflow spans multiple services.


This Is Not a Distributed Transaction

The architecture is deliberately not trying to do this:

BEGIN GLOBAL TRANSACTION

    Payment Database

    Gateway Database

    Kafka

    External Payment Provider

COMMIT EVERYTHING
Enter fullscreen mode Exit fullscreen mode

Instead:

Payment Service
      │
      └── Local Transaction

            │
            ▼

       Durable Event

            │
            ▼

Gateway Service
      │
      └── Local Transaction

            │
            ▼

    External Provider
      │
      └── Independent System
Enter fullscreen mode Exit fullscreen mode

Each component owns its local consistency boundary.

Durable events connect those boundaries.

That means the system accepts an unavoidable reality:

A distributed workflow can be partially complete.

Once we accept that, we can design explicitly for recovery.


What Happens If the Application Crashes?

Let's look at several failure windows.

Crash after the payment commits but before Kafka receives the event

The Outbox record still exists.

The dispatcher can publish it later.

Crash after the Gateway Service receives the event

The Inbox provides durable processing state.

External gateway temporarily fails

The retry policy determines whether and when processing should continue.

Crash after creating the result event but before publishing it

The result remains in the Outbox.

The dispatcher can recover it after restart.

The architecture does not prevent crashes.

It does something more practical:

It reduces the amount of important information that disappears when a crash happens.


Retries Require Idempotency

Reliable systems retry things.

And once retries exist, duplicate execution becomes possible.

Imagine this scenario:

Application
    │
    ├── Send payment request
    │
    ▼
External Provider
    │
    ├── Accept request
    │
    ▼
Network timeout
Enter fullscreen mode Exit fullscreen mode

From the application's perspective, the request failed.

From the provider's perspective, it may have succeeded.

Blindly retrying could now perform the operation twice.

This is why retries and idempotency must be designed together.

For events, we might use:

  • stable event IDs
  • Inbox deduplication

For external payment providers:

  • idempotency keys

For local business state:

  • valid state-transition checks
  • unique business constraints
  • processed-operation records

The important rule is:

Retries without idempotency can turn a reliability feature into a correctness bug.


Observability Becomes Part of the Workflow

Once the process becomes asynchronous, one question becomes much harder:

Where is my payment?

A synchronous request may be relatively easy to follow through a stack trace.

A long-running workflow can span:

Multiple transactions
        +
Multiple services
        +
Multiple events
        +
Retries
        +
Broker delivery
        +
External providers
        +
30 minutes of elapsed time
Enter fullscreen mode Exit fullscreen mode

We therefore need to answer operational questions such as:

  • What is the current payment state?
  • Which event started this step?
  • Was the Outbox event dispatched?
  • Did the consumer receive it?
  • What is the Inbox processing status?
  • How many times has processing been attempted?
  • Why did the previous attempt fail?
  • Is another retry scheduled?
  • Which downstream event was produced?

This is one reason I prefer durable Outbox and Inbox state over hiding all reliability behavior inside broker configuration.

Persistent state gives us an operational window into the workflow.


Putting Everything Together

The architecture eventually looks something like this:

┌────────────────────┐
│   Payment Service  │
│                    │
│ Payment + Outbox   │
└─────────┬──────────┘
          │
        COMMIT
          │
          ▼
       Broker
          │
          ▼
┌─────────────────────────┐
│ Payment Gateway Service │
│                         │
│ Inbox                   │
└────────────┬────────────┘
             │
             ▼
    External Gateway Call
    (No DB Transaction)
             │
             │ potentially long-running
             ▼
      Gateway Response
             │
             ▼
┌─────────────────────────┐
│ Short Local Transaction │
│                         │
│ Business State          │
│ Inbox Completion        │
│ Outbox Result Event     │
└────────────┬────────────┘
             │
           COMMIT
             │
             ▼
          Broker
             │
             ▼
      Payment Service
             │
             ▼
       Payment Updated
Enter fullscreen mode Exit fullscreen mode

There is no 30-minute database transaction.

There is a 30-minute business workflow composed of short transactions and durable handoffs.

That distinction is the heart of the architecture.


The Lesson Isn't Really About Payments

Payments make the problem easy to see because correctness matters so much.

But the same architecture applies to many long-running workflows:

  • order fulfillment
  • identity verification
  • document processing
  • external approval workflows
  • provisioning
  • subscription activation
  • fraud checks
  • shipping
  • asynchronous reporting

Whenever a business process crosses service boundaries and may take significant time, it is worth asking:

Are we modeling a workflow, or are we accidentally trying to stretch a database transaction across it?

Those are very different things.


What Building This Changed for Me

Spring makes @Transactional so convenient that it is easy to think about transaction boundaries in terms of methods.

Distributed workflows force us to think differently.

A transaction boundary should follow the local consistency requirement, not necessarily the entire business operation.

A payment may take 30 minutes.

That doesn't mean its transaction should take 30 minutes.

A workflow may cross five services.

That doesn't mean those services should share one transaction.

An external provider may fail after local state has already committed.

That doesn't mean the architecture is broken.

It means failure and partial progress need to become first-class parts of the design.

This is where Outbox and Inbox become much more interesting than simple messaging patterns.

Together, they let us transform:

One fragile long-running transaction
Enter fullscreen mode Exit fullscreen mode

into:

A sequence of short,
durable,
observable,
recoverable steps
Enter fullscreen mode Exit fullscreen mode

The Bigger Lesson

The goal of a reliable distributed workflow is not to make the entire system atomic.

In most real-world architectures, it can't be.

The goal is to make every transition understandable and recoverable.

Commit local state.

Record what needs to happen next.

Release the transaction.

Perform slow external work outside the transaction.

Persist the result.

Reliably hand responsibility to the next component.

Make every step observable.

For a 30-minute payment workflow, that architecture is much more valuable than simply increasing a timeout.

Because the real problem was never the 30 minutes.

The real problem was asking one transaction to represent a process that was never truly transactional in the first place.


About NERV Event

NERV Event is an open-source event-driven infrastructure library for Spring Boot focused on reliable event delivery and processing.

It provides infrastructure for:

  • Transactional Outbox
  • Inbox processing
  • retries
  • idempotency
  • ordering
  • Kafka and SQS integration
  • multi-instance processing
  • scheduler resilience
  • operational visibility

NERV Event is part of NERV — Next-Generation Engineering for Runtime Velocity.

If you're interested in the implementation, architecture, or want to contribute:

👉 NERV Event on GitHub


This article is part of my ongoing series about building production-ready event-driven systems with Spring Boot and NERV Event.

Top comments (0)