Author: Mohammad Shakhtour

Status: Accepted

Context

The Problem

ADR-003 established that we use IIntegrationEventHandler for asynchronous event processing outside the transaction scope and in async way. But we also need to guarantee that integration events are reliably delivered even in the face of system failures?

Consider this failure scenario:

1. Transaction begins
2. Work order completed (database updated)
3. Transaction commits successfully
4. Send email notification via external API
5. Application crashes before email is sent

Result: Work order is completed in the database, but the notification email is never sent and there's no record that it should have been.

The Dual-Write Problem

This is known as the dual-write problem in distributed systems:

> "If you need to update two different systems (a database and a message broker), you cannot make both updates atomically. One might succeed while the other fails."

> — Chris Richardson, Microservices Patterns (2018)

Traditional Approaches Fail:

Approach 1: Publish then Commit

1. Publish event to message bus
2. Commit database transaction
Problem: If database commit fails, event already published

Approach 2: Commit then Publish

1. Commit database transaction  
2. Publish event to message bus
Problem: If publish fails or app crashes, event is lost

Approach 3: Two-Phase Commit

1. Prepare both database and message bus
2. Commit both atomically
Problem: Complex, poor performance, requires distributed transaction coordinator

None of these approaches guarantee both atomicity and reliable delivery.

Why This Matters

The Outbox Pattern emerged as the industry-standard solution to this problem, documented by Chris Richardson:

> "The Transactional Outbox pattern uses the database as a temporary message queue. Services that send messages insert them into an outbox table as part of the database transaction that updates business entities."

> — Chris Richardson, Microservices Patterns (2018)

---

Decision Drivers

  1. Guaranteed Delivery - Integration events must not be lost, even during system failures
  2. Transactional Consistency - Event publishing must be atomic with business state changes
  3. System Resilience - Must survive application crashes and restarts
  4. At-Least-Once Semantics - Events delivered at least once (idempotency handled by consumers)
  5. Retry Capability - Failed deliveries automatically retried with backoff
  6. Microservices Readiness - Supports evolution to distributed architecture with message bus

---

Options Considered

Option 1: Fire-and-Forget (Task.Run)

Pattern:

Execute integration event handlers in background tasks without persistence.

Pros:

Cons:

Verdict: Rejected - No reliability guarantees

---

Option 2: In-Memory Queue with Background Service

Pattern:

Queue events in memory, background service processes them.

Pros:

Cons:

Verdict: Rejected - Not restart-safe

---

Option 3: Transactional Outbox Pattern (Chosen)

Pattern:

Write integration events to a database table within the same transaction as the business operation. A separate background process reads from this outbox table and publishes events.

Business Transaction:
1. Update work order status
2. Write integration event to Outbox table
3. Commit transaction (both updates atomic)

Background Process:
1. Poll Outbox table for unprocessed events
2. Publish event to handlers/message bus
3. Mark event as processed
4. Retry on failure

Pros:

Cons:

Verdict: Accepted - Industry-standard solution balancing reliability and complexity

---

Decision

Implement the Transactional Outbox Pattern with the following design:

Core Components

1. Outbox Table

Stores integration events within the same database as business entities

2. Event Writer (DomainEventsPipeline)

Writes integration events to outbox within the command's transaction

3. Outbox Processor (Background Service)

Polls outbox table and processes undelivered events:

4. Outbox Store (Repository)

Abstracts database operations for the outbox

Separate Database Context

The outbox uses a dedicated OutboxDbContext separate from the main application context:

Why Separate?

---

Key Design Decisions

Serialization Strategy:

Polling vs Push:

Retry Strategy:

Idempotency:

---

Future Evolution

Phase 1: Default Implementation (Simple)

Characteristics:

Key Point: The IOutboxStore abstraction makes this a simple starting point that's easy to replace later.

Phase 2: Message Bus Integration

Changes:

Phase 3: Microservices

Characteristics:

Migration Path

No Breaking Changes:

---

Operational Considerations

Monitoring

Key Metrics:

Alerts:

Maintenance

Cleanup Strategy:

Successfully processed events should be archived/deleted:

Dead Letter Queue (Future):

Events exceeding max retries moved to dead letter table:

---

Guidelines

Writing Integration Event Handlers

Handler Requirements:

Event Design

Best Practices:

---