Skip to content
All studies

Payments / Independent study

Payment Orchestration Platform

A timeout is not a failed payment.

Separating payment intent from provider attempts, so retries and late confirmations don't create a second charge.

Scope

Architecture proposal, OpenAPI sketch and tested browser model. No production backend or employer data.

Proposed backend: Java · Spring Boot · PostgreSQL · AWS SQS

Executable model: TypeScript

The problem

A checkout request crosses a boundary the application cannot control. If a provider times out after accepting it, blindly retrying through another provider can turn a network problem into a duplicate charge.

The approach

Give the payment intent one owner. Persist the merchant-scoped idempotency key and request fingerprint before starting a provider attempt. Treat an ambiguous response as awaiting confirmation, then use a verified webhook or reconciliation result to resolve it.

Architecture

One intent. Explicit uncertainty.

Proposed system boundariessync / durable / async
  1. 01

    Merchant API

    Authenticate merchant; validate amount and key.

  2. 02

    Payment domain

    Own intent, attempt history and legal transitions.

  3. 03

    Provider adapter

    Translate requests; classify definite vs ambiguous failures.

Durable boundary

PostgreSQL: payment intent + request fingerprint + attempt + outbox record. Unique merchant/key constraint. Never hold a database transaction open across a provider call.

Asynchronous work

Outbox relay → SQS → attempt worker. Verified webhooks enter an inbox; reconciliation resolves missing confirmations. Each consumer deduplicates independently.

Logical boundaries, not a deployed topology. The interactive model below covers state transitions only.

Exercise the failure path

The response was lost. Now what?

Submit a simulated CAD 75.00 payment, retry the request, then deliver its provider confirmation twice.

Interactive model
Payment state
ready
Provider attempts
0
Confirmed amount
Not confirmed

key: checkout-001 / event: evt-001 / currency: CAD

Ready. Submit the payment to simulate a provider timeout.

Model limits: Single payment, in-memory TypeScript model. The webhook is assumed verified. No provider calls, persistence, concurrency or signature verification.

API design

An explicit command boundary.

Download OpenAPI sketch

Illustrative request and response. This contract specifies one command, not the entire proposed API.

POST /v1/payments

Idempotency-Key: command_demo

{
  "amountMinor": 7500,
  "currency": "CAD",
  "providerToken": "tok_demo"
}

Response

202 Accepted
Location: /v1/payments/pay_demo

{ "id": "pay_demo", "status": "awaiting_confirmation" }

Technical decisions

What I would choose. What it costs.

01Separate intent from attempts

A merchant payment can have several traceable attempts, but an unresolved attempt prevents blind rerouting.

Trade-off: Lower apparent availability during uncertainty; operational reconciliation is required.

02Provider-specific idempotency

Reuse a provider key for the same attempt, only where the provider guarantees that behaviour. Store the relationship locally.

Trade-off: Providers have different key retention and retry contracts. An abstraction cannot erase those differences.

03Start with a modular service

Keep domain, adapter and webhook modules explicit inside one Spring Boot service before introducing more deployable services.

Trade-off: Less independent scaling initially, but simpler transaction boundaries and incident investigation.

Operating the system

What needs attention in production.

Unknown outcomes

Alert on age of awaiting-confirmation payments, not just HTTP error rates. Route unresolved items to a reconciliation queue with evidence.

Safe recovery

Bound retries for definitively retryable failures. Dead-letter exhaustion for review. Never replay a financial operation without checking its current state.

Trust boundary

Verify provider signatures and replay windows before inbox insertion. Redact tokens and sensitive payloads; retain correlation IDs and audit metadata.

Outcome & limits

What this study establishes.

The model keeps the provider attempt count at one when the request is retried, applies a confirmation once, and rejects reuse of an intent with a different amount. Those are tested local invariants, not a production reliability claim.

View study source

Next steps

Before calling it production-ready.

  • Implement PostgreSQL constraints and concurrency tests against a provider stub.
  • Test crash recovery around the outbox relay and webhook inbox.
  • Define retention, reconciliation SLAs and per-provider retry policies before a real integration.

Start a conversation

Building a backend team?

I'm interested in senior backend and Tech Lead roles across fintech, payments, SaaS and platform engineering in Canada.

Next study: Money Movement & Ledger