When did we forget transactions are critical sections?

I don’t know when it became fine to pass an open database transaction around like some sort of ambient value, but I was raised by grouchy neckbeards who taught me to treat it with respect. Open it, do the work that needs it, give it back. I’ve seen codebases thread open transactions through every call like a trace ID, then complain that SQL doesn’t scale.

I honestly thought we left transaction-per-request behind in the late noughties, along with XML-configured Hibernate.

I’ve been doing more Go lately, and it turns out this isn’t the settled truth I thought it was.

I worked in a Go codebase where a middleware opened a transaction for every HTTP request, stashed it in context.Context, and committed it when the handler returned. A transactor has since replaced the middleware, but it still puts the transaction in the context, and the transaction stays open as long as the caller wants.

It isn’t just one codebase. Django has ATOMIC_REQUESTS. Spring Boot enables Open Session In View by default, which holds the persistence context, often with a connection, for the whole request. Go has libraries built around carrying the transaction in ctx.

A transaction is a critical section. Keep it short, block on nothing outside the database, and give it one owner. I think the repository is where those boundaries belong.

The pattern

Starting the transaction in the handler is still common in APIs, and behind a lot of slow ones. The handler begins a transaction, calls whatever services it needs, and commits at the end. Every query, every bit of business logic and every outbound call in between runs with the transaction open.

Once the handler owns the transaction, the next problem is a service underneath opening its own. The usual answer is a transactor. That’s a reasonable tool: it owns begin, commit and rollback so nobody forgets one. This one also stores the transaction in the context and joins it if one is already there:

func (t *Transactor) RunTx(ctx context.Context, fn func(ctx context.Context) error) error {
	if _, ok := txFrom(ctx); ok {
		return fn(ctx)
	}
	tx, err := t.pool.Begin(ctx)
	if err != nil {
		return err
	}
	defer tx.Rollback(ctx)

	if err := fn(withTx(ctx, tx)); err != nil {
		return err
	}
	return tx.Commit(ctx)
}

Then you wrap whatever section needs it:

err := s.tx.RunTx(ctx, func(ctx context.Context) error {
	if err := s.accounts.Debit(ctx, accountID, amount); err != nil {
		return err
	}
	order, err := s.broker.PlaceOrder(ctx, req)
	if err != nil {
		return err
	}
	return s.orders.Save(ctx, order)
})

It reads well, compiles, passes the tests, and works in a quiet environment. But PlaceOrder is a network call to another system, and it runs while the transaction is open.

accounts.Debit calls RunTx itself and expects a short transaction of its own. Inside this closure it has joined the caller’s, and nothing at the call site can tell the difference.

Wrapping sections in a transactor and storing the transaction in the context is a hack around lax transaction management. It flattens nested transactions into one and does nothing about the root issue: nobody decided where the transaction should start and end.

What an open transaction holds

You wouldn’t make a network call while holding a mutex, and an open transaction is worse. It holds a connection from a pool sized for short queries, it holds any row locks your writes have taken, and once it has written anything, it holds back vacuum’s horizon, so Postgres can’t remove dead rows until it ends.

The broker’s latency becomes your database’s lock hold time. With 20 connections and a provider that takes 500ms, you top out at 40 requests a second, and that’s with nothing else using the pool. When the provider has a bad day, the pool drains, requests queue on pool.Acquire, and pg_stat_activity fills with idle in transaction sessions.

Slowness is the mild outcome. If PlaceOrder succeeds and orders.Save fails, the rollback erases the debit, but the order still exists at the broker. Your database has no record of something that happened in the real world, and you’ll find out during reconciliation if you’re lucky.

Get in, do the writes that must succeed together, get out.

Context is the wrong place for an open resource

The best-known Go write-up of this design is Thibaut Rousseau’s SQL Transactions in Go: The Good Way, which grew into the transactor library. I strongly disagree with it. Rousseau argues transactions can be part of the business logic, and keeps database/sql out of the service layer by putting the open transaction in the context. Every store then fetches its connection through a getter that checks the context first. The article admits that’s sometimes controversial. I think it’s naive: it assumes every call site knows the transaction is there and treats it carefully, and the context gives them no way to know.

The context docs say values are for request-scoped data that transits processes and APIs. An open transaction is a resource with a lifecycle: it must be committed or rolled back, it can’t be used concurrently, and it’s dead after commit. Context can express none of that. A goroutine spawned inside RunTx inherits a pgx.Tx that isn’t safe for concurrent use. A context that outlives the commit carries a dead transaction. A repository that forgets to check, or a caller that passes the wrong context, runs on the pool and silently escapes the transaction. The compiler won’t catch any of it.

Keep the transaction in the repository

None of this is new. The Go docs’ own transaction example is one function that begins, defers the rollback, does its reads and writes, and commits. The same page warns that plain sql.DB calls made while a transaction is open run outside it, which is exactly the mistake a context-carried transaction hides.

A transactor is still the right tool, as long as the transaction goes to the closure as an argument. Go 1.27’s generic methods let RunTx hand the closure a query type bound to the transaction, and return whatever the closure returns:

type Transactor[Q any] struct {
	pool *pgxpool.Pool
	bind func(pgx.Tx) Q
}

func New[Q any](pool *pgxpool.Pool, bind func(pgx.Tx) Q) *Transactor[Q] {
	return &Transactor[Q]{pool: pool, bind: bind}
}

func (t *Transactor[Q]) RunTx[R any](ctx context.Context, fn func(Q) (R, error)) (R, error) {
	var zero R
	tx, err := t.pool.Begin(ctx)
	if err != nil {
		return zero, err
	}
	defer tx.Rollback(ctx)

	res, err := fn(t.bind(tx))
	if err != nil {
		return zero, err
	}
	if err := tx.Commit(ctx); err != nil {
		return zero, err
	}
	return res, nil
}

Each service wires it up once, with a small adapter that lives next to sqlc’s generated code:

// in the db package
func NewTx(tx pgx.Tx) *Queries { return New(tx) }

txr := transactor.New(pool, db.NewTx)

The adapter is there because db.New takes sqlc’s DBTX interface, and Go won’t convert a func(DBTX) to a func(pgx.Tx). On Go versions before 1.27, make RunTx a package-level generic function that takes the transactor as its first argument. pgx.Tx appears in the transactor and that adapter, and nowhere else. It never goes in the context, and nothing outside the repository can reach it.

Here’s a payment that needs to be idempotent and publish an event:

func (r *PaymentRepo) CreatePending(ctx context.Context, p Payment, idemKey string) (Payment, error) {
	return r.tx.RunTx(ctx, func(q *db.Queries) (Payment, error) {
		rec, err := q.ClaimIdempotencyKey(ctx, "payment", idemKey, p.ID)
		if err != nil {
			return Payment{}, err
		}
		if rec.Existing {
			return q.GetPayment(ctx, rec.ResourceID)
		}

		if err := q.InsertPayment(ctx, p.ID, p.AccountID, p.Amount, p.Status); err != nil {
			return Payment{}, err
		}
		if err := r.outbox.Write(ctx, q, p.ToEvent()); err != nil {
			return Payment{}, err
		}
		return p, nil
	})
}

The claim, the payment and the outbox event commit together or not at all. The outbox writer takes q as an argument, so it’s plainly inside the transaction, and nothing reaches it through the context. Nothing leaves the database until commit; a relay delivers the event after it.

A concurrent request with the same key waits on the claim, then gets the committed payment back. Because the transaction is short, that wait is milliseconds, not the length of the broker call.

A failed transaction leaves nothing behind, not even the key, so the client can retry with the same key. That only works because the transaction is short and has no side effects. A request-scoped one can’t be safely retried, because the handler may already have called the broker.

If a method needs a row lock, SELECT ... FOR UPDATE inside it is fine. The lock lasts milliseconds and never outlives the method. The trouble is a lock whose hold time depends on whoever holds the context.

External calls

When you need another system’s answer before responding, it’s tempting to keep the transaction open around the call. Don’t. Store pending, make the call, store the result:

payment, err := s.payments.CreatePending(ctx, NewPayment(req.AccountID, req.Amount), req.IdempotencyKey)
if err != nil {
	return err
}
if payment.Status != StatusPending {
	return nil
}
res, err := s.rail.Submit(ctx, payment.ID, req.IdempotencyKey)
return s.payments.RecordResult(ctx, payment.ID, res, err)

Two short transactions, and no connection held during the network call. If the process dies after Submit, the payment stays pending and a reconciler asks the provider what happened.

“Just use it carefully”

The pushback is always the same: why are you making this so complicated? YAGNI. You need long transactions for consistency when you’re working with money, so just be careful.

But that’s how you end up with transactions opened in the request handler and passed around in the context, transactions inside transactions, and the database acting as a distributed mutex for the whole system.

And “be careful” isn’t a plan. It’s a Sev 1 waiting to happen. We need to design systems so the wrong thing is hard to do, and passing an open transaction around in ctx is the exact opposite. It isn’t obvious to the new hire that they’re in an open transaction. They add a quick call to another service to check something, it gets through code review, and the next time that service is slow, your pool is empty.

Putting an open connection in the context hands every developer a noose and tells them to have fun. And since the pool is shared, it’s everyone’s neck.

If an invariant really spans two repositories, that’s usually an aggregate nobody has named. Write the method that owns it. Three Dots Labs make the same case in Database Transactions in Go with Layered Architecture: transactions in the logic layer are an anti-pattern, and moving them into the context only hides it.

A transaction is a critical section. Treat it like a mutex: never hold it across a network call, and know exactly who owns it.

If your transactor stores the transaction in ctx, change RunTx so the closure takes the queries as an argument and let the compiler show you every place that was silently relying on the context. Then look at what runs inside each closure. I’d bet at least one of them makes a network call.