Layers, daemons, services - one boundary drawn at three depths, and each depth has its own bill.
👋 Hi, I'm Anton - a software engineer working mostly in PHP/Symfony and Go, currently carving a live PHP monolith into Go services. This block is about what a service is to everyone else: whose it is, who calls it, what it promises. Part 5 listed the four lines of a storage requirement; this last part picks one of them - the ratio of reads to writes - and follows it as far down as it goes. Running notes are on my GitHub: github.com/brilliant-almazov.
This is how I do it right now, with the price attached - maybe you already do it better, maybe you see it differently.
Where the boundary gets drawn
Short primer, so this reads without the earlier parts of the block.
Each of our Go services raises itself from one declaration - a manifest listing its daemons, its databases, its queues, its schedules. The runtime and the deploy pipeline read that file; nothing about a service is true because a page somewhere says so. One line of the requirement feeding that file is the ratio of reads to writes, and it is written before the first table exists.
What that line buys is not only a schema decision. It decides which executor runs a read and which one runs a write - and "executor" turns out to mean three different things, depending on how far down you take it.
The thesis
Splitting reads from writes is not a decision about whether. It is a decision about how deep.
The same boundary can sit between two packages inside one process, between two daemons built from one image, or between two services with a stream of events between them. All three are the same decision taken at a different depth. Each step down buys something specific, and each step down is paid for in consistency, in latency and in the number of moving parts.
The case: 44 copies of four read shapes
Before any of that is worth discussing, the first depth has to actually be one thing rather than forty-four. Here is the audit that started this, run on the service in question in August:
multi-row loop .............. 18 copies
single-row read .............. 6
COUNT(*) .............. 8
existence probe .............. 12
────────────────────────────────────
total .............. 44 of which 12 differ only
in the text of the error wrapper
Four shapes of reading rows out of a database, written by hand in every domain that needed them. Nothing in that list was wrong on its own. Every copy compiled, every copy was tested, every copy did what its caller wanted.
The problem is what the four shapes differed by: the result type and the dependency, and nothing else. That is the definition of a generic - written out with copy-paste instead of a type parameter. And the correct form already existed in the tree: a reader carrying an executor, a statement and a scanner, with one Query(ctx, in) method. It was sitting inside a domain package, where no other domain could reach it, so everyone else wrote their own.
How it got to 44 is the least interesting part and the most familiar one. Each copy was written by somebody who needed a list of rows that afternoon, could see three similar loops nearby, and had no reason to believe the fourth one was a problem. Nobody made a bad call. The count is what makes it visible: one loop is a loop, eighteen loops are a missing abstraction, and nothing in between produces a moment where anyone notices.
The fix was not clever. The reader moved into one cross-cutting package, the 44 copies became calls into it, and the question was closed by a forbidding test: for rows.Next(), rows.Err(), rows.Close() and QueryRow( occur in exactly one package of the tree. Not "should occur" - the test fails the build if they occur anywhere else. A convention that is only written down decays; a convention with a failing build behind it does not.
That is the connection to everything below. The first depth is only worth having when the read side is one primitive. A read layer that is forty-four hand-written variants is not a layer; it is a naming convention. Splitting it from the write side buys nothing, because there is nothing coherent on either side of the line.
How this is usually done
Fairly, because all three of these are reasonable and I have used the first two.
The common framing is command-query separation inside one application: the write path and the read path get their own models and their own objects, in one process and one deployment - the shape described in Fowler's note on CQRS. One level up, the read side becomes its own model, kept current by events emitted from the write side; that is the version in the Azure architecture centre's CQRS pattern, usually paired with a materialised view. And one level up again, the read model becomes a separate deployable, with its own scaling and its own store.
Those are three points on the same axis, not three competing schools. What follows is where I place each one and what it costs me.
Depth 1 - layers inside one process
Two layers touch the database, and only two. repository reads and only reads. manager writes and only writes, find-or-create included, because that is a write with a read in front of it.
There is no store, no writer, no dao, no service in between. Those names exist to avoid deciding which of the two things a type does, and every one of them ends up doing both. When a manager needs to read, it calls a repository; it does not carry its own copy of the SQL.
The package layout follows the component type, with the domain nested inside it:
internal/repository/order/methods/ one method = one file = one type with Execute
internal/manager/order/methods/
internal/repository/account/methods/
internal/manager/account/methods/
Two details make the split hold rather than erode:
-
One method is one file and one type with an
Executemethod. A package that reads eleven things has eleven files, not one file with eleven functions on a struct that slowly grows a write path. -
The transaction is a decorator on the outside, and only over mutations.
tx pgx.Txnever appears in a signature; the executor comes from the context. A domain method does not know whether it is running inside a transaction, which is exactly why the read side can never quietly acquire one.
Worth being precise about what this is not. Two layers is not two models. There is one representation of the data, one schema, one transaction boundary; what is split is which code is allowed to write and which code is allowed to read, not what the data looks like on each side. The distinction matters because the word "CQRS" gets attached to both, and only the deeper version carries the consistency bill.
The price of this depth is close to zero. It is a rule about where files live, not a piece of infrastructure. Nothing new runs, nothing new can fail, and there is no consistency question because it is all one process and one connection pool.
Depth 2 - two daemons of one image
One image, two binaries: server answers synchronous calls, worker drains the outgoing-event queue, runs background computation, and executes partition retention on the audit tables.
Both are declared in the manifest and deployed by the same pipeline with a parameter saying which of them to bring up. That declaration is the whole mechanism - there is no second repository, no second image, no separate release.
One rule about the composition of that list, and it is not a technical one: the set of daemons is decided by the owner of the service, and by nobody else. New background work goes into an existing daemon. A new cmd/* is not something an implementer introduces on the way past, because a binary that nobody declared is a binary that nothing deploys.
What this depth buys is that a long job stops sharing a request budget with a synchronous handler. What it costs:
- A second process, with its own resource configuration, its own restarts and its own place in every dashboard. Two processes is where "the service is up" stops being a single fact.
-
One publisher per gauge. Queue depth and partition age are published by
workeronly, never byserver. Two processes publishing one gauge produce a graph that flips between their two versions of the truth, which is worse than having no graph, because it looks like data.
Depth 3 - separate services
The third step is a write-side service and a stream of events, with the read side living somewhere else entirely, on its own store and its own schedule.
We do not have this in production. What we have is depth 1 and depth 2. I am describing the third because the mechanism is worth writing down and because its price is what keeps the decision honest - not because there is an installation behind it.
Everything that was free at the first two depths becomes an explicit obligation here:
- Delivery is at-least-once. The broker does not deduplicate. Idempotency by event identifier is the consumer's job, and it is not optional - it is the thing that makes the stream usable at all.
-
The relay is the single drainer. Events leave in strict order by
id, andFOR UPDATE SKIP LOCKEDis deliberately not used. Parallel draining is faster and it breaks order, and order is precisely what lets a consumer rebuild a correct replica. -
Redelivery is normal, a spike of it is the incident. If the process dies between publishing a message and deleting its row, the same
message_idgoes out again on the next pass. That is designed behaviour, not a fault. The alert is on the rate, not on the event.
Read that list again as a cost sheet rather than a design. Three obligations, all of them permanent, all of them on code that did not previously need to exist. And note where they land: not on the service that made the split, but on every consumer downstream of it. That is the part that makes this depth expensive in a way the first two are not - you are spending somebody else's budget as well as your own.
The depth is chosen by a line of the requirement
This is the part I would defend hardest, and it is not about reads and writes specifically.
An argument about "CQRS or not" is settled by a line of a non-functional requirement, not by preference. If no such line exists, there is nothing to argue about - the requirement has not been gathered yet, and two people are comparing instincts.
For this particular decision the mapping is short enough to fit in four lines:
read/write ratio near 1, small load → depth 1, and stop
different load, different shape → depth 1, from the first table
heavy background work beside handlers → depth 2
different owner, different response target → depth 3, with the bill below
The read-to-write ratio is one of the four lines of the storage requirement from Part 5. That is where this decision is made - months before anybody opens an editor to make it.
The one conclusion
Every step deeper is paid for in three currencies: consistency, latency and the number of moving parts. Until a line of the requirement shows there is something to pay with, the decision stays at the depth above.
That is the whole rule. It is deliberately not "start at depth 1 and grow into depth 3" - growth is not the trigger, a named requirement is.
What each depth costs
| depth | what it buys | what it costs |
|---|---|---|
| layers in one process | a read is written once, not 44 times | moving between two package trees to read one feature |
| daemons of one image | background work stops sharing a request budget | a second process, a second configuration, one gauge owner |
| separate services | read and write scale and fail separately | eventual consistency, idempotent consumers, ordering by design |
The third row is the one worth staring at, and it has one more entry than the cell holds: debugging a problem across two systems instead of one. Nothing in that row is a one-off cost. Every line of it is a permanent obligation on everyone who touches that code afterwards.
What a skipped depth looks like later
The failure mode never announces itself as "we picked the wrong depth". It arrives wearing something else.
A background recalculation that shares a process with request handlers is not reported as a topology problem. It is reported as latency, intermittently, on endpoints that have nothing to do with it - and it gets three separate investigations before anyone looks at what else is running in that process. That is a missing depth 2.
The opposite direction is quieter and more expensive. A team goes to depth 3 for a domain whose read and write have the same owner and the same response-time target. Nothing breaks. There is simply a stream, a relay, an idempotency key and a second store to reason about forever, in exchange for a separation that nobody ever needed. Nothing about that shows up as an incident - it shows up as every subsequent change taking longer than it should.
When not to go deeper
When the ratio is close to one and the load is small. Depth 1 is the whole answer. Going further adds parts and removes no problem, and "we might need it later" is not a requirement line.
When there is no separate response-time target for reads. Without that, depth 3 buys nothing that depth 2 did not already buy, and it charges you eventual consistency for the privilege.
And the honest boundary, stated once more because it is easy to read this piece as a report rather than an analysis: one domain split across separate read and write services does not exist in my production. Layers exist. Daemons exist. The third depth is written up here as a mechanism and a price, and that is all it is.
The multiplier line
Collapsing 44 copies into one reader is exactly the kind of work an automated executor does well - mechanical, well-specified, verifiable by a test that either passes or does not. What it did not do was notice that four shapes differing only by result type and dependency were a generic written by hand, or decide that the right answer was one cross-cutting primitive rather than a tidier copy in each domain. That judgement is the whole exercise, and speed applied before it is made produces forty-four beautifully refactored copies. Speed amplifies whoever set the boundary; it does not draw it.
Services and commitments - Part 6. That closes this block: a service exists when it is declared, its owner is a field, its call map is observed rather than claimed, its promises are metrics with targets on them, its volume is a requirement line, and that same line decides how deep the read/write boundary is cut.
If you do this better, tell me what made you go one depth further than you had to. If you have been through this, what did the extra depth cost you that nobody had costed up front? If you see it differently, say where separating reads and writes at the service level pays off earlier than I claim. How is it solved on your side, and what broke there?




Top comments (0)