A couple of years back, I ran a series of posts which tried to explain what it meant to achieve mastery in software engineering, including a tongue-in-cheek answer to the question of what it meant to work on hard problems. But looking back, none of these really answered the question of what kinds of hard problems experienced engineers face, and how this shifts their perspective. This is the first of what will hopefully be a series of posts that fill this gap.
Exactly Once
One of the most common “hard” system design problems is making sure that something happens exactly one time. The credit card must be charged exactly once, the payment disbursed exactly one time, the merchandise shipped once and only once. Please believe me when I tell you that it really sucks to try to convince a customer to return merchandise that you accidentally double-shipped.
In order to solve this problem, and, I would argue, work on any sufficiently complex system, you need to understand the concept of idempotence.
Idempotence is the idea that no matter how many times a function is run, the output and side effects will be the same. The following commands are idempotent:
x = 1
UPDATE my_table SET my_column = 5;
echo "foo" > my_file.txt
No matter how many times you run the above commands, they will have the same result. x will always equal 1. my_column will always equal 5. my_file.txt will always have the same contents (though the creation date will be different). Run them once or a thousand times, it makes no difference.
By contrast, the following commands are not idempotent:
x += 1
UPDATE my_table SET my_column = my_column + 1;
echo "foo" >> my_file.txt
Running the above commands twice will give you different results from running them once. Running them three times will give you a different result from twice, and so on.
The reason this is important is that idempotence is a requirement for “exactly once systems”. If the call to the credit card processor fails, but you can’t safely re-run the “charge credit card” function again without the danger of double-charging, then your design is bad and you’re going to get in trouble.
But why, you may ask, is this such a big deal? Just be careful when writing the code, create an appropriate set of unit tests, QA all the edge cases, and everything should work out… Right?
LOL. No.
The unfortunate reality of working on real-world systems (as opposed to academic exercises) is that you have to assume that your code can fail on any line, at any time, for any reason or no reason. Maybe you have a bug. Or maybe your code is great, but a library upgrade subtly changes the API contract, and the code starts failing under some complex set of circumstances. Maybe the credit card processor is down. Or someone unplugs the server between when you make the http call and when you get a response. Or AWS chooses that moment to terminate the process. Or the server suddenly and catastrophically crashes for absolutely no reason.1 Shit happens. With this in mind, senior engineers understand that you have to design your system with the expectation that it could fail between any two lines of code, and that it needs to be resilient where it counts.
A Simple Example
Consider the following simplified e-commerce flow:
- Reserve inventory
- Charge credit card
- Create shipping label2
- Ship merchandise
Conceptually, you can imagine that there’s a function in every e-commerce platform that performs these steps in order. If you reserve the inventory and charge the credit card, but then the process dies due to a memory leak or server crash, then you need to call the function again and skip the first two steps. The simplest way (conceptually) to think about this is to redefine the above list like this:
- Orchestration step: Skip to the appropriate step, based on recorded status
- Reserve inventory step:
- Reserve inventory
- Update status
- Charge credit card step:
- Charge the credit card
- Update status
- Create shipping label step:
- Create the shipping label
- Update status
- Ship merchandise step:
- Ship the merchandise
- Update status
OK, but what if your machine crashes between charging the credit card (3.I) and updating status (3.II)? If a new process picks up the job, it will just go through and charge the credit card again. So we have to be careful: first recording that we’ve started a step, then trying to complete it, then recording that we finished.
- Orchestration step: Skip to the appropriate step, based on recorded status
- Reserve inventory step:
- Check status to see if we’ve already reserved inventory. If not:
- Record that we’re going to reserve inventory
- Reserve inventory
- Update status
- Charge credit card step:
- Check status to see if we’ve already charged the credit card. If not:
- Record that we’re going to charge the credit card
- Charge the credit card
- Update status
- Create shipping label step:
- Check status to see if we’ve already created the shipping label. If not:
- Record that we’re going to create the shipping label
- Create the shipping label
- Update status
- Ship merchandise step:
- Check status to see if we’ve already shipped the merchandise. If not:
- Record that we’re going to ship the merchandise
- Ship the merchandise
- Update status
Now, we have an indication of intra-step status: if we get to a step where we see that we already tried to execute the action, but didn’t record completion, then we know something’s wrong. We’ll get to that in a little bit.
An Aside: Orchestration
In the above example I’ve sidestepped the questions of how to manage the list of tasks, and how to manage transition between steps in a single task. In modern systems one or both of these are commonly handled by message queues and emitted events, but to keep this post short I’m going to defer that conversation to a future post.
Different Challenges
Back to idempotency. Multi-step processes can typically be broken down into steps which can be handled independently. And while idempotency might be a requirement for the full process, this can be accomplished by focusing on the idempotency of the individual steps.
Following are a couple examples of how you might handle different types of problems.
Databases
Idempotency with databases is usually easy, since 1) you can either use transactions or do clever things with upserts, sub-queries, etc., and 2) completing the action and recording the action are frequently the same thing. In the e-commerce example, you don’t actually need to run a step to check if inventory has been reserved, or to record it after the fact. Creating a row to reserve the inventory can be run idempotently (assuming a unique index on transaction_id, merchandise_id), and the row serves as a record that the action has been completed.3
INSERT INTO reservations
(transaction_id, merchandise_id, num_units)
VALUES ($1, $2)
ON CONFLICT (transaction_id, merchandise_id)
DO NOTHING
Files
It’s common to pull data down from an external source and persist it in cloud-based storage that you control. The problem is that you can’t just use the file’s existence as proof that the action completed correctly—you could end up with a job that failed halfway through and left a file that’s empty, incomplete, or corrupted. It’s also possible that “no data” is a reasonable result, in which case the absence of a file wouldn’t indicate job status.
In this case, you might do something like this:
- Check the status of the step
- If “unstarted” (or null), then go to step 2
- If “started”, then you’re in an unknown state – don’t worry about it, just go to step 3
- If “error” then terminate the process
- if “done” then move to the next step
- Set the status of the step to “started”
- Retrieve the file
- Save the file
- Update the status to “done”
We expect that saving a file with a deterministic filename should be idempotent. If we write it multiple times, it will just overwrite any previous copies. The timestamp will be different, but the content of the file should be the same.
Calls to External Services
There are many ways that calls to external systems might fail. For instance, when calling a payment processor, your process could fail before the API call returns (did it succeed? who knows?). Or the request could time out—but does that mean that the payment failed to go through? Or that the charge was successfully made and that something prevented the API from sending a result? Or stopped your code from receiving it? You have to think through all possibilities, and you might end up with something like this:
- Check the status of the step
- If “unstarted” (or null), then go to step 2
- If “started”, then you’re in an unknown state – make a call to the payment processor to see if it has a related payment (important to make sure that the search criteria are carefully defined—you could have someone who bought the same product twice in rapid succession).
- If there is a payment matching the criteria, then go to step 4
- If there isn’t, then go to step 3
- If there was an error, then update the status to “error” and put the transaction in a queue for special handling
- If “error” then terminate the process
- if “done” then move to the next step
- Set the status of the step to “started”
- Send a request to the payment processor to execute the payment
- Update the status to “done”
The challenge is when we know that we started previously, but that for some reason the process was interrupted. In this case, we need to be able to find some evidence of what happened. If the failure happened while we were querying an external service (e.g., the server suffered a catastrophic failure after we sent a request to a credit card processor), then we can’t know what happened without some kind of follow-up investigation.
Maybe we can send a request to the credit card processor with a unique ID, and get a record of the transaction (or proof that it didn’t happen). Or perhaps we can configure a web hook to receive confirmation of each transaction, or set up an hourly job that pulls down all completed transactions. The key thing is to think through what it will look like in the case of success or failure, and to understand how you can determine true status.
What it all means
This may seem dense, tedious, and meticulous—the opposite of the rough and tumble world of software engineering you thought you were signing up for. It doubtless seems ridiculous that an unproblematic line of code should be followed by another unproblematic line of code (with error handling, and a finally clause!), and yet somehow that the program could fail catastrophically between one and the other. I agree. It does seem ridiculous, and deeply unfair, and yet it will happen. You must think deeply about how your code might will fail, minimize dependent steps, and think through how to handle failures at critical junctions.
If a failed data import can be trivially fixed by re-running the process, then you’ve just written idempotent code. If, on the other hand, restarting a failed checkout flow could charge the customer’s credit card twice, then you’ve created a dangerous situation that will bite you. You might be able to get away with a hack for a small scale prototype, but it won’t work for a production system.
Worrying about idempotence might feel like a distraction from your “real job” of writing new features, but this is junior thinking. Idempotence is a discipline you need to build, a muscle that you need to constantly exercise. In many cases, the feature isn’t “do X”—it’s “provably do X exactly once.” Building the habit of evaluating designs for idempotence—identifying whether they need it, and how to achieve it—is a critical requirement for anyone who wants to grow in their craft and their career.
Updated to add: Final Word
As mentioned in the comments, there’s an additional problem you need to worry about: making sure that two different jobs don’t pick up the same task simultaneously. This issue will be addressed in the next post.
- Please believe me when I tell you that this happens regularly. ↩︎
- Creating shipping labels costs money (they’re like stamps in e-commerce), so you don’t want to accidentally create extras. ↩︎
- Yes, of course it’s more complicated than this—you first have to check that the inventory can be reserved! And you’re probably using an inventory management system that you connect to via API, instead of just recording things in a database. Yes, yes, you’re very clever. ↩︎
Note: All em-dashes were artisanally created using alt-shift-hyphen. No part of this post was written using AI.
Great post! I fixed an idempotency issue like this some time ago. We had a workflow where one of step would read a number, increment it and write it back to a dyanmodb table. When that step occasionally fails, the workflow would stuck forever, manual intervention was required which was a pain for oncall. The fix was essentially take the read operation out and put it in the previous step, increment it and pass it onto the next step to write it.
Also, in your last example, you must ensure that a single process is retrying this charge request, otherwise you can have a race condition and end up charging twice.
These things are hard enough that they’ve kept CS researchers busy for a few decades. I concur that developers need to be aware of them!
But by now we should have systems (libraries, frameworks, services) that abstract most of it under a clear contract (“if your system using my system maintains these invariants, then my system ensures good things happen”).
I never worked in payment processing, maybe there are good tools in that domain? If not, my guess is that they could be assembled using XA transactions (or some other form of 2PC), probably with a consensus-based (Paxos, Raft, etc.) transaction manager for availability.
I’ve found that learning about concurrency and distributed systems (with TLA+ and from Lamport’s course) helped immensely, and so did reading about immutability (see “Immutability Changes Everything”). But you just can’t reinvent everything, much like you don’t recreate your OS or compiler: we need abstractions we can safely build on.
Lol, I had an entire section on thread safety, but decided the post was already a bit long. Great catch!