n8n · Zapier · Make — repair work
An order-intake automation that lost one order in eight — and said nothing
A small ceramics studio ran orders from their shop into Google Sheets, a confirmation email and a Slack channel. It worked for months, then started dropping orders and sending some customers two confirmations. Nobody noticed until a customer asked where their mug was.
This page is the repair: what was broken, why, and what the fix changed — with both versions of the workflow replayed over the same shift so the difference is counted, not claimed.
Self-initiated demo. Palewood Ceramics is an invented company; the orders, customers and outages are generated by a scenario file that ships with the project. The workflow JSON, the defects and the repairs are the kind of defects and repairs this service covers.
What the owner reported
Four complaints, in her words. Every one of them turned out to have a different cause.
“Some orders just never show up in the sheet.”
No pattern she could see — a few every day, always discovered from the customer’s side.
“A few customers got the same confirmation twice.”
Embarrassing, and two of them thought they had been charged twice.
“Sometimes the Slack channel goes quiet for an hour.”
The automation had stopped mid-run. n8n knew; nobody else did.
“I don’t trust it any more, so I check every order by hand.”
Which is the real cost: the automation was still running, but it had stopped saving time.
Replay the same shift through both versions
One five-hour shift: 120 orders, 60 scheduled runs, and the same interruptions in both runs — a seven-minute shop-host restart, one rate limit and one bad gateway from the shop API, two Gmail quota blocks, three orders the shop published seven to nine minutes late, four customers who corrected their order a few minutes after placing it, one order without a surname, one without a phone number, one with an empty email field, and one address that bounces permanently. The two workflow files are the only difference.
Before — “Order intake v3”
order-intake-BROKEN.jsonAfter — “Order intake v4”
order-intake-FIXED.jsonWhat was broken, and why
Six defects. The “before” and “after” snippets are read out of the two workflow files on this page — not retyped — so they cannot drift from what actually runs.
Run log
Every scheduled run of the shift, as n8n would list its executions. In the broken version the failures are the rows in red — and none of them produced a message to anyone.
| Run | Time (UTC) | Status | Saved | Parked | Emails | Alerts | What happened |
|---|---|---|---|---|---|---|---|
| Run the comparison above to fill the log. | |||||||
How this page runs the workflows
So that nothing here has to be taken on trust.
- The two JSON files are written in n8n’s export format — node types, typeVersions, resource locators, retry settings and error outputs — and are meant to be loaded with Workflows → Import from file. To be exact about it: they were written for this demo and have not been opened inside a running n8n instance, so treat the canvas as untested and the logic as tested. The nodes, the expressions and the JavaScript inside the Code nodes are what you see in the snippets above.
- The page executes those files. A small runner in
engine/runner.jswalks the same nodes and connections, evaluates the same{{ … }}expressions and actually runs the JavaScript stored in the Code nodes. The shop API, Google Sheets, Gmail and Slack are local mocks with scripted outages. - Both versions get the identical shift. Same orders, same timestamps, same outages, same virtual clock — so retrying two seconds apart cannot escape a seven-minute outage.
- It is not n8n itself. The runner covers the nine node types these workflows use — Schedule Trigger, Error Trigger, HTTP Request, Code, Split Out, IF, Google Sheets, Gmail and Slack — and does not reproduce n8n’s queue mode, concurrency or its UI.
- One n8n quirk to know before testing it there: the cursor and the memory of handled orders live in the workflow’s static data, which n8n keeps for production executions of an active workflow, not for manual runs from the editor. Pressing “Execute workflow” by hand will therefore look as if the cursor never moves.
- One difference worth naming: the runner retries the item that failed. n8n retries the whole node, which can repeat items that already succeeded. For a node with a side effect — sending email — the production build therefore loops over orders one at a time, so a retry cannot email anyone twice. That loop is the next step recommended in the written repair report (available on request), not something this replay pretends to prove.
- The numbers in the panels come from the run you just triggered, not from
text written into the page. In the project,
node tests/run-tests.jschecks the same figures from the command line: 38 assertions over the files and both runs.
What the monthly part covers
- Watching the runs. Failed executions and error alerts land in a channel I read, so a broken automation is found by me and not by a customer.
- Keeping up with the services. Shop, Google and Slack change their APIs and their limits; an automation nobody maintains fails quietly in the middle of a busy week.
- Small changes as the business changes. A new product type, a second sheet, an extra notification — included, within the agreed monthly hours.
- A short monthly note: how many orders went through, what failed and what I changed. One page, no jargon.