Writing / Hunter Harris
One Line, 25M Revenue
“I mean, it’s one bug, Michael. What could it cost, ten dollars?”
Do you have some white whale stories? You know, the unbelievable ones. The “I can’t believe that actually happened” ones. The “I was there for it, and I barely believe it”. Today I want to share one of those stories. I’ve been incredibly hesitant to share, in part because it’s so big. Well, big to me. On the scale of the economy and enterprise, it’s not small potatoes, but it is potatoes. But I live in Startup Land, where every dollar counts.
I also want to share this, because it was one of the catalyzing events of my career. I would not be here writing this blog if not for this event. I might not be running my business if not for this event. This event was the one where I went from “I think there’s an opportunity” to “I was incredibly pessimistic and naive about the available opportunity”.
Gearing Up
Let’s start, as many stories do, at a beginning. It wasn’t the beginning, but it will suffice. I joined a well-regarded restaurantTech, and as I prefer to do, I settled myself into an engineering team in need of assistance. In this case, it was Growth - responsible predominantly for things like coupons, promotional emails, and the like. I settled down, and looked for a good place to start. I was promptly told that a good entrypoint ticket was a simple card. Take these designs, restyle the marketing emails, easy peasy. 1 ticket, 1 point, you should be done by this time tomorrow. Normally, they’d be right. In this case, that estimate was wrong.
We’re not here to talk about the email designs, but you probably want some context. “Change colors in the email? That’s not even a 5 minute fix with AI. What’s the holdup?” And you’d be right. Unfortunately, the colors were not supported in the system. (This is where you buckle up, unless you want to skip to the next section).
To add colors, you can’t just add them to the email. They have to be added to a separate repo, managed by another team. No, they don’t use this repo often. No, it actually requires some amount of work to do that.
Oh, those color values are outside the design system for the dependency. We have to do a refactor on a different repo to support just these colors.
Now have to translate the designs to achieve design fidelity. That’ll take a couple weeks with Design’s backlog.
Oh, we also require different font requirements than expected, that’s further refactoring in the repo.
Oh, it turns out this repo is a shared dependency. It, too, haphazardly includes the shared dependency, and now requires a refactor to accommodate the colors and fonts.
Now we’ve got one email rendering correctly between email and the email builder! Done, right? No, next we have a design framework that must also be extended along these refactors, begetting a third refactor.
OK great, now the theme builder is integrated right? Well, congratulations, but no. You’ve done the template for one email. You now need to backfill “many” records. Also, you cannot just change the record in this backfill. You have to traverse a directed graph of emails and perform operations, live, a LARGE number of times.
I’ll pump the brakes here, after adding: this is not a comprehensive list of issues. Requirements changed over time. We were obligated to do refactors more than once, and these migrations more than once. Whew. Thanks for sitting through that.
Now that we’re on the other side, that 1 point 5 minutes card? We’ve involved no less than 3 teams, the Growth team is at about 80% capacity (Product, Design, 3 engineers), and we are on month 4. To slightly change how emails feel. Number of customers who have asked for this: 0. Associated revenue: 0. Output? We did a lot of it. Customer impact? I don’t think anyone noticed or cared, inside or outside the building. Total spend: At least 500k in salary. Classic Ship to Burn. Oh, by the way, we aren’t done yet.
Digging in
Now, I’m a human, and I can only bear to code so many hours per day. Support tickets came in, and I was personally invested in getting them solved. I noticed quickly that we had quite the backlog, so being the Product aligned individual I am, I burned that thing down. Total size, hundreds. Less than a thousand. So large, but not insurmountable.
Two trends emerged. There were hundreds of tickets that were just empty, no context, author had forgotten. Noise. Close, no fix. There were a number of trivial bugfixes with clear consequences + repro steps - just ship. There were a bunch of things though that were… interesting.
My approach for interesting ones is one I hope most people do. I took the tickets, tagged them, and essentially PRD’d scoped chunks of “here be dragons, there’s a bunch of issues here”. Doing this, you quickly learn things like “Wow, this has been reported 72 times over 6 months” or “this seems severe, but actually no one seems to care”. It also lets you know things like “Hey, our coupons UI is quite a bit better than we expected (or no one is using it)” or “Wow, parts of our email delivery have a literal menagerie of bugs. Like a whole bug circus. Like we need to invent new taxonomies of bugs, there are so many”.
This menagerie of bugs is the sort of thing I like poking into.
In Too Deep
So, this is where it’s probably useful to lay out some of the behaviours we saw. This is a company that helps restaurants market to their customers. One major way they do this is via coupons. We wanted to test a variety of different coupons, via email and text.
We noticed that our deliverability for coupons and other promotions was weird. In a lot of circumstances, a user would schedule a coupon for say 11:50am Tuesday, with a coupon like HOTLUNCH.
You would, naturally, expect this to go out to customers around 11:50am on Tuesday. You would also be wrong.
Frequently, it would go out somewhere between 1:33am and 2:13am on Saturday. It would then proceed to be forgotten about by the time it could be redeemed, on Tuesday. This was a constant source of support tickets, and backlog tickets. It’s been years, so I don’t remember all the specifics, but there was a varied set of behaviours around this - not reproducible outside production, intermittent, uneven. A classic tangled knot of tech debt.
But an interesting one. I got involved first, because it stunk - specifically, of opportunity. This was an area where lots of customers were feeling pain. There was opportunity to do some good for them and for the business here. So I started with something I was familiar with - open hours. In this system, open hours were calculated by stacking availability/calendars, and calendars was something I was imminently familiar with from my time at Calendly. It turns out that we were combining opening hours wrong algorithmically, and changing the equation meant that the problematic sending behaviour shaped up into the wrong time, but the right day.
It was clear this was fixable, but that there were layers of bug investigation and archeology to be done.
Light At the End of The Tunnel
Now, we’re getting to the meat of the story. And honestly? This is maybe the most boring part of the story. First, I fixed a bug, then I fixed another bug, then I fixed another one. Behaviour got better by better, in small increments. Until one day.
On this particular day, I had just had to revert some code. I had copied a pattern from elsewhere in the code, with how we use SQS, and it didn’t work as I expected in production. It turned out that for the test suite, we had a SQS mock that behaved differently than how AWS actually behaved. I had to go to the documentation, and run through a bunch of parameters and check the contract.
This is when I noticed something extremely interesting. We consistantly misused a particular parameter tied to SQS (Simple Queue System)’s method for queueing. This meant a very interesting thing: We were sending out emails sequentially by tenant. Normally, you want to blast out email (not like, blast-blast, but sending 100k emails does take time).
We were not.
We were sending emails sequentially. One at a time. Load email, send email, wait, repeat. Exacerbating this, we wanted to send emails at a given time, so after time passed (I believe an hour) the send would terminate (without errors, natch, so no visibility).
I want to dwell on the ramifications of this for a moment. Which I know may spoil the surprise at the end, but you already know the number. This meant that if I was a shack with a 1000 person customer list, I might comfortably send out my entire email list. This also meant that is I was a multi-location chain, with 10k+ emails at each of my 100 locations, I might send out an average of 10 per location (or possibly worse, I might send 1000 for one location and leave the others high and dry).
The fix for this was “one line of code” - actually less than that, it was on the order of “#{key}” if I recall, splitting keys into discrete email sends. We shipped this change, and thankfully we had great visibility into order performance.
The next day, we had devops come to us sweating. “The machines are under load but they’re holding”. In another world, this would have been a SEV0 outage, but our infrastructure was just barely scaled to the new load we were putting on the system. But the system was creaking and groaning under the load. Honestly, we thought maybe we had broken something.
Then we checked the coupon redemption numbers.
Friends, I want you to understand something. This company’s job is to do marketing for restaurants. That means, sending emails to people. Generally with coupons. They do a lot more, but that’s an important thing they do. What you may not fully get is that people respond to incentives. When you send someone a coupon at the right time, when they’re hungry, they order food and use the coupon. It WORKS. And when you send coupons to a lots of people, a lot of people buy food with those coupons. And if you make money by having people use those coupons, you stand to make a metric ton of money (shocking, I know).
So we immediately saw heavy up-and-to-the-right motion on every graph possible. It turns out that for the vast majority of our customers, we were sending the rough equivalent of a rounding error of their total email send - well below 1%. We fixed that, in spades. People converted.
By the end of the first month, we could attribute a difference of about 2.2M in additional revenue for the company month over month to the change to the queue behaviour (other fixes obviously contributed, but that was the catalyzing event) - and it sustained, and increased as the year progressed. 2.2M+ * 12 is a little over 25M per year. And this isn’t all sunshine and rainbows - we git blames around these bugs, and discovered that they had been festering in the system for a little over 4 years.
I’d like to mention, the aforementioned color change? It was still in flight.
So where does that leave us?
Well, I’d like to highlight that the color change was classic ship to burn. Non revenue tied objective, sunk cost fallacy. We should have cut tackle on it day 2, but every day it felt like “just one more day and it’s done”. I’d also like to mention, modern AI agents were available and did not help here. This was a large, sprawling, complex system and AI does not perform well in systems like this.
Second, the bugfix was ship to learn -> ship to earn. We didn’t formally go in with hypotheses (we should have), but I strongly suspected that there was something here. I will also add - AI tools existed, but were blind to this category of mistake. In the docs and in the code, we lied to the agents. And those lies (really misunderstandings) we told AI made it blind, and in turn we blinded ourselves. We were regularly traversing this area looking for bugs with AI and it never found this one.
One of my takeaways from this was that technical debt can have a monumental impact on the business if you compromise in the wrong place. This was a build vs buy decision gone wrong. This company could easily have adopted a third party service, gotten blocked out of a few cool features, and eliminated every part of this story. No crazy story, tighter backlog, more in the bank.
I’d also like you to take away how Product misalignment can sink you. This was a team pouring effort into something that did not matter while a separate massive issue festered. A single person had to go, against the organizational directives, and go spelunking in their free time to uncover this. By default, this issue would have persisted for a few years more. It is critical to get crisp with your priorities.
I’d also like to add a disclaimer. Engineering is a team sport. I did not do this work in a silo or a vacuum. I had help from others, and I helped others in executing on it. It is extremely hard to have this level of impact in a vacuum.
So anyways, that was a big thing that happened in my life, and now I want to do that for basically everyone. I think you have a pot of gold buried in your backyard. Go do a little spelunking for once.
