Trustworthy Online Controlled Experiments: Book Review

This book is a rare jewel that I would recommend to anyone, running A/B tests. Pretty happy with the purchase and my time spent with it.

After I figured out that something with my understanding of how experiments should run was off, my instinct was that it’s math and stats skills that were lacking. So I looked into improvements in the area of probability and statistics first and my last two self-improvement books where in that area. Trustworthy Online Controlled Experiments isn’t about math. It contains condensed experience from people, wrangling experiments at Google, Microsoft, and LinkedIn, and then has some advanced chapters, intended to make the book complete.

The area that stuck with me the most was how to pick metrics for evaluating experiments and what to do with them.

For example, one of my most beloved revenue metrics (average revenue per user) is considered a health metric by the authors and their reasoning very solid. They argue any primary metric should be around user journey, usability, and satisfaction, around the purpose for the specific service. Revenue is an indicator how well the overall system performs but better indicators exist that are also easier to move with an experiment.

Many other metrics I currently frequently check, the book categorizes under drill-downs of existing metrics. The book also highlights the importance of having classical log data, like server errors, in the experiment dashboard. The authors list 20-ish industry standard metric ideas but also state that their internal systems have 1000s of metrics. What could these be? This statement also makes me wonder, how do they even make any calls, if they have so much data? Sounds like a good problem to have.

Regarding the number of metrics, the next book on the subject I picked, called Experimentation for Engineers, has a lengthy chapter about optimizing for just one metric. Then another lengthy chapter about optimizing for two metrics. Is that just a matter of taste? Being spoiled by the ideas of Trustworthy Online Controlled Experiments, I didn’t like it.

Imagine a page with two buttons, and we make one of the buttons orange in the treatment variation. Two clear metrics here can be clicks on the modified button and clicks on the non-modified button, together with the general metrics, available for all experiments. If the orange button gets more clicks, the other one will get fewer clicks. Such is life and decisions should be informed based on both changes. It’s also possible that the total number of clicks goes down because people perceive the page as spammy due to the orange color, despite orange clicks going up.

Optimizing for one thing without knowing the others can work but only in scenarios where there is really just one thing. For example, algorithmic trading of a specific stock. Optimizing for one metric while staying oblivious about the others and can produce errors and eventually erode the trust in experimentation.

So, thanks for reading this brain dump. It’s not about a cat, it’s not a quite a book review either because it covers just one area the book touches, but what can I do. The things that fascinate me are sometimes unusual.

A metric for an experimentation system’s maturity

Experiments in the context of software involve showing different groups of people different UIs or features and measuring the impact. This is one way to make informed decisions based on data rather than opinions. If experiments aren’t used, decision-making is usually subjective and prone to errors.

Humans are not good at predicting the future. We may be able to guess where the bus will stop next, but try guessing the price of Bitcoin next year. Being able to properly evaluate the outcome of decisions has real value, while relying on opinions can lead to irreversible losses.

Experiments are one way to get that data, and they’re something I’ve practiced a lot over the last two decades. I can say I have some experience, good or bad. But how far did I get with my own understanding of what a mature experimentation system should be?

The book I’m reading introduces a very interesting metric for experimentation maturity, based on the number of experiments, run per year. Metrics like this can certainly be criticized as superficial, it’s comparable to lines of code written or numbers of PRs. But having a metric is better than not having one, so here’s the maturity model from Trustworthy Online Controlled Experiments, Chapter 4. Numbers apply for large online services, like Google.

  1. Crawl phase: <10 experiments/year, few metrics, no internal system, manual post-experiment analysis
  2. Walk phase: up to 50/year, establishing standard metrics and some setup and data validation procedures, some internal systems and scorecards are available
  3. Run phase: up to 250/year, adoption at a level where most decisions are verified; the focus is on scaling, ease of use, increased number of standard metrics (100s to 1000s)
  4. Fly phase: 1,000s/year, a memory system is developed, the process is built around experimentation, shipping is easy, proper learnings and knowledge are preserved, metric drill downs (mentioned later in the book), thousands of metrics

The book isn’t exactly clear about when you’re out of Run and into Fly, given that one is around 250 experiments and the other is 1,000+, but it at least gives you a way to estimate roughly where you are if you’re running experiments.

It also doesn’t talk much about metric classification. As experimentation matures, metrics should be grouped into categories, with drill-downs for important characteristics. Some of the examples of what should be available felt very strong, like things I wish I knew sooner.

This book is great. I’m about 20 pages away from the end, and I have enough notes to write a chapter rather than a blog post. The last chapters are full of formulas and take time, as some of the formulas are not something I’ve encountered before. I learned stats with pen and paper, making countless calculations of averages, standard deviations, p-values, and such by hand. When I encounter a new stats formula, my brain wants to see the tables behind it, the charts behind it, the history behind it, and it refuses to move on by just accepting it.

Commentary on Brandolini’s Law

Brandolini defined the following law:

The amount of energy needed to refute bullshit is an order of magnitude bigger than that needed to produce it.

It normally applies to arguments and misinformation. Now that generating content with AI is so easy, the orders of magnitude have changed. Both the statement and the refutation can be forged very quickly and people believing stuff on the Internet just look more and more foolish.

So, while trying to figure out what to even think about this, I came up with the following commentary on Bardolini’s Law.

The probability that any new online content is AI, BS, or both increases over time.

With the growth of DC power, personal AI orchestrators, GEO, and SEO, this probability will eventually be trending towards 1. What would be the point to refute, or even read anything, if it’s one of a million clear BS pieces of information? So here’s my prediction:

The authenticity of content will become more important than its quality.

Rusty the cat

A story about the 33-year-old cat Rusty took Reddit by storm. Rusty passed away at the age of 33, making him into top 10 of the oldest cats in the recorder cat history. The poster submitted proof to the mods of r/cats that Rusty was real. People rushed to update Wikipedia’s list of oldest cats, expressed condolences, sent love and wishes. The post generated 134K likes and was probably viewed by most of the non-bot Redditors.

Unfortunately, Rusty was AI slop. Rusty:

The post and the bot that submitted it got deleted.

I think my attempt to restrict Reddit to a few essential subreddits like r/cats is not very successful and I still have exposure to something that I wouldn’t even call AI slop. More like farming for free human-generated text for the purpose of training LLMs. Rusty helps me understand why the sudden rush to gather personal IDs and verify humans on social media. All the social networks are vulnerable to slop and risk losing engagement if they don’t put it under some level of control.

Here’s a real orange cat for you, blissfully unaware about the decay of r/cats.

UPDATE: this is a male cat named Pesho, also known as The Son of a Mother. Likes cuddles and bites for no reason.