A metric for an experimentation system’s maturity

Experiments in the context of software involve showing different groups of people different UIs or features and measuring the impact. This is one way to make informed decisions based on data rather than opinions. If experiments aren’t used, decision-making is usually subjective and prone to errors.

Humans are not good at predicting the future. We may be able to guess where the bus will stop next, but try guessing the price of Bitcoin next year. Being able to properly evaluate the outcome of decisions has real value, while relying on opinions can lead to irreversible losses.

Experiments are one way to get that data, and they’re something I’ve practiced a lot over the last two decades. I can say I have some experience, good or bad. But how far did I get with my own understanding of what a mature experimentation system should be?

The book I’m reading introduces a very interesting metric for experimentation maturity, based on the number of experiments, run per year. Metrics like this can certainly be criticized as superficial, it’s comparable to lines of code written or numbers of PRs. But having a metric is better than not having one, so here’s the maturity model from Trustworthy Online Controlled Experiments, Chapter 4. Numbers apply for large online services, like Google.

  1. Crawl phase: <10 experiments/year, few metrics, no internal system, manual post-experiment analysis
  2. Walk phase: up to 50/year, establishing standard metrics and some setup and data validation procedures, some internal systems and scorecards are available
  3. Run phase: up to 250/year, adoption at a level where most decisions are verified; the focus is on scaling, ease of use, increased number of standard metrics (100s to 1000s)
  4. Fly phase: 1,000s/year, a memory system is developed, the process is built around experimentation, shipping is easy, proper learnings and knowledge are preserved, metric drill downs (mentioned later in the book), thousands of metrics

The book isn’t exactly clear about when you’re out of Run and into Fly, given that one is around 250 experiments and the other is 1,000+, but it at least gives you a way to estimate roughly where you are if you’re running experiments.

It also doesn’t talk much about metric classification. As experimentation matures, metrics should be grouped into categories, with drill-downs for important characteristics. Some of the examples of what should be available felt very strong, like things I wish I knew sooner.

This book is great. I’m about 20 pages away from the end, and I have enough notes to write a chapter rather than a blog post. The last chapters are full of formulas and take time, as some of the formulas are not something I’ve encountered before. I learned stats with pen and paper, making countless calculations of averages, standard deviations, p-values, and such by hand. When I encounter a new stats formula, my brain wants to see the tables behind it, the charts behind it, the history behind it, and it refuses to move on by just accepting it.

Commentary on Brandolini’s Law

Brandolini defined the following law:

The amount of energy needed to refute bullshit is an order of magnitude bigger than that needed to produce it.

It normally applies to arguments and misinformation. Now that generating content with AI is so easy, the orders of magnitude have changed. Both the statement and the refutation can be forged very quickly and people believing stuff on the Internet just look more and more foolish.

So, while trying to figure out what to even think about this, I came up with the following commentary on Bardolini’s Law.

The probability that any new online content is AI, BS, or both increases over time.

With the growth of DC power, personal AI orchestrators, GEO, and SEO, this probability will eventually be trending towards 1. What would be the point to refute, or even read anything, if it’s one of a million clear BS pieces of information? So here’s my prediction:

The authenticity of content will become more important than its quality.

Rusty the cat

A story about the 33-year-old cat Rusty took Reddit by storm. Rusty passed away at the age of 33, making him into top 10 of the oldest cats in the recorder cat history. The poster submitted proof to the mods of r/cats that Rusty was real. People rushed to update Wikipedia’s list of oldest cats, expressed condolences, sent love and wishes. The post generated 134K likes and was probably viewed by most of the non-bot Redditors.

Unfortunately, Rusty was AI slop. Rusty:

The post and the bot that submitted it got deleted.

I think my attempt to restrict Reddit to a few essential subreddits like r/cats is not very successful and I still have exposure to something that I wouldn’t even call AI slop. More like farming for free human-generated text for the purpose of training LLMs. Rusty helps me understand why the sudden rush to gather personal IDs and verify humans on social media. All the social networks are vulnerable to slop and risk losing engagement if they don’t put it under some level of control.

Here’s a real orange cat for you, blissfully unaware about the decay of r/cats.

UPDATE: this is a male cat named Pesho, also known as The Son of a Mother. Likes cuddles and bites for no reason.

Muted Reddit’s auto-subscription after it became intolerable

Reddit somehow detected that I’m interested in AI and Data and auto-subscribed me to several hundred AI and Data subreddits. This turned their feed to an endless slop machine. Every single pre-IPO hype post or or doom prophecy would pop on my feed multiple times per subreddit, 5 if it was Sam Altman’s. Multiply that by at least 100 AI subreddits and you get the idea. You are scrolling and scrolling without seeing a single cat. Claude this, Sam Altman that, Anthropic this, I vibe coded this genius thing, and then Sam Altman again.

Rather than adding Reddit to the list of sites I ignore, I followed the subreddits where I’ve recently posted comments, and then disabled the following setting:

I greatly recommend this approach. Cats, rockets, and cars showed up again. The apocalypse is muted.