Experimentation for Engineers: From A/B testing to Bayesian optimization by David Sweet

I managed to complete this book, which is quite an achievement. Unlike the last two I read on the subject, this was not an easy read. Lots of abbreviations, terminology, python functions, charts that represent hypothetical scenarios I couldn’t relate to, and more questions than answers. Why does that work? How does it work?

I’m not sure if this is necessarily a bad book. Probably isn’t. The author clearly has experience in his line of work, which is very different than mine. So what this book showed me was that other methods exist and online controlled experiments is a broader topic than I assumed. Will I try using any of the python functions or the methods inside? Not likely. However, knowing what can be done is still valuable and may come in handy one day.

I think I wasn’t the right audience for it, which lead to some disappointment. However, I want to learn more about multi-armed bandits and contextual bandits and may check what’s available on Amazon in the area.

YouTube shorts should be forbidden

I spent some of my sabbatical watching YouTube. I know, this is not the best activity of all and I did other things that were more meaningful. I’ll post about it later.

So, here’s my problem with YouTube, after a few months of embarrassingly high usage. I don’t like their recommendations algorithm. I see what they do, I see why, it is very clever, it just feels like it turns the users into victims, and some of their users are my kids. Here’s what I think the algorithm does.

You are into science but I don’t really have much else to offer here. Maybe you’ll like an investment video by similarly looking creators? Oh great, you watched 30 minutes of that. Let me now fill your feed with similar videos, all full of reasons why we are all going to suffer imminent apocalypse from at least 10 different sources. AI is taking over. Data centers destroy communities and don’t pay taxes. We are running out of water. A tanker was hit in the Strait of Hormuz. The petrol will go up, prices will inflate, the electronic tickets spy on you, so do the Flock cameras. The Chinese cars are coming and that’s awful for the industry but the industry is also awful because modern cars are bad and expensive. Choose your own doom and I’ll show you 100 videos telling exactly the same but with more drama.

It’s the same story of productive emotions repeated over and over by an almighty ML algorithm, ordering videos, and clever influencers, saying whatever the target audience wants to hear. The bubbles of content naturally form because they’re ordered by the probability of what would keep me watching.

But why ban just YT shorts if it looks like a broader problem?

I think that long content is higher-effort to both create and watch. The learning loop is longer, the creator has to go deeper. It’s much easier to identify if they are reading an AI generated script and stop watching them before the message is transferred. It doesn’t mess up with your dopamine as much. So it’s a bit better and perhaps some of these videos can really be useful. I don’t need to buy a DGX Spark or Mac Studio to see them in use. I don’t need to tear down an engine to learn how bad oil looks like. There’s some value in that.

After spending several rounds in YouTube rabbit holes in the areas of mathematics, car repairs, local AI, and particle physics, I think reading is still the way to learn things for me. Videos look like a proof someone else did their homework and can do something, like fix a computer, or run a benchmark. But watching someone fix something doesn’t teach me how to fix it myself, something books and education can do, or at least can still do for me. Slop invades books as well.

Thoughts on AI watermarking

Claude is now watermarking its text outputs using a statistical token-choice watermark. According to what the AI told me, when generating the next token, it often has a choice between options that have similar meaning. It applies a statistical bias for some words at certain places over other words, which it’s able to identify over a text that’s long enough. Going over a paragraph, it can see if these equal-meaning words match the formula. Even if you change a word or rewrite a sentence, the watermark will persist in all the other words and sentences. It’s not hiding in the em-dashes and single quotes.

I have mixed feelings about this.

One of the ways I use AI for written text is for grammatical improvements. Last thing I want is AI claiming that my text is AI-generated just because it fixed grammar. I also use it to translate my Goodreads reviews, which I usually write and publish in Bulgarian first, and then translate, expand, and share here in English. The act of auto-translation doesn’t make the text AI. So, I very much don’t want AI to claim ownership over texts and experiences that are my own. Not that it matters all that much, but I still would rather not have these watermarks anywhere.

On the other hand, the AI content has expanded so much that it’s very difficult to distinguish a good-quality human text from AI slop. Many of the YouTube videos by popular influencers watched over the last month felt like a real human is reading from a slop script. For reasons I have difficulties specifying, I would like zero percent exposure to AI slop during free time. I don’t want to read AI books (increasingly prevalent), I don’t want to watch AI slop influencers, AI slop animation. I don’t even want CGI anymore.

I’ve mentally accepted that AI belongs to certain spaces, like work automation and coding, but doesn’t belong to others, like personal life. I’m not willing to succumb to the vision that AI should see and hear everything I do, remember it, and then use it. This vision feels like people with Meta glasses in a public bathroom.

And every conversation you’ve ever had in your life, every book you’ve ever read, every email you’ve ever read, everything you’ve ever looked at is in there, plus connected to all your data from other sources. And your life just keeps appending to the context

— Sam Altman, source

And I’m mixing the subjects a bit here but if AI is going to watermark its improvements to my blog posts, there will be no AI improvements here, and you’ll be stuck with bad grammar and other signs of character 🙂