Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines a perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
this seems wildly naive
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike