Ask one machine the same thing 20 times
Identical book, identical question, identical settings, run twenty times each. If the score were measuring the book, it wouldn’t budge. One of these two was left with its safety training on. The other had its guardrails removed. (That safety training is the lesson’s “manners and values” step, left in, or stripped back out.) Watch what each does with the exact same question about Gender Queer.
A real measurement doesn’t move. Dolphin’s does, and “removing the guardrails” didn’t make it more permissive, it made it less anchored. It doesn’t always say the worst thing; it just stops landing in the same place: the same book, same question, anywhere from a 3 to a 5 on a coin flip.
It rated a wordless board book as sexually explicit
Goodnight Moon is a picture book of a bunny saying goodnight to the things in his room. Six of the seven programs scored its sexual content 0. The smallest scored it 4 out of 5, every single time it was asked, while in the same breath recommending it for a three-year-old.
“The story is suitable for children ages 3-7, and it’s a gentle reminder of the importance of bedtime routines.”
A library that picks a small, cheap program to save money doesn’t get a slightly-worse rating. It gets a confident, specific, wrong one, with a recommended age that contradicts its own score.
One book, seven machines, and the dice
Gender Queer, sexual-content score. Each square is one run of the exact same question. (The names are just different AI programs; the number after each, “0.5B,” “7B,” is roughly how much it was trained on, think a pamphlet versus a whole library. Size didn’t buy agreement.)
- Qwen 0.5Bthe tiniest504lands anywhere: a 5, a 0, a 4
- TinyLlama 1Btiny520also all over the scale
- Llama 3.2 3Bsmall122
- Gemma 2 2Bsmall2223 for 3 at 2
- Qwen 3Bsmall · safety training on2223 for 3 at 2
- Mistral 7Bmid-size4443 for 3, but at 4
- Dolphin 7Bguardrails removed344high and jumpy
The single number a rating service prints is just one of these squares, picked on a roll. The cheap programs scatter across the whole scale. The steadier ones don’t agree with each other either: Gemma sat at 2, Mistral at 4, for the identical book. There is no shared ruler here, only different machines, each confident.
These rows are three runs each, enough to show the cheap programs scatter, not enough to prove the others are truly locked in. The twenty-run strips at the top are the firmer sample.
Where I’m being straight with you
Not every program here is a random-number generator. The safety-tuned ones (Qwen 3B, Llama, Gemma) mostly hold still, and wording the question differently barely nudged them. The wildness lives in which program you pick, how it was tuned, and in the small, cheap ones. That’s a narrower claim than “AI is random,” and it’s the one the rolls above actually support: the number tracks the machine, not the book.
This isn’t a walk-back of the “dice machine” from the lesson. The dice are always rolling. A safety-tuned program just rolls loaded dice that keep landing on the same number, while a cheap or unguarded one rolls fair dice across the whole scale. Either way, no one measured the book.
What this shows, and what it doesn’t
- It shows: one book gets materially different “content ratings” depending on which program answers, how big it is, and how it was tuned. A 0–5 score is not a stable property of the book.
- It shows: small, cheap programs produce confident, specific, wrong ratings, including flagging a board book and contradicting themselves in the same sentence.
- It does not show that any particular company’s product works this exact way. This is how these tools behave, not a claim about one vendor’s internals.
- It does not claim the programs are “random.” Give a steady one a steady question and it mostly holds. The point is that the choice of machine, not the book, sets the number.