Annotation by Dev Malhotra on Large Language Models explained briefly

Dev MalhotraDev Malhotra@devmalhotraSample Account?Sep 24, 2026Technology
Clip transcript
Explainer

The sentence to pause on is at 1:20: “even though the model itself is deterministic, a given prompt typically gives a different answer each time it's run.” Same weights and same prompt give the same probabilities every time. The variety is added afterwards, on purpose: the software sometimes picks a less likely word because the result reads more naturally. That's a setting in the code serving the model, not the model changing its mind.

So two different answers to the same question aren't two opinions. They're two draws from one set of odds. Worth remembering before anyone screenshots one of them as “what the AI thinks.”

Ellie BrennanEllie Brennan@elliebrennanSample Account?Sep 24, 2026

Small correction to “deterministic”: in practice, even with the randomness turned all the way off, you don't reliably get the same answer twice. Thinking Machines asked one open model the same question 1,000 times at temperature 0 and got 80 different completions. The cause is in the serving, not the sampling: how many other people's requests yours is batched with changes the arithmetic slightly.

“Surprisingly, we generate 80 unique completions, with the most common of these occuring 78 times.”

Defeating Nondeterminism in LLM Inferencethinkingmachines.ai
Dev MalhotraDev Malhotra@devmalhotraSample Account?Sep 24, 2026

@elliebrennan fair, and that's a better version of my point. “Deterministic” is true of the math, not of the server. Either way the answer you got depends on things that have nothing to do with your question, which is the part I'd want people to know.

Nadia HaddadNadia Haddad@nadiahaddadSample Account?Sep 24, 2026

This is the exercise I run with ninth graders now. Everyone asks the chatbot the same research question, then we put the answers side by side. Once they see three confident, different answers, “the AI said so” stops working as a citation.

Hana SatoHana Sato@hanasatoSample Account?Sep 24, 2026

Checked the 2,600 years at 1:34. It works if you read GPT-3's 300 billion training tokens as words at about 220 a minute, nonstop. A token is closer to three-quarters of a word, so in words it's nearer 2,000 years. Same order of magnitude, same point.

“All models were trained for a total of 300 billion tokens.”

Language Models are Few-Shot Learnersarxiv.org
Priya VenkatPriya Venkat@priyavenkatSample Account?Sep 24, 2026

@devmalhotra “two draws from one set of odds” is doing a lot of work there. If the odds put 70% on one answer, that is the model's view in any sense that matters to the person reading it. Different answers don't mean no opinion; they mean you saw one sample of it.

Dev MalhotraDev Malhotra@devmalhotraSample Account?Sep 24, 2026

@priyavenkat agreed, which is the case for asking twice. One sample tells you almost nothing about where the 70% sits.

Kwame AsanteKwame Asante@kwameasanteSample Account?Sep 24, 2026

Same idea as a season sim. You run it 10,000 times and report how often each team wins the league, not the one run where the relegation favourites won it.

Owen PriceOwen Price@owenpriceSample Account?Sep 24, 2026

Asked one of these what off-road diesel would run me this winter, twice, and got two different numbers. Went and called the co-op.