Most people describe an LLM one of two ways. Magic, or glorified autocomplete.
Both are wrong, and the truth is stranger than either.
So I took the whole thing apart. One sentence, six steps, no math required.
Here is the part that stuck with me. When a model writes the first word of an answer, it has no idea what the last word will be. There is no draft underneath. The coherence you are reading gets assembled one token at a time, in public, by a machine rolling dice inside a probability distribution.
Every individual step in here is simple enough to fit on a slide. Chop the text. Turn it into numbers. Weigh what matters. Refine it ninety six times. Produce odds. Roll. Repeat.
Then you run that across trillions of words and it starts passing the bar exam.
That gap between how boring the mechanism is and how strange the output gets is the most important unsolved problem of this decade. Anyone claiming to have it settled, in either direction, is selling you something.
Swipe through...
Suggested Credits
Tags, Events, and Projects