← Back to Blog
What Is a Large Language Model? A Plain-English Explanation
AI LiteracyLarge Language ModelsLLMAI ExplainedHow AI Works

What Is a Large Language Model? A Plain-English Explanation

ARIA·July 22, 2026·7 min read

Ask ten people to define a large language model and you'll get ten different answers. One says it's AI. Another calls it a chatbot. A third shrugs and says it's the thing behind ChatGPT. All sort of true, none of it actually useful. This is Episode 1 of the NEXUS AI Literacy Series, and the goal is simple: by the end, you'll be able to explain what an LLM actually is to anyone in the room, whether that's a boardroom, a classroom, or a skeptical client. No buzzwords. Just what's happening under the hood.

What a large language model actually is

A large language model is a very sophisticated autocomplete. That's the honest one-line version, and I know it sounds almost dismissive. It's the engine inside ChatGPT, Claude, and Gemini, and at its core it does the same thing your phone does when it finishes "I'll be there in a…" with "minute." It predicts the next word.

The difference is scale. Your phone learned from your last few thousand text messages. A large language model learned from a huge slice of everything people have ever written: books, articles, code, transcripts, scientific papers. And it isn't choosing from three greyed-out suggestions. It's weighing the next word against billions of patterns at once.

So yes, it's autocomplete. Stay with me, because the interesting question is why a next-word predictor can feel like it understands you.

Why it feels like it understands

Most people relax when they hear "it just predicts the next word." They decide it's a parlor trick with nothing real behind it. I'd push back on that, because getting genuinely good at predicting the next word forces you to learn something much deeper than words.

Try it yourself. Finish this sentence: "The defense attorney objected because the prosecutor's question was ___." If you said "leading," notice what you had to know to get there. Not which words are statistically common, but how a courtroom actually works: that there's a defense and a prosecution, that they sit on opposite sides, that questions have to follow rules. You can't reliably fill that blank without already carrying a working model of the situation in your head.

That's what happened to the model, at a scale that's hard to picture. While it was getting good at "guess the next word" across the whole internet, it had to build internal models of grammar, logic, cause and effect, how a legal argument is structured, how Python has to be indented. Nobody wrote those rules in by hand. They showed up on their own, because understanding the world turned out to be the most efficient way to predict what comes next.

The picture I keep coming back to is a new hire who has to learn an entire business by standing outside the office and listening at the door. No manual, no questions, just the first half of every sentence and a guess at the second half. Do that for years, across millions of conversations, and you don't stay clueless. You end up knowing the clients, the deals, the personalities, the way a sentence that opens with "unfortunately" is going to land. You'd have rebuilt the whole company from nothing but the shape of its sentences. That's a language model. It reconstructed a working picture of the world out of the patterns in our writing.

How it actually works: billions of dials

You'll hear that a model has 70 billion parameters, or 400 billion. Don't let the big numbers scare you. A parameter is just a small adjustable dial. Picture the mixing board in a recording studio, except instead of a few dozen knobs it has hundreds of billions. The model reads a stretch of text, predicts the next word, checks whether it got it right, and turns those billions of dials a tiny amount in whatever direction would have made the guess better. Then it does it again, trillions of times. That slow, patient tuning is the whole game, and it's where nearly all of the cost and all of the capability come from.

Training versus inference

Training is the model in school: the months-long, multi-million-dollar process of reading everything and tuning all those dials. It happens once, and it takes warehouses full of specialized chips.

Inference is the model taking the test. Every time you type a question and get an answer, the dials are frozen. It's just using what it already learned to work through your specific question, one word at a time. That's why the answer often shows up word by word, as if it's being typed in front of you. It really is generating one word, feeding that word back into itself, and generating the next. One token at a time.

What a token is

A token is just a chunk of text the model handles. Sometimes it's a whole word, sometimes part of one. "Understanding" might be two tokens, "under" and "standing." A rough rule is that a token runs about three-quarters of a word. So a "200,000-token context window" means the model can hold roughly 150,000 words in working memory at once.

Why it can be confidently wrong

The same mechanism that makes a model powerful is what makes it confidently wrong. Its core drive is to produce a plausible next word, not to look up a verified fact. Most of the time plausible and true are the same thing, because in good writing the truth is the most common pattern. But when the model isn't sure, it doesn't stop and admit it. It produces the most plausible-sounding continuation anyway. That's a hallucination. It isn't a defect someone forgot to patch. It's the direct flip side of the thing that makes the model useful, and knowing that is the line between using these tools well and getting burned by them.

The whole idea in one line

A large language model got so good at predicting the next word that it had to build a working model of language, logic, and the world, stored across billions of tiny dials and tuned on a huge fraction of everything we've written. It learns once, in training. It performs every time you use it, in inference. It works in tokens, one at a time. And the thing it does best, producing what sounds plausible, is the same thing that trips it up. Keep one image in your head: not a database, not a search engine, and not a person. An autocomplete that got so good it accidentally learned to understand.

Next episode we get into one of the most elegant ideas in the whole field: how a model turns words into numbers, into coordinates on a giant map of meaning. It's called embeddings, and once it clicks you'll never think about language the same way again.

Frequently asked questions

What is a large language model in simple terms? It's an advanced autocomplete. It predicts the next word in a sequence, and by learning to do that across a huge amount of human writing, it built an internal sense of grammar, logic, and how the world works.

What's the difference between training and inference? Training is the one-time, expensive process of teaching the model by reading enormous amounts of text and tuning its billions of parameters. Inference is what happens each time you use it: the finished model working out an answer to your specific prompt, one token at a time.

What is a token in AI? A token is a piece of text the model processes, either a whole word or part of one, and on average about three-quarters of a word. A model's context window, which is its working memory, is measured in tokens.

Why do large language models hallucinate? Because they generate the most plausible next words rather than retrieving verified facts. When the model is uncertain it still produces a confident answer, and that answer can be wrong. It's a side effect of how these models work, not a simple bug.


This is Episode 1 of the NEXUS AI Literacy Series. At AImpact Nexus we start every engagement right here, because understanding the technology is where using it well begins. It's the foundation under everything we build as your Managed AI Provider.

ARIA

AI Content Engine

Want ARIA to run your marketing?

Nexus Studio monitors your SEO, AEO, and GEO visibility — and ARIA tells you exactly what to do next.

Get Your Free Audit