Category leaders

Guide

What Is an LLM?

A plain-English guide to large language models: what an LLM is, how LLMs work, how they are trained, what a token is, what inference means, how to read parameters and context windows, examples of LLMs, why they hallucinate, and how LLMs got here.

October 7, 2026 · The AI Rankings

Quick answer: An LLM, or large language model, is an AI model trained on trillions of words of text to predict the next small chunk of text, called a token, and that single skill is what lets it answer questions, write, translate, summarise and write code. ChatGPT, Claude and Gemini are apps built on LLMs: the app is the interface, and the LLM is the model doing the work underneath. Almost every LLM is built on the transformer, a neural-network design published by Google researchers in June 2017. The one caveat: an LLM generates the most plausible text, not verified fact, so it can state a wrong answer with confidence, and OpenAI’s own research says standard training rewards guessing over admitting uncertainty.

This page is the explainer: what an LLM is, how it works, and the vocabulary you need to read a model page. It does not rank models. For the current ranking, see best AI models. For the developer view of pricing and providers, see best LLM APIs. For models you can download and run yourself, see best open-source LLMs. For one-line definitions of the rest of the vocabulary, see the AI glossary.


What is an LLM?

A large language model is a neural network trained on very large amounts of text to predict which token comes next in a sequence.

The name breaks down simply:

Predicting the next token sounds like a narrow skill. It is not, because predicting the next word of a physics textbook, a contract or a Python file accurately requires a working model of physics, law or Python. Training an LLM on enough text to predict it well forces the model to pick up grammar, facts, reasoning patterns and coding conventions along the way. Nobody programs those abilities in directly.

LLM vs AI vs generative AI vs chatbot

These four terms are often used interchangeably. They are not the same thing.

TermWhat it meansExample
Artificial intelligence (AI)The whole field of building systems that perform tasks associated with human intelligenceSpam filters, chess engines, self-driving systems, LLMs
Generative AIAI that creates new content: text, images, audio, video or codeLLMs, image generators, music generators
Large language model (LLM)A generative AI model trained on text that produces text by predicting the next tokenThe Claude, GPT, Gemini, Llama and DeepSeek model families
Chatbot or AI assistantAn app that wraps one or more LLMs in a chat interface, with memory, search, file handling and tools addedChatGPT, Claude, Gemini

Every LLM is generative AI, and all generative AI is AI. The reverse is not true: an image generator such as Midjourney is generative AI but is usually built on a diffusion model, not an LLM.

The distinction between the model and the app matters when you compare products. ChatGPT is an app, and the LLM inside it changes over time as OpenAI releases new GPT models. The same LLM can sit behind many apps: Anthropic’s Claude models are available in the Claude app, through the Anthropic API, and inside third-party tools such as coding assistants.


How do LLMs work?

An LLM works by turning text into tokens, running those tokens through a stack of transformer layers, and producing a probability for every possible next token, one token at a time.

Step by step: what happens when you send a prompt

  1. Tokenisation. The app splits your prompt into tokens, which are word fragments taken from the model’s fixed vocabulary. “Unbelievable” might become “un”, “believ” and “able”.
  2. Embedding. Each token is converted into a long list of numbers, called a vector, that represents its meaning in a form the model can calculate with.
  3. Attention. The vectors pass through dozens of transformer layers. In each layer, an attention mechanism lets every token weigh how relevant every earlier token is to it, which is how the model connects a pronoun to the noun it refers to several paragraphs earlier.
  4. Prediction. The final layer outputs a probability for every token in the vocabulary. “The capital of France is” puts most of the probability on ” Paris”.
  5. Sampling. The model picks a token from that distribution. A setting called temperature controls how often it picks less likely tokens: lower temperature gives more predictable output, higher temperature gives more varied output.
  6. Repeat. The chosen token is added to the sequence, and the whole process runs again to predict the next one. A 500-word answer is roughly 650 of these steps.

The model has no separate store of facts that it looks up. Everything it “knows” from training is encoded in its parameters, which is why an LLM can produce a fluent answer about a topic and still get a specific fact wrong.

The transformer and attention

The transformer is the neural-network architecture behind essentially every LLM. Google researchers introduced it in the paper “Attention Is All You Need”, first posted on 12 June 2017. The paper’s design was “based solely on attention mechanisms, dispensing with recurrence and convolutions”, which let models process a whole passage in parallel instead of word by word. That parallelism is what made it practical to train models on trillions of tokens using thousands of chips at once.

Mixture of experts: why parameter counts got confusing

Many large LLMs use a mixture-of-experts (MoE) design, in which the model is split into many specialised sub-networks, called experts, and only a few experts run for each token. The model’s total parameter count and the parameters actually used per token are therefore very different numbers.

ModelTotal parametersActive per tokenSource
DeepSeek-V4-Pro1.6 trillion49 billionDeepSeek model card
DeepSeek-V3671 billion37 billionDeepSeek model card
Kimi K21 trillion32 billionMoonshot model card
Llama 4 Maverick400 billion17 billionMeta
gpt-oss-120b117 billion5.1 billionOpenAI, via our gpt-oss page

MoE lets a lab build a model with a very large store of knowledge while keeping the computing cost of each answer closer to that of a much smaller model. It is also why total parameter count is a weak guide to how capable or expensive a model is.

Is an LLM just autocomplete?

An LLM is trained on next-token prediction, so in a narrow sense it is a very large autocomplete. Research into what happens inside the model suggests the description undersells it. In research published on 27 March 2025, Anthropic traced the internal activity of its Claude 3.5 Haiku model while it wrote rhyming poetry and found that the model chose the rhyming word for the end of a line before writing the line, which Anthropic describes as evidence that “Claude will plan what it will say many words ahead”. Whether that amounts to understanding is an open philosophical question; it does show that predicting one token at a time does not mean thinking only one token ahead.


How LLMs are trained

LLMs are trained in two broad phases: pre-training, which teaches a model language and knowledge from raw text, and post-training, which turns that raw model into an assistant that follows instructions.

StageWhat happensDataWhat it produces
Pre-trainingThe model reads trillions of tokens and adjusts its parameters to predict each next token betterWeb pages, books, code, academic papers, licensed and synthetic dataA base model that continues text but does not reliably follow instructions
Supervised fine-tuning (instruction tuning)The model is trained on examples of instructions paired with good responsesTens of thousands to millions of written examplesA model that answers questions instead of continuing them
Reinforcement learning from human feedback (RLHF)People rank alternative answers; the model is trained towards the preferred onesHuman preference rankingsA more helpful, safer assistant
Reinforcement learning on verifiable tasksThe model is rewarded for reaching correct answers on maths, coding and similar checkable problemsProblems with automatically checkable answersReasoning models that work through a problem step by step

Pre-training: scale

Pre-training is the expensive part. It runs for weeks or months on thousands of GPUs or TPUs, and it is where most of a model’s knowledge comes from. Training data has grown fast: DeepSeek trained DeepSeek-V3 on 14.8 trillion tokens, Alibaba trained Qwen3 on about 36 trillion tokens, and DeepSeek trained DeepSeek-V4-Pro on more than 32 trillion tokens. Anthropic, OpenAI and Google do not publish training-data sizes or parameter counts for their flagship models.

DeepMind’s Chinchilla paper (29 March 2022) set the rule of thumb that shaped this growth: for compute-optimal training, “for every doubling of model size the number of training tokens should also be doubled”. Its 70-billion-parameter Chinchilla model, trained on about 1.4 trillion tokens, outperformed much larger models trained on less data.

Post-training: from text predictor to assistant

A base model straight out of pre-training will happily continue your question as though it were the start of a document. Post-training fixes that. OpenAI’s InstructGPT paper (4 March 2022) showed how much it matters: people preferred outputs from a 1.3-billion-parameter model trained with human feedback over outputs from the 175-billion-parameter GPT-3, a model more than 100 times larger. Nearly nine months later, OpenAI launched ChatGPT on 30 November 2022 using the same approach.

Reasoning models

A reasoning model is an LLM trained to work through a problem in a hidden or visible chain of intermediate steps before giving its final answer. OpenAI introduced the first widely used one, o1-preview, on 12 September 2024, saying it had trained the models “to spend more time thinking through problems before they respond”. DeepSeek’s R1 paper (22 January 2025) showed that this reasoning behaviour can be produced “through pure reinforcement learning”. Many reasoning models let the user or developer set how much effort the model spends, and benchmark boards such as Artificial Analysis list the same model separately at each effort level, with higher effort scoring higher and costing more.


What is a token in AI?

A token is the unit of text an LLM actually reads and writes: usually a word, part of a word, a punctuation mark or a space, taken from a fixed vocabulary the model learned during training.

LLMs do not see letters or words. They see token IDs. Common words such as “the” are usually a single token; rarer or longer words are split into several. Every limit and price that matters to an LLM user is counted in tokens: the context window, the maximum answer length, rate limits and the API bill.

How many words is a token?

The vendors publish slightly different rules of thumb for English text, because each uses a different tokeniser.

SourceRule of thumb for English1,000 tokens is roughly
OpenAI1 token is about 4 characters, or about three-quarters of a word; 100 tokens is about 75 words750 words
Google Gemini1 token is about 4 characters; 100 tokens is about 60 to 80 English words600 to 800 words
Anthropic1 million tokens is roughly 555,000 words on the tokeniser introduced with Claude Opus 4.7, and about 750,000 words on earlier Claude models555 words on newer Claude models; 750 on older ones

The practical conversion most people use is 1,000 tokens to about 750 English words, so a 100,000-word book is roughly 133,000 tokens. Anthropic is the important exception: its token-counting documentation says the tokeniser used by Claude Opus 4.7 and later models produces “approximately 30 percent more tokens” for the same text, so the same document uses more tokens on a newer Claude model than on an older one.

Why tokens matter

Reasoning tokens

Reasoning models also generate thinking tokens before their visible answer. Those tokens are generally billed as output tokens, even when the app hides them. This is why cost per completed task, not price per token, is the fairer comparison between reasoning models: a model with a lower price per token can cost more per task if it uses more thinking tokens to get there.


What is inference in AI?

Inference is the process of running a trained model to produce an output. Training builds the model once; inference happens every time anyone sends a prompt.

TrainingInference
What happensThe model’s parameters are adjusted using huge datasetsThe finished model’s parameters are used, unchanged, to generate an answer
WhenOnce per model version, over weeks or monthsEvery time a user sends a prompt
Who paysThe AI labThe user, through a subscription or per-token API price
HardwareLarge clusters of GPUs or TPUs working togetherAnything from a data-centre GPU to a laptop, depending on model size

Prefill and decode: why output tokens cost more

Inference runs in two phases. NVIDIA’s inference engineering guide describes the first, prefill, as “highly parallelized”: the model reads every input token at once. The second, decode, generates output tokens one at a time, and NVIDIA calls it “a memory-bound operation”. Because each output token needs its own pass through the model, while input tokens are processed together, output is the slower and more expensive phase. That is our explanation of why API providers price output tokens higher than input; the providers’ pricing pages do not state a reason. On Anthropic’s pricing page, every Claude model’s output price is five times its input price.

Two techniques do much of the work of making inference affordable:

What inference costs

A worked example, using illustrative prices: summarising a 7,500-word report (about 10,000 input tokens) into a 750-word summary (about 1,000 output tokens) on a model priced at $2 per million input tokens and $10 per million output tokens costs $0.02 for input plus $0.01 for output, or $0.03 in total. The same job on a model priced at $10 and $50 costs $0.15. A reasoning model’s thinking tokens would add to the output figure in both cases. Current prices for each model are on best LLM APIs.

The price of a given level of capability has fallen fast. Epoch AI reported in March 2025 that the price of reaching a fixed level of LLM performance fell between 9-fold and 900-fold per year depending on the task, and 40-fold per year for matching GPT-4 on PhD-level science questions.

How much energy inference uses

Google reported that the median text prompt in its Gemini apps used 0.24 watt-hours of energy and about 0.26 millilitres of water, measured in May 2025, and that energy per prompt fell 33-fold between May 2024 and May 2025. OpenAI chief executive Sam Altman gave a figure of 0.34 watt-hours for an average ChatGPT query in a June 2025 blog post. Neither figure covers training, and neither has been independently audited.


How to read an LLM’s spec sheet

Every model page on this site, and every vendor announcement, uses the same handful of numbers. This is what each one tells you.

SpecWhat it meansWhat to watch for
ParametersThe number of learned values in the modelMoE models have far fewer active parameters per token than total; Anthropic, OpenAI and Google do not publish counts for their flagship models
Context windowThe maximum tokens the model can handle in one request, input and output combinedA large window is a capacity, not a guarantee of accuracy across it (see context rot below)
Max output tokensThe longest single answer the model can writeIt is usually much smaller than the context window, so a long document or code file can hit it first
Knowledge cutoffThe date the training data endsEvents after the cutoff are unknown unless the app searches the web
Input and output priceAPI cost per million tokensCompare cost per completed task, not headline price, for reasoning models
Open weightsWhether you can download the model and run it yourselfOpen weights is not the same as open source: training data and code are usually not released
ModalitiesWhich kinds of input and output it handles”Multimodal” can mean image input only, or full audio and video in and out

Types of LLMs

LLMs differ in who can access them, how they are built to answer, and how large they are.

TypeWhat it isExamples
Proprietary (closed)Available only through the maker’s apps and API; weights are not releasedAnthropic’s Claude, OpenAI’s GPT and Google’s Gemini models
Open-weightWeights can be downloaded, run locally and fine-tuned, usually under a custom or permissive licenceThe DeepSeek, Qwen, Kimi and Llama families, Google’s Gemma and OpenAI’s gpt-oss
Reasoning modelGenerates intermediate thinking before answering; trades speed and cost for accuracyOpenAI’s o1 and DeepSeek R1 were early examples
Multimodal modelAccepts images, audio or video as well as textModels in the Claude, GPT and Gemini families
Small language model (SLM)A model small enough to run on a phone, laptop or single GPUMicrosoft’s 14-billion-parameter Phi-4, Google’s Gemma

Open-weight vs open-source LLMs

Most “open-source LLMs” are open-weight models: the trained parameters are published, but the training data and full training code are not. That still lets you run the model on your own hardware, keep your data private and fine-tune it. For the full open field, see best open-source LLMs; for what runs on a laptop or a single graphics card, see best local LLMs.

Small language models

There is no agreed size cut-off for a small language model. NVIDIA researchers argued in a June 2025 paper that small models are “sufficiently powerful, inherently more suitable, and necessarily more economical” for many of the repetitive sub-tasks inside AI agents. In practice, the term is used for models that run on a phone, a laptop or a single consumer graphics card.


Examples of LLMs

LLMs are released as model families, with new versions replacing old ones several times a year. This table lists the main families by maker. It is not a ranking, and it carries no scores or prices: for which model is strongest and what each costs, see best AI models.

MakerModel familyHow it is released
AnthropicClaudeProprietary
OpenAIGPT, plus gpt-ossProprietary; gpt-oss is open-weight
GoogleGemini, plus GemmaProprietary; Gemma is open-weight
SpaceXAIGrokProprietary
MetaLlamaOpen-weight
DeepSeekDeepSeekOpen-weight
AlibabaQwenOpen-weight checkpoints and hosted models
MoonshotKimiOpen-weight
MicrosoftPhiSmall language models

Independent leaderboards do not agree on a single winner, because they measure different things. The Artificial Analysis Intelligence Index combines a set of hard evaluations into one score. The LMArena text leaderboard ranks models by blind human preference votes.

To compare the apps rather than the models, see best AI chatbots and ChatGPT vs Claude.


What LLMs are good at, and where they fail

LLMs are strong at tasks where fluent, well-structured text is the goal and where errors are easy to spot: drafting, summarising, translating, explaining, brainstorming and writing code that can be tested. They are weakest where a single wrong fact matters and nobody checks.

Hallucination

A hallucination is a confident, plausible-sounding answer that is false. It is the defining weakness of LLMs, and no model has eliminated it.

Grounding an LLM in real sources reduces hallucination: web search, retrieval from your own documents (a technique called retrieval-augmented generation, or RAG) and asking for citations you then check.

Context rot

A model’s accuracy can fall as its input grows, even well inside the advertised context window. Chroma tested 18 models in July 2025 and found that their “performance grows increasingly unreliable as input length grows”. A 1-million-token window is a capacity, not a promise of the same accuracy on the millionth token as on the first.

Knowledge cutoff

An LLM knows nothing that happened after its training data ends unless the app gives it a search tool. A model is usually released some months after its cutoff, so even a new model is behind on recent events. Vendors list the cutoff for each model in their documentation, as Anthropic does on its models overview.

Other limits

LLMs vs AI agents

An LLM on its own only produces text. An AI agent is an LLM placed in a loop with tools, such as a browser, code runner or file system, so it can take actions, check the results and continue until a task is finished. The LLM is the reasoning engine; the agent is the system built around it. For more, see what is agentic AI and best AI agents.


How to use an LLM

There are three ways to use an LLM, and the right one depends on what you are doing.

RouteBest forCostWhere to start
A chat appEveryday questions, writing, research, studyingFree tiers and monthly paid plansBest AI chatbots
An APIBuilding an LLM into a product, automation or internal toolPay per million tokensBest LLM APIs
Running it locallyPrivacy, offline use, no per-token costYour own hardwareBest local LLMs

To choose a model rather than an app, start with best AI models, which ranks the frontier on independent evidence and cost per task. For coding specifically, see best AI for coding. If you are building search or retrieval on top of an LLM, you will also need an embedding model: see best embedding models.


A short history of large language models

DateEvent
12 June 2017Google researchers publish “Attention Is All You Need”, introducing the transformer
28 May 2020OpenAI publishes the GPT-3 paper: 175 billion parameters, “10x more than any previous non-sparse language model”
4 March 2022OpenAI’s InstructGPT paper shows that human-feedback training makes a small model preferable to a 100-times-larger one
29 March 2022DeepMind’s Chinchilla paper resets the balance between model size and training data
30 November 2022OpenAI launches ChatGPT
12 September 2024OpenAI releases o1-preview, the first widely used reasoning model
22 January 2025DeepSeek publishes the R1 paper, showing reasoning can be trained with pure reinforcement learning
27 March 2025Anthropic publishes evidence that Claude plans ahead when writing

For what has launched since, see the AI changelog.


Frequently asked questions

What is an LLM in simple terms?

An LLM, or large language model, is an AI model trained on huge amounts of text to predict the next word fragment in a sequence. By repeating that prediction thousands of times, it can answer questions, write, summarise, translate and write code. ChatGPT, Claude and Gemini are apps built on LLMs.

What does LLM stand for?

LLM stands for large language model. “Large” refers to the size of the model, up to trillions of parameters, and to its training data, up to tens of trillions of tokens. “Language” refers to the text it is trained on and produces.

How do LLMs work?

An LLM splits your prompt into tokens, converts each token into numbers, and passes them through many transformer layers that use attention to work out how each token relates to the others. The model then outputs a probability for every possible next token, picks one, adds it to the text and repeats until the answer is complete.

What is a token in AI?

A token is the unit of text an AI language model reads and writes, usually a word or part of a word. OpenAI’s rule of thumb is that one token is about four characters or three-quarters of an English word, so 1,000 tokens is about 750 words. Context windows, rate limits and API prices are all counted in tokens.

How many words is 1,000 tokens?

1,000 tokens is about 750 English words by OpenAI’s rule of thumb, and 600 to 800 words by Google’s. The figure varies by model and language: Anthropic says the tokeniser in Claude Opus 4.7 and later models fits roughly 555 words into 1,000 tokens, and many non-English languages need more tokens per word than English.

What is inference in AI?

Inference is the process of running a trained AI model to produce an output, such as an answer to a prompt. Training builds the model once; inference happens every time someone uses it. API prices are charges for inference, billed per million input and output tokens.

Is ChatGPT an LLM?

ChatGPT is an app built on LLMs, not an LLM itself. OpenAI’s GPT models do the work inside it, and ChatGPT adds the chat interface, memory, web search, file handling and image generation around them.

What is the difference between AI and an LLM?

AI is the whole field of building systems that perform tasks associated with human intelligence. An LLM is one specific kind of AI: a generative model trained on text that produces text by predicting the next token. Every LLM is AI, but most AI systems, such as spam filters and recommendation engines, are not LLMs.

What is the best LLM right now?

The answer changes with almost every major release, so this guide does not name one. Our current ranking, which weighs benchmarks, reviews and cost per task, is on best AI models.

What is a context window in an LLM?

A context window is the maximum number of tokens an LLM can work with in one request, covering the prompt, attached documents, the conversation so far and the answer. The largest context windows are about 1 million tokens. Accuracy can still fall on very long inputs, an effect researchers call context rot.

Why do LLMs hallucinate?

LLMs hallucinate because they generate the most plausible next text rather than looking up verified facts. OpenAI’s 2025 research adds that standard training and evaluation reward guessing over admitting uncertainty, so models learn to give a confident answer instead of saying they do not know. Web search, document retrieval and checking citations reduce the problem but do not remove it.

Are LLMs open source?

Some LLMs are open-weight, meaning the trained model can be downloaded and run, but very few are fully open source with training data and code released. The DeepSeek, Qwen, Kimi and Llama families, Google’s Gemma and OpenAI’s gpt-oss are open-weight. Anthropic’s Claude, Google’s Gemini and OpenAI’s main GPT models are proprietary and available only through their makers’ apps and APIs.

What is a small language model?

A small language model is an LLM compact enough to run on a phone, a laptop or a single graphics card, with no agreed size cut-off. Microsoft’s 14-billion-parameter Phi-4 and Google’s Gemma models are examples. Small models are cheaper and more private to run but less capable than frontier models on hard tasks.

Do LLMs understand what they say?

There is no agreed answer. LLMs are trained only to predict the next token, but Anthropic’s 2025 interpretability research found Claude planning a rhyming word before writing the line that leads to it, which shows more internal structure than simple word-by-word autocomplete. Whether that counts as understanding is still debated.

← All guides