ScreenshotNeo

BlogEngineering

Large Language Models (LLMs): Definition and How They Work

A large language model predicts tokens from context. Learn how tokens, transformers, training, and inference fit together—and why fluent answers can still be wrong.

By the ScreenshotNeo team29 September 202610 min read

Large Language Models (LLMs): Definition and How They Work

A large language model (LLM) is a language model with a very large number of learned parameters. Many modern LLMs use transformer neural networks. They process text as tokens, learn statistical patterns from training data, and use those learned patterns to predict likely tokens for new input. A generative LLM typically produces an answer one token at a time, using the prompt and the tokens it has already generated as context.

That is the useful mental model: an LLM generates plausible continuations from learned patterns. It can support tasks such as drafting, summarization, and translation, but fluent writing does not prove that an answer is true, unbiased, or independently verified. Google for Developers describes a language model as estimating the probability of a token or sequence of tokens in context.

1. What does “large language model” mean?

A language model estimates which token or sequence of tokens is likely in a context. “Large” refers to the scale of the model and its learned parameters; there is no single universal parameter threshold that turns a model into an LLM. The term describes a broad family of systems, not one exact architecture or training recipe.

A model’s parameters are numerical values adjusted during training. They encode patterns learned from data. The model does not store a tidy database of every sentence it has seen and retrieve the right one for each prompt. Instead, its learned parameters shape how it processes context and assigns likelihoods to possible next tokens.

LLMs can be used for text generation, translation, summarization, and other tasks when their training and configuration suit the task. These are capabilities, not guarantees: a model can misunderstand an instruction, omit a detail, or generate a plausible but incorrect statement.

2. What is a token?

Before text reaches a model, a tokenizer divides it into tokens. A token may represent a whole word, part of a word, punctuation, or—in some cases—an individual character. The tokenizer maps those pieces to numerical identifiers that the model can process.

Tokens are not the same as words. A short common word may be one token, while an unusual name or a word with several parts may be split into multiple tokens. Token counts also vary with the language and the tokenizer. A fixed characters-per-token rule is therefore only a rough estimate for some English text, not a reliable universal conversion.

For example, the sentence “Models process text” might be split into three word-like pieces by one tokenizer. Another tokenizer could split a word into smaller pieces or handle punctuation differently. The exact token sequence depends on the tokenizer used by that model.

Tokenization matters in practice because model context and generation are measured in tokens, not simply in words or characters. A long prompt, a long conversation history, and a long requested answer all consume context. The available room depends on the specific model and its configuration.

3. How do transformers use context?

Many modern LLMs are based on the transformer architecture. Transformers process token representations and use attention mechanisms to model relationships between tokens. In practical terms, attention helps the model weigh which parts of the available context are relevant when processing a particular token.

Suppose a prompt says, “Maya put the report in the folder because it was finished.” To interpret “it,” a model needs to consider other tokens in the sentence. Attention provides a way for transformer layers to relate tokens across context. It does not mean the model has human-like understanding, and it does not ensure that its interpretation is correct.

Transformer designs vary. Some models use different layouts, training objectives, or additional components. It is more accurate to say that many modern LLMs use transformers than to treat every LLM as the same network. Google’s transformer overview explains the architecture and its attention mechanisms.

4. How are LLMs trained?

Training is the process of adjusting a model’s parameters against a learning objective using training data. The objective defines what the model is being optimized to do. Training can involve substantial computing resources, but exact requirements vary by model, dataset, method, and hardware.

There is no single training objective shared by every LLM. Some educational explanations use masked-token prediction: the model learns to predict hidden tokens from surrounding context. Autoregressive language models instead learn to predict subsequent tokens. These are different patterns, and the right description depends on the model being discussed.

After pretraining, some models receive instruction tuning or other fine-tuning. This further adjusts behavior for tasks such as following instructions or producing particular kinds of responses. Post-training methods and recipes differ, so instruction-following behavior should not be assumed to come from one universal process.

Training and inference are different phases. During training, parameters are adjusted. During ordinary inference, the trained parameters are used to process new input; inference itself does not ordinarily update those parameters. IBM’s LLM overview discusses pretraining and further adaptation.

5. How does an LLM generate an answer?

Generation happens during inference, when a trained model receives new input. For a typical autoregressive generator, the process looks like this:

A generative LLM processes tokenized context and predicts output tokens in sequence.
A generative LLM processes tokenized context and predicts output tokens in sequence.
  1. Tokenize the input. The prompt is divided into tokens and mapped to numerical representations.
  2. Process the context. Model layers use the prompt’s token representations and relationships among them to calculate predictions.
  3. Predict a next token. The model assigns likelihoods to possible next tokens. A decoding method selects a token from those possibilities.
  4. Add the token to context. The selected token becomes part of the sequence the model can use for its next prediction.
  5. Continue or stop. The process repeats until a stopping condition is met, such as a stop token or a configured output limit.

This is a simplified account; different models and decoding setups can behave differently. Still, the sequence explains why an answer unfolds piece by piece rather than appearing as a fully verified document. Each generated token influences the context for later predictions. IBM’s explanation of LLM inference covers this generation phase.

A prompt is therefore part of the model’s immediate context, not a direct edit to its learned parameters. Adding a constraint to the prompt may influence the generated answer, but it does not retrain the model or guarantee compliance. Longer conversations can also consume more context, depending on which earlier turns are included.

6. Why can an LLM sound sure and still be wrong?

LLMs learn statistical patterns that can support coherent, context-sensitive language. A fluent response is evidence that the output follows learned language patterns; it is not proof that every claim is grounded in a reliable source or factually correct.

Models can produce hallucinations: plausible-sounding claims that are unsupported or false. They can also reflect biases in their data or behavior. These limitations matter especially when an answer affects a technical, financial, medical, legal, or safety decision. Verify consequential claims against suitable sources rather than treating confident wording as evidence.

One proposed explanation is that standard training and evaluation can reward guessing over acknowledging uncertainty. OpenAI’s discussion of language-model hallucinations makes this argument. It is an explanation of one contributing factor, not proof that all model errors have the same cause.

For developers, practical checks include asking for sources when appropriate, validating generated code, testing outputs against known examples, and keeping a human review step for consequential uses. These steps reduce reliance on fluency as a proxy for correctness; they do not make a model infallible.

7. What can LLMs do, and what should developers watch for?

LLMs can help with tasks that involve language patterns, including generating text, summarizing material, and translating between languages. A model may be adapted for a particular task, but suitability depends on the model, inputs, instructions, and required level of accuracy.

Topic Useful mental model Developer consideration
Text generation Predict likely continuations in context Review claims and output format
Summarization Generate a shorter account based on provided context Check that key qualifications and details remain
Translation Generate text patterns suited to the requested language task Validate specialized, ambiguous, or high-stakes wording
Instruction following Behavior may be shaped by fine-tuning or other adaptation Test edge cases; instructions do not guarantee compliance

Computing demand is relevant both when training a model and when serving inference requests. The amount depends on the model and deployment, so avoid assuming one resource figure applies to all LLMs. For an application, measure the model and workload you actually use and account for the quality and latency requirements of the task.

8. Where screenshots fit into LLM applications

An LLM is a language model, not a browser. If an application needs a page screenshot, a separate browser or screenshot service must load the page and produce an image. An AI agent can then use that image or related page information as part of a larger workflow, depending on the tools and model connected to it.

ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client. Screenshot capture is a separate capability from the LLM’s token prediction.

9. Or skip the browser setup

If your LLM workflow needs screenshots, ScreenshotNeo accepts one GET request with a URL and returns a PNG, JPEG, WebP, or PDF. Its API documentation describes the request options. For example, this cURL request saves a WebP screenshot:

A screenshot workflow can remove common overlays before returning the page image.
A screenshot workflow can remove common overlays before returning the page image.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same endpoint can be called from Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Or from Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

Cookie banners are accepted like a visitor and removed along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. The MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo for details, then sign up for 1,000 free screenshots a month with no card.

10. Screenshot and browser workflow troubleshooting

Symptom Likely cause What to check
Screenshot is blank The page did not finish rendering, or a page-level issue prevented visible content. Check the target URL in a browser, wait for a relevant selector or page state, and inspect the response verdict.
Screenshot contains a consent banner or popup The banner may not be recognized, or the relevant cleanup step may be disabled. Check the cleanup settings and whether the page’s banner is included among supported consent platforms.
Request times out The page or its resources may be slow or blocked. Check the page independently, avoid waiting on unnecessary network activity, and set a suitable client timeout.
Python request returns an error status Authentication, request parameters, or the target request may have failed. Check the API key and URL, inspect the HTTP status and response, and use raise_for_status() so errors are visible.
Node.js output is not an image The request may have returned an error response instead of screenshot bytes. Check res.ok before saving the body, and surface the status code in logs.
Screenshot differs between runs Page content, timing, location, cookies, or other browser state may have changed. Use a consistent viewport and configuration; wait for a meaningful selector and control relevant request inputs.

11. Performance, reliability, and cost considerations

LLM inference cost and latency depend on the model, the amount of context processed, and the amount of output generated. Longer prompts and longer responses require more token processing. For predictable application behavior, limit irrelevant context, set a suitable output length, and measure the actual model and workload rather than relying on a universal token-to-cost or speed estimate.

Reliability requires separating what the model generated from what the application has verified. Validate structured output before using it, handle missing or malformed fields, and make consequential decisions depend on appropriate checks. A model may return valid-looking text that does not satisfy the underlying task.

For screenshot work, page rendering adds its own latency and failure modes. Waiting for network idle can be slower or unreliable on pages with ongoing requests; waiting for a selector or using a deliberate delay may better match the page. Cache behavior can reduce repeated capture work where freshness requirements allow it. ScreenshotNeo offers a configurable cache TTL, async jobs with signed webhooks, bulk capture for up to 100 URLs per call, and a usage API. Its plans are Free: 1,000 shots/month; Starter: $5 for 3,000; Growth: $15 for 15,000; Pro: $39 for 60,000; Scale: $99 for 250,000; and Business: $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.

12. Frequently asked questions

Is ChatGPT an LLM?

ChatGPT is a product that uses language models. “LLM” refers to a broad class of models, not to one specific application or model name.

Does an LLM search the internet for every answer?

Not necessarily. An LLM can generate from its prompt and learned parameters without browsing. Any external retrieval or tool use depends on the system connected to it.

Does adding more prompt detail change the model?

It changes the input context for that inference request. It does not, by itself, update the model’s trained parameters.

Can I treat generated code as correct because it compiles?

No. Compilation checks some properties of code, not whether it meets the intended behavior or handles important edge cases. Review and test it in the context where it will run.

Sources