ScreenshotNeo

BlogEngineering

AI Inference: Definition and How It Works

AI inference applies a trained model to new input. Learn the LLM request path, deployment modes, performance metrics, costs, and practical examples.

By the ScreenshotNeo team1 October 20267 min read

AI inference is the execution of a trained model on new input to produce an output. The output might be a class label, prediction, generated paragraph, image, embedding, or recommendation. Training changes the model’s parameters; inference uses those parameters.

For a large language model (LLM), inference turns a prompt into tokens, processes the prompt context, generates output tokens, and converts them back into readable text. The request may run in a cloud service, a data center, on-premises hardware, or an edge device.

What AI inference means

A model contains learned parameters produced during training. During inference, new data passes through the model’s computation, usually called a forward pass, and the model returns a result. NVIDIA’s definition of AI inference describes this as using a trained model to make predictions or generate outputs.

Lifecycle term What happens Typical input Typical output
Training Parameters are learned from examples. Large training dataset Trained model
Fine-tuning An existing model is adapted with specialized data. Task-specific examples Adapted model
Inference The trained or adapted model runs on new input. One request or batch Prediction, classification, or generation
Serving Infrastructure deploys and operates an endpoint that accepts inference requests. API or application request Managed inference response

Inference is therefore not synonymous with an API. An API can be part of serving, while inference is the model computation that produces the answer. Google Cloud’s lifecycle overview explains these distinctions and how deployment location affects the design.

How an LLM inference request works

  1. Request preparation: the application assembles the system instructions, conversation history, user prompt, tools, and generation settings.
  2. Tokenization: a tokenizer converts text into tokens. A token can represent a word, part of a word, punctuation, or another text unit; one token is not necessarily one word.
  3. Prefill: the model processes the prompt tokens and builds the context needed for generation. This phase often depends heavily on prompt length.
  4. Decode: the autoregressive model selects output tokens, commonly one at a time, using the context and decoding settings.
  5. Detokenization: generated tokens are converted to displayable text and returned, either all at once or as a stream.

This is a simplified text-generation path. Observed request time can also include authentication, queueing, safety checks, retrieval, tool calls, network transfer, and postprocessing. AWS describes these stages and common inference modes.

Why the first token can be slow

Time to first token (TTFT) includes request handling, queueing, and prompt prefill before any generated text is available. Once streaming starts, inter-token latency measures the gaps between subsequent tokens. A short answer can have an acceptable total time but a poor TTFT, while a long answer can have a good TTFT but take longer overall.

A runnable inference example in Python

The following standard-library script demonstrates inference with fixed, already-trained parameters. The weights are illustrative; the important distinction is that the script applies parameters to new text instead of updating them.

import math
import re

# Parameters produced earlier by a training process.
WEIGHTS = {
    "refund": 2.2,
    "charged": 1.8,
    "invoice": 1.4,
    "error": 0.9,
    "hello": -0.8,
    "thanks": -0.5,
}
BIAS = -0.4

def tokenize(text):
    return re.findall(r"[a-z]+", text.lower())

def infer(text):
    tokens = tokenize(text)
    score = BIAS + sum(WEIGHTS.get(token, 0.0) for token in tokens)
    probability = 1.0 / (1.0 + math.exp(-score))
    label = "support_request" if probability >= 0.5 else "other"
    return {"label": label, "probability": round(probability, 4), "tokens": tokens}

if __name__ == "__main__":
    print(infer("I was charged twice for my invoice"))

The function infer performs the forward computation and returns a prediction. It does not modify WEIGHTS or BIAS; changing those values would be a training or fine-tuning operation.

Inference modes

Mode How it works Good fit Main constraint
Batch Groups inputs and processes them on a schedule. Document classification, nightly scoring, offline analysis Results are delayed until the batch runs.
Real time Processes one request for an interactive response. Chat, search ranking, fraud decisions Latency and burst capacity matter.
Streaming Consumes an ongoing stream and emits results continuously. Events, telemetry, live transcription Backpressure, ordering, and partial results need handling.
Edge Runs near the user or data source. Intermittent connectivity, local control loops, low network travel Device memory, compute, power, and model size are limited.

Cloud, data-center, on-premises, and edge placement are architectural choices. Edge placement can reduce network travel, but it does not automatically improve privacy, speed, or cost; those outcomes depend on hardware, connectivity, data handling, and operations. See AWS’s edge inference overview.

What affects inference speed, capacity, and cost

  • Prompt length: longer context increases prefill work and memory use.
  • Generated length: more output tokens increase decode time and usage.
  • Concurrency: simultaneous requests compete for memory and compute.
  • Batching: improves utilization in some workloads but can increase waiting time for an individual request.
  • Model size and context capacity: larger models generally require more memory and may need partitioning or quantization.
  • Hardware and runtime: accelerators, kernels, precision, and scheduling affect throughput and latency.
  • Network and queueing: service overhead can dominate short requests.

Inference is often bursty, memory-sensitive, and latency-sensitive in ways that scheduled training is not. AWS documents these differences between training and inference.

Metrics to record

Metric Definition Report with
TTFT Time until the first generated token. Prompt length, queueing policy, streaming settings
Inter-token latency Gap between streamed output tokens. Output length and concurrency
End-to-end latency Total time from request start to completed response. Clear benchmark boundaries and percentiles
Throughput Tokens or requests processed per unit of time. Concurrency, model, prompt/output profile
Cost Resource or service cost for a defined workload. Quality target, retries, idle capacity, and request mix

Do not compare a throughput result from a large batch with an interactive latency result and call one universally faster. NVIDIA’s benchmarking guidance recommends stating the model, workload, hardware, software, concurrency, and measurement method.

Reliability and production design

  • Set request, queue, and total deadlines separately so a stuck dependency does not consume workers indefinitely.
  • Use bounded retries with backoff for transient transport failures; avoid blindly retrying non-idempotent actions or rate-limit responses.
  • Return partial streamed output deliberately. Decide whether clients can resume, discard, or persist an interrupted response.
  • Track timeouts, cancellations, queue depth, token counts, error rates, and model-version changes.
  • Validate input size and output limits before dispatch to protect memory and cost budgets.
  • Keep a fallback plan for model, region, or device failure, and verify that the fallback meets quality and data-location requirements.
  • Monitor data and concept drift. A model can keep running successfully while its predictions become less useful.

Common mistakes and troubleshooting

Symptom Likely cause Fix
High TTFT Long prompt, queueing, cold start, or overloaded runtime. Measure each phase, trim context, warm capacity, and control concurrency.
Output stops mid-response Output limit, client timeout, dropped stream, or server cancellation. Check finish reason and limits, increase read timeout, and handle reconnects explicitly.
Out-of-memory errors Model, context, batch, or KV cache exceeds available memory. Lower concurrency or context, use a smaller/quantized model, or add suitable memory.
Throughput falls as users increase Contention, excessive batching delay, or saturation. Load-test at the target concurrency and tune batching and admission control.
Predictions degrade after deployment Input distribution or label meaning changed. Compare production data with evaluation data and retrain or fine-tune only with evidence.
Edge device misses deadlines Insufficient compute, thermal throttling, power limits, or intermittent connectivity. Reduce model size, optimize the runtime, change the deadline, or move work to a nearby service.
Cost is unexpectedly high Large prompts, verbose outputs, retries, idle capacity, or unsuitable batch size. Track tokens and retries by route, enforce limits, and compare cost at the same quality target.

Or skip the browser setup

If you need screenshots of inference documentation, dashboards, or model result pages, ScreenshotNeo provides a single screenshot API request. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are never billed, and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

There is a free plan with 1,000 screenshots a month and no card. Paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

Frequently asked questions

Is inference the same as prediction?

Prediction is one possible inference output. Generative models produce text, images, audio, or other outputs through inference as well.

Does inference require a GPU?

No. GPUs and TPUs are common accelerators, but inference can run on CPUs or other devices when the model and latency target fit.

Is fine-tuning performed during inference?

No. Fine-tuning changes model parameters before deployment. Inference applies the resulting model to new inputs.

Why can two systems with the same model have different latency?

Prompt and output lengths, batching, concurrency, hardware, runtime configuration, queueing, and network paths can all differ.

When should inference run at the edge?

Consider edge placement when network travel or connectivity is a constraint and the device can meet the model’s memory, power, and quality requirements.