How to Pad a Dataset
Learn how to pad variable-length sequences and arrays for batching, choose safe fill values, preserve masks, and avoid wasted memory or silent truncation.

Padding a dataset means extending variable-length samples to a common length or shape by adding a fill value. It makes sequences and arrays stackable into batches. Padding does not add real observations, balance classes, or create new information.
The safest workflow is:
- Decide whether the target should be the longest item in each batch or a fixed maximum.
- Choose a fill value that your model can distinguish from real data, or keep a mask/length for every sample.
- Make an explicit policy for inputs longer than the target: reject, truncate, or increase the target.
- Apply the same policy to features, labels, attention masks, and any aligned metadata.
- Validate shapes, dtypes, padding direction, and a short and long example after transformation.
What padding changes—and what it does not
Suppose one sample has length 3 and another has length 5. A batch tensor needs a rectangular shape, so the shorter sample can become length 5 by appending two fill values. The original sequence length remains 3; the padded representation is simply easier for tensor libraries to store and process.
Padding is different from oversampling, augmentation, and class balancing. It does not create additional records. It also does not repair missing values or make a biased dataset representative.
Choose a padding strategy
| Strategy | Use it when | Main trade-off |
|---|---|---|
| Longest in each batch | Sequence lengths vary and your loader supports dynamic shapes | Less wasted memory, but batch shapes vary |
| Fixed maximum | The model, accelerator, export format, or serving API requires predictable shapes | Stable shapes, but more filler and a truncation decision |
| No padding | The downstream operation accepts lists or ragged tensors | Minimal filler, but fewer dense tensor operations |
Batch-longest padding
For each batch, find its longest sample and pad the other samples to that length. This usually reduces unnecessary computation compared with padding an entire dataset to a global outlier. Length bucketing—grouping similarly sized samples together—can reduce waste further.

Fixed maximum padding
Choose a documented maximum, such as 256 tokens or 2,000 time steps. Every sample shorter than the maximum is padded to that shape. Inputs longer than the maximum need a separate rule. Truncate only when losing the tail is acceptable; otherwise reject the item or choose a larger maximum.
Left or right padding
Right padding appends values after the real sequence and is common for audio, time series, and many token pipelines. Left padding puts fill values before the sequence and can be useful when the most recent tokens must align at the right edge. The model and tokenizer must agree on the direction.
Pad one-dimensional NumPy arrays in Python
The following helper right-pads an array and refuses to hide an overlong input. It uses a configurable fill value instead of assuming that zero is always correct.

import numpy as np
def right_pad_1d(values, target_length, fill_value=0.0):
values = np.asarray(values)
if values.ndim != 1:
raise ValueError("values must be one-dimensional")
if len(values) > target_length:
raise ValueError(
f"input length {len(values)} exceeds target {target_length}"
)
padding = target_length - len(values)
padded = np.pad(
values,
(0, padding),
mode="constant",
constant_values=fill_value,
)
mask = np.concatenate([
np.ones(len(values), dtype=np.int8),
np.zeros(padding, dtype=np.int8),
])
return padded, mask
x = np.array([0.4, 0.7, 0.2], dtype=np.float32)
padded, mask = right_pad_1d(x, target_length=5, fill_value=0.0)
print(padded) # [0.4 0.7 0.2 0. 0. ]
print(mask) # [1 1 1 0 0]
The mask has one entry per position: 1 means real data and 0 means padding. Keep the original length as well when it is useful for pooling, loss calculation, or restoring the unpadded output.
Pad a complete variable-length dataset
For a small in-memory dataset, compute a target and return both the dense matrix and lengths. Computing the target from the training partition avoids allowing a held-out sample to determine preprocessing for the training data.
import numpy as np
def pad_dataset(sequences, target_length=None, fill_value=0.0,
truncate=False):
arrays = [np.asarray(s) for s in sequences]
if not arrays:
raise ValueError("sequences cannot be empty")
if any(a.ndim != 1 for a in arrays):
raise ValueError("every sequence must be one-dimensional")
lengths = np.array([len(a) for a in arrays], dtype=np.int32)
target = int(lengths.max()) if target_length is None else int(target_length)
if target < 0:
raise ValueError("target_length must be non-negative")
output = []
masks = []
for sequence in arrays:
if len(sequence) > target:
if not truncate:
raise ValueError(
f"sequence length {len(sequence)} exceeds target {target}"
)
sequence = sequence[:target]
padded, mask = right_pad_1d(sequence, target, fill_value)
output.append(padded)
masks.append(mask)
return np.stack(output), np.stack(masks), lengths
samples = [
np.array([1.0, 2.0]),
np.array([3.0, 4.0, 5.0, 6.0]),
np.array([7.0]),
]
batch, mask, original_lengths = pad_dataset(samples)
print(batch.shape) # (3, 4)
print(original_lengths) # [2 4 1]
print(mask)
For production data, process batches incrementally rather than loading the entire dataset into memory. If you use a fixed maximum, record the value with the model configuration so training and inference use the same shape.
Token sequences and tokenizer settings
Tokenizers commonly expose three conceptual modes: pad to the longest sequence in the batch, pad to a specified maximum length, or leave sequences unpadded. Truncation is a separate setting. Configure both explicitly.
# Pseudocode: exact keyword names depend on the tokenizer library
encoded = tokenizer(
texts,
padding="longest", # or "max_length" / False
truncation=True,
max_length=256,
return_attention_mask=True,
)
Use the tokenizer’s configured pad-token ID. Do not assume integer zero is padding: zero may be a valid token or feature value. Confirm that the model has a pad token and that its attention-mask convention ignores padded positions.
Padding multidimensional arrays
Images, spectrograms, and spatial features may need padding on more than one axis. Define the target shape per axis and keep the channel or feature axis unchanged unless the model expects otherwise.
import numpy as np
def pad_matrix(matrix, target_rows, target_cols, fill_value=0.0):
matrix = np.asarray(matrix)
if matrix.ndim != 2:
raise ValueError("matrix must be two-dimensional")
rows, cols = matrix.shape
if rows > target_rows or cols > target_cols:
raise ValueError("target shape is smaller than the input")
return np.pad(
matrix,
((0, target_rows - rows), (0, target_cols - cols)),
mode="constant",
constant_values=fill_value,
)
When padding a feature array and a label array, pad them consistently. A sequence-labeling example needs a label mask as well as a feature mask; otherwise the loss may be calculated on artificial labels.
Framework batch APIs
Many dataset frameworks provide a padded-batch operation. MindSpore’s versioned API references use padded_batch and pad_info to describe padded shapes and values; unspecified shape entries can be padded to the largest sample shape. API names and defaults vary by release, so check the documentation for the version installed in your environment.
In PyTorch-style loaders, a custom collate_fn is often the right place to compute the longest length for the current batch, create a dense tensor, and return lengths or masks. Keep this operation deterministic so workers produce the same shape and fill policy.
Choosing a fill value
- Numeric signals: zero is convenient and is used by documented NumPy examples, but it may also be a real measurement.
- Tokens: use the tokenizer’s pad-token ID and its attention-mask rules.
- Floating-point features: a sentinel such as a value outside the valid range can work only if every downstream operation handles it safely.
- Images: use the same normalization space as the model; a black pixel is not automatically zero after normalization.
- Labels: use an ignore index supported by the loss, and mask those positions.
If no fill value is unambiguously distinct, preserving lengths and masks is mandatory. A mask is also safer than trying to infer padding later by looking for zeros.
Truncation, rejection, and outliers
Padding cannot solve an input longer than the target. Choose one policy:
- Reject: fail validation and send the sample to a review queue.
- Truncate: keep the first, last, or a task-specific window, and record that truncation occurred.
- Increase the target: preserve information while accepting more memory and compute.
- Chunk: split long sequences into windows, carrying overlap and labels carefully.
Never silently truncate. For time series, truncating the beginning and truncating the end have different meanings; for language, removing a middle span can destroy context. Write the rule into the dataset version and model configuration.
Validation checklist
- Print or assert the final batch shape and dtype.
- Check that every original length is less than or equal to the target.
- Inspect one shortest, median, and longest sample.
- Verify left/right padding direction.
- Verify masks contain real positions only where data exists.
- Confirm labels, timestamps, and features remain aligned.
- Measure the proportion of padded positions.
- Run one training and one inference batch through the full model.
- Serialize the target length, pad value, pad token, and truncation policy.
Performance, memory, and reliability
Padding increases tensor size, so the cost is proportional to the number of padded positions. Batch-longest padding and length bucketing reduce wasted work. A global maximum is simpler for compiled graphs and serving systems but can consume substantially more memory when a few outliers are long.
For distributed training, make sure every worker follows the same batching and padding policy. Dynamic shapes can increase compilation or kernel-selection overhead. Fixed shapes can improve repeatability but should be chosen from observed training data rather than an arbitrary extreme.
Keep masks with the batch rather than recreating them in later layers. Log counts of rejected and truncated samples. If padding-related failures appear intermittently, inspect batch composition, worker seeding, and whether one partition uses a different tokenizer or dtype.
Common errors and fixes
| Error | Cause | Fix |
|---|---|---|
| “all input arrays must have the same shape” | Stacking happened before padding | Pad inside the collator or preprocessing step, then stack |
| Index or shape mismatch | Features and labels received different padding | Use one target length and separate masks where needed |
| Model attends to filler tokens | Mask was omitted or has the wrong polarity | Return the framework’s expected attention mask and test it |
| Unexpected accuracy change | Zero or another fill value is a meaningful feature | Use a model-appropriate sentinel or an explicit mask |
| Samples lose their endings | Fixed maximum caused silent truncation | Reject, chunk, or document an explicit truncation rule |
| Out-of-memory errors | Target was set by a very long outlier | Bucket by length, cap the target, or process long items separately |
| Tokenizer reports no pad token | Model configuration lacks one | Configure a valid pad token and confirm the model supports it |
Or skip the browser setup
If your workflow needs screenshots of padded-data reports, validation dashboards, or generated documentation, ScreenshotNeo can return an image or PDF from one request. It accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page and element capture, dark mode, device presets, custom viewports, retina scale, PDF settings, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture, and a usage API. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
How do I pad sequences to the same length?
Choose a target length, append the configured pad value to shorter sequences, and retain lengths or a mask. Use the longest item in each batch when dynamic shapes are acceptable.
What value should I use for padding?
Use the tokenizer’s pad token for text. For numeric data, zero is one option, but only when zero cannot be confused with a real value or a mask is always applied.
Should I pad before or after splitting the dataset?
Split first, then derive training preprocessing parameters from the training partition. Apply that documented policy to validation and test data.
Does padding balance classes?
No. Padding changes representation shape. Use sampling, class weights, or another explicitly designed balancing method for class imbalance.
Can I avoid padding entirely?
Yes, if your framework supports ragged tensors, packed sequences, lists, or per-example execution. Dense accelerators and many batch APIs still require rectangular tensors.


