ScreenshotNeo

BlogGuides

What Are Scrapy Items and Item Loaders, and How Do You Use Them?

Learn how Scrapy Items and Item Loaders structure, clean, validate, and export scraped data with complete examples and troubleshooting.

By the ScreenshotNeo team30 September 202610 min read

What Are Scrapy Items and Item Loaders, and How Do You Use Them?

Short answer: a Scrapy Item is the structured container that holds data extracted by a spider. An Item Loader is an optional collection and processing layer that gathers values from XPath, CSS selectors, or direct values, applies input processors as values arrive, and applies output processors when you call load_item(). The resulting item can then pass through item pipelines and exporters.

Scrapy does not require an Item Loader. For a small spider, assigning fields directly can be clearer. Loaders become useful when several selectors feed one field, when values need consistent normalization, or when different spiders should share parsing rules.

1. The data flow: response to exported record

A typical Scrapy data flow has four stages:

  1. Spider extraction: a callback reads a Response and finds raw values.
  2. Item construction: the spider populates a dictionary, scrapy.Item, dataclass, attrs object, or Pydantic model.
  3. Pipeline processing: enabled pipeline components validate, deduplicate, enrich, store, or drop items.
  4. Export: an exporter serializes items to JSON, CSV, feeds, or another storage format.

Scrapy’s documentation summarizes the distinction precisely: items provide the container of scraped data, while Item Loaders provide the mechanism for populating that container. See the official Items documentation and Item Loaders documentation.

2. What is a Scrapy Item?

An item represents one scraped record, such as a product, article, job, or property. The current Scrapy ecosystem supports several representations through itemadapter:

Representation Strength Trade-off
Dictionary Minimal and flexible No declared schema; typos create unexpected keys
scrapy.Item Declared fields and field metadata More framework-specific
Dataclass Readable Python type declarations and defaults Annotations alone do not enforce runtime types
attrs class Explicit attributes and validation options Requires the attrs package and conventions
Pydantic model Runtime validation and coercion Additional dependency and model rules

Using scrapy.Item

import scrapy


class Product(scrapy.Item):
    name = scrapy.Field()
    price = scrapy.Field()
    url = scrapy.Field()
    in_stock = scrapy.Field()

Declaring fields catches assignments to undefined names and gives you a place to store metadata such as processors. A plain dictionary is appropriate when the schema is deliberately loose or generated dynamically.

Using a dataclass

from dataclasses import dataclass
from typing import Optional


@dataclass
class Product:
    name: Optional[str] = None
    price: Optional[str] = None
    url: Optional[str] = None
    in_stock: Optional[bool] = None

Defaults matter when a loader fills fields incrementally. A dataclass whose constructor requires every field can be awkward because the loader usually starts with only a partially known record. Optional fields with defaults, or a pre-populated instance, solve that construction problem.

Using Pydantic

from pydantic import BaseModel


class Product(BaseModel):
    name: str
    price: float
    url: str
    in_stock: bool = True

Pydantic performs runtime validation. Dataclass annotations by themselves are documentation and editor support; they do not automatically reject a string in a field annotated as float.

3. Direct assignment without an Item Loader

For a single selector per field, direct assignment is often the simplest approach:

import scrapy
from myproject.items import Product


class ProductSpider(scrapy.Spider):
    name = "products"
    start_urls = ["https://example.com/products"]

    def parse(self, response):
        for card in response.css("article.product"):
            item = Product()
            item["name"] = card.css("h2::text").get(default="").strip()
            item["price"] = card.css(".price::text").get(default="").strip()
            item["url"] = response.urljoin(card.css("a::attr(href)").get())
            yield item

This style keeps extraction visible and has little machinery. It becomes repetitive when every field needs trimming, type conversion, fallback selectors, or aggregation from multiple nodes.

4. What an Item Loader adds

An ItemLoader stores values internally by field. Each call to add_xpath(), add_css(), or add_value() contributes values to that field. Input processors run immediately for each addition. The processed values accumulate as a list. When load_item() runs, the output processor receives the accumulated values and returns the final field value.

An Item Loader gathers raw selector values, processes them, and produces one structured item.
An Item Loader gathers raw selector values, processes them, and produces one structured item.
from scrapy.loader import ItemLoader
from myproject.items import Product


def parse(self, response):
    loader = ItemLoader(item=Product(), response=response)
    loader.add_xpath("name", '//h1[@class="product-name"]/text()')
    loader.add_css("stock", "p.stock")
    loader.add_value("last_updated", "today")
    return loader.load_item()

The same field can receive values from multiple sources. This is useful for a title split across several text nodes, a price with a separate currency node, or fallback selectors for two page templates.

5. Input and output processors

Input processing

Input processors clean or convert one value at a time as it enters the loader. Use them for whitespace removal, case normalization, numeric conversion, URL cleanup, or dropping empty strings.

Output processing

Output processors decide what to do with the accumulated values when the item is loaded. Use them to select one value, join fragments, return a list, or perform a final conversion.

from itemloaders.processors import Join, MapCompose, TakeFirst
from scrapy.loader import ItemLoader


def strip_text(value):
    return value.strip()


class ProductLoader(ItemLoader):
    default_output_processor = TakeFirst()
    name_in = MapCompose(strip_text)
    price_in = MapCompose(strip_text)
    description_out = Join(" ")

TakeFirst is convenient for scalar fields, but it discards later values. Do not use it for tags, image URLs, or any field that must retain multiple results. For those fields, leave the output as a list or use a processor that removes duplicates while preserving order.

Processor precedence

Scrapy resolves processors from strongest to weakest precedence:

  1. Loader class attributes such as name_in and name_out.
  2. Field metadata such as input_processor and output_processor.
  3. Loader-wide defaults: default_input_processor and default_output_processor.

Processors receive an iterable as their first argument. A direct scalar passed to add_value() is treated as a one-value input, so the same processor rules apply to selector results and manually supplied values.

Field metadata processors

import scrapy
from itemloaders.processors import MapCompose, TakeFirst


def clean(value):
    return value.strip()


class Product(scrapy.Item):
    name = scrapy.Field(
        input_processor=MapCompose(clean),
        output_processor=TakeFirst(),
    )

Use field metadata when the rule belongs to one item schema. Use loader attributes when two loaders need different behavior for the same item. Use defaults for broad conventions shared by most fields.

6. A complete spider using an Item Loader

import scrapy
from itemloaders.processors import Join, MapCompose, TakeFirst
from scrapy.loader import ItemLoader


class Product(scrapy.Item):
    name = scrapy.Field()
    price = scrapy.Field()
    description = scrapy.Field()
    tags = scrapy.Field()
    url = scrapy.Field()


def clean_text(value):
    return " ".join(value.split())


def clean_price(value):
    return value.replace("$", "").replace(",", "").strip()


class ProductLoader(ItemLoader):
    default_output_processor = TakeFirst()
    name_in = MapCompose(clean_text)
    price_in = MapCompose(clean_price)
    description_in = MapCompose(clean_text)
    description_out = Join(" ")
    tags_in = MapCompose(clean_text)


class ProductsSpider(scrapy.Spider):
    name = "products"
    start_urls = ["https://example.com/products"]

    def parse(self, response):
        for card in response.css("article.product"):
            loader = ProductLoader(item=Product(), selector=card)
            loader.add_css("name", "h2::text")
            loader.add_css("price", ".price::text")
            loader.add_css("description", ".description ::text")
            loader.add_css("tags", ".tag::text")
            loader.add_value("url", response.urljoin(card.css("a::attr(href)").get()))
            yield loader.load_item()

Run it with:

Scoping a loader to one card prevents values from neighboring records from being mixed.
Scoping a loader to one card prevents values from neighboring records from being mixed.
scrapy crawl products -O products.json

The loader receives a card selector, so all CSS queries are scoped to one product. Scoping prevents selectors from accidentally collecting values from neighboring cards.

7. Handling missing, repeated, and malformed values

  • Missing selector: add_css() contributes no values. Decide whether the output should be an empty list, None, or a default.
  • Whitespace-only text: clean it in an input processor and return nothing for empty results if the field is optional.
  • Repeated values: keep a list for tags and images; use TakeFirst only for fields that are truly scalar.
  • Multiple templates: add a primary selector and a fallback selector, then choose or merge values in the output processor.
  • Locale-specific prices: preserve the original string when currency rules are ambiguous, or pass locale context to a processor.
  • Relative URLs: call response.urljoin() before adding the value.
  • JavaScript-rendered content: an ordinary Scrapy response contains server HTML only. Use an appropriate rendering integration or an upstream capture service when the required content is generated in a browser.

8. Item Loaders versus item pipelines

These components operate at different stages. A loader is local to extraction: it gathers and normalizes values while a callback is processing one response. A pipeline receives the completed item after the spider yields it.

from scrapy.exceptions import DropItem


class RequiredNamePipeline:
    def process_item(self, item, spider):
        if not item.get("name"):
            raise DropItem("missing name")
        return item

Use pipelines for cross-item work such as deduplication, database writes, validation that depends on several fields, or sending records to an external queue. Keep selector-specific parsing in the loader or spider.

Scrapy runs pipeline components sequentially. A component must return the item to continue processing or raise DropItem to stop that item. See the official pipeline documentation.

9. Exporting items

Feed exporters serialize the item representation to formats such as JSON and CSV. Scrapy’s item adapter lets exporters and pipelines handle supported item types consistently. Field values are normally passed to the underlying serializer; custom field serialization can be added when a target format needs special handling. See Item Exporters.

scrapy crawl products -O products.json
scrapy crawl products -O products.csv

Before exporting, inspect the final shape. A field intended as a scalar should not accidentally remain a one-element list, and a multi-value field should not be collapsed by TakeFirst.

10. Troubleshooting common errors

Symptom Likely cause Fix
AttributeError: ... add_css The loader variable is actually an item or response Instantiate ItemLoader(item=..., response=...) and call methods on that object
Field is always empty Selector does not match the response HTML Print response.text, verify the selector in a browser, and check whether content is client-rendered
Only the first tag is exported TakeFirst is the default output processor Override that field’s output processor so it returns the complete list
Numbers contain currency symbols No input normalization Use MapCompose to strip symbols and convert after handling locale rules
Dataclass construction fails Required fields have no defaults Give incrementally loaded fields defaults or create a complete instance before loading
Processor receives unexpected type Processor assumes a scalar instead of an iterable Write processors with an iterable first argument; use MapCompose for per-value functions
Pipeline never sees an item The spider returned a loader instead of load_item(), or a prior pipeline dropped it Yield the completed item and inspect logs for DropItem

11. Performance, reliability, and cost considerations

Item Loaders add small Python-level processing overhead. Network latency, response size, rendering, and database operations usually dominate crawl time. Keep processors deterministic and inexpensive, avoid repeated regular-expression compilation, and do expensive enrichment in a pipeline or a separate batch job.

For reliability, make processors tolerant of missing values, keep schema decisions explicit, and log the URL and field when validation fails. Preserve raw source values when a conversion could lose information. Test loaders with representative HTML fixtures for every page template, including empty fields and duplicate nodes.

Scrapy itself is open-source software, but a crawl can still incur infrastructure costs: proxy traffic, browser rendering, storage, database writes, and retries. Measure response volume and downstream work rather than assuming that parsing is the main expense.

12. Or skip the browser setup

If your Scrapy workflow needs screenshots of rendered pages for audits, visual regression, documentation, or AI-assisted extraction, ScreenshotNeo provides a website screenshot API and MCP server. It can capture PNG, JPEG, WebP, or PDF from one GET request. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options, including full-page capture, element selectors, dark mode, device presets, custom CSS and JavaScript, cookies, headers, blocking rules, waits, caching, asynchronous jobs, signed links, bulk capture, and usage reporting.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" \
  -d access_key=YOUR_API_KEY \
  --data-urlencode url=https://example.com \
  -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', data);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Create a free ScreenshotNeo account.

13. FAQ

Do I have to use scrapy.Item?

No. Dictionaries, dataclasses, attrs objects, and Pydantic models are supported through the item adapter.

Can I use an Item Loader without a custom item class?

Yes. Pass a dictionary or another supported object to ItemLoader.

When should I use add_value()?

Use it for constants, values computed from the response, normalized URLs, timestamps, or data supplied by spider logic rather than a selector.

Why does my input processor not join text?

Input processors handle values as they arrive. Joining accumulated fragments is an output concern, so use an output processor such as Join(" ").

Where should database code go?

Usually in an item pipeline. The loader should focus on extracting and normalizing one item from a response.

Are dataclass annotations runtime validation?

No. Use explicit validation or a Pydantic model when incorrect runtime types must be rejected.

14. Practical checklist

  • Choose an item representation that matches your schema and validation needs.
  • Use direct assignment for simple extraction and a loader for reusable normalization or aggregation.
  • Put per-value cleanup in input processors.
  • Put selection, joining, and list shaping in output processors.
  • Remember the precedence order: loader attributes, field metadata, then loader defaults.
  • Do not apply TakeFirst to fields that must retain multiple values.
  • Use pipelines for validation, deduplication, persistence, and cross-item work.
  • Test selectors and processors against missing, repeated, and template-specific content.
  • Inspect the final exported shape before sending data downstream.