How to Save Web Scraper Data to a File
Learn how to export scraped records to JSON, JSON Lines, CSV, or XML with Scrapy, Python, and Node.js, including append-safe workflows.

Saving scraper output is the step that turns temporary crawl results into data you can inspect, share, analyze, import, or send to another system. The exact command depends on your scraper, but the general process is consistent: define the fields your scraper yields, choose a file format that matches the next consumer, write the records, and verify the resulting file.
If you use Scrapy, the shortest solution is:
scrapy crawl myspider -O results.json
Uppercase -O creates a fresh file by overwriting an existing one. Lowercase -o appends to an existing feed. For append-heavy jobs, JSON Lines is usually safer than ordinary JSON because every record is independent:
scrapy crawl myspider -o results.jsonl
Choose the right output format
Your destination should determine the format. Scrapy supports JSON, JSON Lines, CSV, and XML feed serialization, and can infer the format from a filename extension.

| Format | Use it when | Important considerations |
|---|---|---|
| CSV | A spreadsheet, database import, or reporting tool expects rows and columns. | Choose a stable field list and order. Nested objects and arrays must be flattened or encoded. |
| JSON | You need structured records, nested fields, or broad application compatibility. | A consumer may need to load the complete document. Do not blindly append separate JSON documents. |
| JSON Lines (.jsonl) | You want incremental writes, appendable runs, or record-by-record processing. | Each line must contain one valid JSON value. The file is not one large JSON array. |
| XML | A downstream system specifically requires XML. | Use it for compatibility requirements rather than convenience. |
For large exports, JSON Lines is practical because a reader can process one record at a time instead of parsing an entire document. It also avoids the invalid-JSON problem that can occur when ordinary JSON output is appended repeatedly.
Save data with Scrapy
1. Find the spider name and yielded fields
List your spiders, then inspect the yield statements in the selected spider. The field names become the keys in JSON output and the columns in CSV output.
scrapy list
scrapy crawl myspider -O results.json
A minimal spider might yield records like this:
import scrapy
class ProductSpider(scrapy.Spider):
name = 'products'
start_urls = ['https://example.com/products']
def parse(self, response):
for card in response.css('.product'):
yield {
'name': card.css('.name::text').get(),
'price': card.css('.price::text').get(),
'url': response.urljoin(card.css('a::attr(href)').get()),
}
2. Export a new file
Use uppercase -O when the file should represent this run only:
scrapy crawl products -O products.json
scrapy crawl products -O products.csv
scrapy crawl products -O products.jsonl
scrapy crawl products -O products.xml
The extension selects the feed serializer. If you need explicit settings, configure the feed export in Scrapy settings instead of relying on the filename.
3. Append intentionally
Lowercase -o appends:
scrapy crawl products -o products.jsonl
Appending is appropriate for JSON Lines because each new record is another complete line. Appending separate JSON arrays or objects to a normal .json file can make it impossible for a standard parser to read the result. If you must maintain one JSON array, read the existing data, merge the new records, and write the complete array back, or use a database designed for incremental ingestion.
4. Keep CSV columns stable
CSV works best when every item has a predictable schema. Declare the fields and their order when downstream tools depend on exact columns. A varying set of item keys can produce missing values, unexpected columns, or a header that does not match later rows. Nested values such as lists should be flattened into separate columns or serialized as JSON text.
Save scraper data with Python
For a custom scraper that returns a list of dictionaries, Python’s standard library is enough. This example writes JSON, JSON Lines, and CSV versions of the same records:
import csv
import json
from pathlib import Path
records = [
{'name': 'Alpha', 'price': 12.50, 'tags': ['new', 'sale']},
{'name': 'Beta', 'price': 9.99, 'tags': ['popular']},
]
Path('results.json').write_text(
json.dumps(records, ensure_ascii=False, indent=2),
encoding='utf-8',
)
with Path('results.jsonl').open('w', encoding='utf-8') as file:
for record in records:
file.write(json.dumps(record, ensure_ascii=False) + '\n')
fields = ['name', 'price', 'tags']
with Path('results.csv').open('w', newline='', encoding='utf-8') as file:
writer = csv.DictWriter(file, fieldnames=fields)
writer.writeheader()
for record in records:
row = dict(record)
row['tags'] = json.dumps(row['tags'], ensure_ascii=False)
writer.writerow(row)
Use ensure_ascii=False when you want readable Unicode characters. Always specify an encoding, and use newline='' for CSV so the CSV module controls line endings correctly.
Appending JSON Lines safely
import json
new_records = [{'name': 'Gamma', 'price': 14.25}]
with open('results.jsonl', 'a', encoding='utf-8') as file:
for record in new_records:
file.write(json.dumps(record, ensure_ascii=False) + '\n')
If a process can stop halfway through a write, write to a temporary file and rename it after completion for replace-style exports. For append pipelines, include a run identifier or source URL so a later job can detect duplicates.
Save scraper data with Node.js
Node.js can serialize an array directly and can stream JSON Lines for larger jobs.
import { appendFile, writeFile } from 'node:fs/promises';
const records = [
{ name: 'Alpha', price: 12.50 },
{ name: 'Beta', price: 9.99 },
];
await writeFile('results.json', JSON.stringify(records, null, 2), 'utf8');
const lines = records.map((record) => JSON.stringify(record)).join('\n') + '\n';
await writeFile('results.jsonl', lines, 'utf8');
await appendFile('append-only.jsonl', lines, 'utf8');
For very large crawls, avoid keeping every record in one array. Write each record as it arrives to a JSON Lines stream, or use a database and export after the crawl.
Validate the file before using it
- Confirm the file exists and is not zero bytes.
- Count records and compare that count with the scraper’s item count.
- Open a sample from the beginning, middle, and end.
- Check Unicode, newlines, delimiters, null values, and numeric fields.
- Parse the output with the same kind of consumer that will read it in production.
For JSON, a quick validation command is:
python -m json.tool results.json > /dev/null
For JSON Lines, validate line by line:
python -c "import json; [json.loads(line) for line in open('results.jsonl', encoding='utf-8') if line.strip()]"
A parser error often identifies the first malformed line. Empty lines, truncated final records, and accidental logging output mixed into the file are common causes.
Common errors and fixes
The output file is empty
Your spider may not yield any items, selectors may no longer match, or the crawl may have been blocked. Check the crawl log, inspect the response status, and temporarily yield the page URL or raw selector result to confirm the parser sees the expected HTML.
JSON is invalid after multiple runs
You probably appended to ordinary JSON with -o. Use JSON Lines for append workflows, or regenerate one complete JSON document with uppercase -O.
CSV columns change between runs
Item dictionaries contain inconsistent keys. Define a fixed schema and field order, fill missing values explicitly, and serialize nested values before writing.
Non-ASCII text is escaped
Use UTF-8 and disable unnecessary ASCII escaping. In Python, set ensure_ascii=False; in Scrapy, verify the feed encoding settings and inspect the consumer’s expected encoding.
Duplicate records appear
Appending does not deduplicate. Store a stable key such as a canonical URL or source ID, then remove duplicates during ingestion or before export. Also check whether a retry or pagination loop is yielding the same item twice.
The file is too large to open
Use JSON Lines and process it in a stream. Split exports by date or crawl run, compress them after writing, or load them into a database or object store designed for large datasets.
Performance, reliability, and cost considerations
Serialization is usually cheaper than fetching pages, but memory use can become the limiting factor when you collect a whole crawl before writing. JSON Lines and streaming CSV keep memory bounded. Writing to a local disk is simple, while network filesystems and remote uploads introduce failure points; upload completed files atomically and retain a checksum or record count.
Choose overwrite or append based on recovery needs. Overwrite exports are easier to reproduce and validate. Append exports preserve earlier runs but require deduplication, schema discipline, and a clear policy for partial files. If a crawl is rerun after a failure, record the run ID and source URL so downstream systems can distinguish a retry from new data.
Scrapy’s feed exporter does not make the scraped data legally permissible to collect. Check the target site’s terms, privacy obligations, access controls, and applicable law before crawling or distributing the output.
Or skip the browser setup
If your workflow also needs screenshots of the pages you collect, ScreenshotNeo can return an image or PDF from one GET request. It removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options, including full-page capture, CSS selectors, device presets, custom headers and cookies, waits, blocking rules, caching, signed links, asynchronous jobs, bulk capture, and usage reporting.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 screenshots a month free with no card. Paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Should I use JSON or CSV?
Use JSON for nested or application-oriented records. Use CSV for rectangular data consumed by spreadsheets, imports, or reporting tools.
Is JSON Lines the same as JSON?
No. JSON Lines stores one JSON value per line, while ordinary JSON is commonly one array or object. JSON Lines is easier to append and stream.
Can I export directly from a hosted scraper?
Some hosted services provide dataset download endpoints in JSON, CSV, or JSON Lines. Follow that service’s API documentation; Scrapy command-line flags do not automatically apply to every hosted scraper.
How do I preserve nested fields in CSV?
Flatten them into columns or encode the nested value as a JSON string. Document the choice so consumers know how to decode it.
How do I make reruns safe?
Use a run identifier, stable record key, and deduplication step. Prefer writing a new file per run when reproducibility matters.


