ScreenshotNeo

BlogHow-to

How to Scrape OpenSea Data With Python: NFT Metadata and Listings

Use OpenSea's authenticated API with Python to collect NFT metadata and listings safely, paginate results, handle rate limits, and monitor events.

By the ScreenshotNeo team1 October 20267 min read

Use OpenSea’s authenticated API rather than scraping its website with a browser. Create an API key, send it in the x-api-key header, call the documented metadata route for each NFT, use documented listing endpoints with cursor pagination, and handle rate limits from response headers. OpenSea’s API covers NFTs, tokens, marketplace data, collections, listings, offers and event streams. Review the OpenSea developer documentation and current Terms before running a large job.

What you can collect

  • NFT metadata: name, description, image, animation URL, external link and traits.
  • Marketplace data: current listings, offers and related events through documented API endpoints.
  • Live activity: listings, sales, transfers, metadata updates and cancellations through the Stream API.

Browser automation is less stable, exposes you to changing page markup and may violate OpenSea rules. OpenSea’s Terms prohibit unauthorized automated extraction, circumventing access controls or rate limits, sharing API keys or API data, and commercializing API data without express written permission. Preserve required OpenSea attribution and link back when displaying NFTs.

1. Create an API key and configure Python

  1. Create a key through OpenSea’s developer flow.
  2. Store it outside source control.
  3. Install the HTTP client: python -m pip install requests.
export OPENSEA_API_KEY='replace-with-your-key'
export OPENSEA_CHAIN='ethereum'
export OPENSEA_CONTRACT='0x0000000000000000000000000000000000000000'
export OPENSEA_TOKEN_ID='1'

Keys expire and limits can change. Never put the key in a browser bundle, notebook shared publicly, Git repository or client-side application.

2. Fetch NFT metadata

The documented metadata route is /api/v2/metadata/{chain}/{contractAddress}/{tokenId}. The following script normalizes nullable fields and expands traits into rows you can write to a database.

import os
import requests

API_KEY = os.environ['OPENSEA_API_KEY']
chain = os.environ.get('OPENSEA_CHAIN', 'ethereum')
contract = os.environ['OPENSEA_CONTRACT']
token_id = os.environ['OPENSEA_TOKEN_ID']

url = f'https://api.opensea.io/api/v2/metadata/{chain}/{contract}/{token_id}'
headers = {'x-api-key': API_KEY, 'Accept': 'application/json'}

with requests.Session() as session:
    response = session.get(url, headers=headers, timeout=30)
    response.raise_for_status()
    payload = response.json()

metadata = {
    'chain': chain,
    'contract': contract,
    'token_id': token_id,
    'name': payload.get('name'),
    'description': payload.get('description'),
    'image': payload.get('image'),
    'animation_url': payload.get('animation_url'),
    'external_url': payload.get('external_url'),
}
traits = []
for trait in payload.get('traits') or []:
    traits.append({
        'trait_type': trait.get('trait_type'),
        'value': trait.get('value'),
        'display_type': trait.get('display_type'),
        'max_value': trait.get('max_value'),
    })

print(metadata)
for trait in traits:
    print(trait)

Keep the original response as well as normalized columns. New nullable fields can appear, while trait arrays may contain different keys or values.

3. Fetch listings with cursor pagination

Use the documented collection or NFT listing endpoint that matches your scope. Endpoint paths and filters vary by the resource you need, so set the exact route from the current OpenSea reference documentation rather than guessing a URL. Most list responses return a result array and a cursor for the next batch.

import os
import time
import requests

API_KEY = os.environ['OPENSEA_API_KEY']
LISTINGS_URL = os.environ['OPENSEA_LISTINGS_URL']
headers = {'x-api-key': API_KEY, 'Accept': 'application/json'}

# Add only filters supported by the endpoint you selected.
params = {'limit': 50}
all_listings = []

with requests.Session() as session:
    while True:
        response = session.get(LISTINGS_URL, headers=headers, params=params, timeout=30)
        if response.status_code == 429:
            retry_after = int(response.headers.get('Retry-After', '5'))
            time.sleep(retry_after)
            continue
        response.raise_for_status()
        body = response.json()
        batch = body.get('listings') or body.get('orders') or body.get('results') or []
        all_listings.extend(batch)
        cursor = body.get('next') or body.get('next_cursor') or body.get('cursor')
        if not cursor:
            break
        params['cursor'] = cursor

print(f'Collected {len(all_listings)} listings')

Persist the cursor after each successful batch. If the process stops, restart from the checkpoint instead of repeating every request. Store a retrieval timestamp because a listing can expire or be fulfilled between pages.

4. A production-grade request helper

Different status codes require different actions:

Status Meaning Action
401 Missing, expired or invalid key Check the environment variable and create a current key.
403 Access or permission problem Check endpoint permissions and policy compliance; do not bypass controls.
404 Resource not found Validate chain, contract and token ID; distinguish this from an expired key.
429 Rate limit exceeded Wait for Retry-After, or until X-RateLimit-Reset, then retry.
5xx Transient server failure Use bounded exponential backoff with jitter.
import random
import time
import requests

RETRYABLE = {500, 502, 503, 504}

def get_json(session, url, headers, params=None, attempts=5):
    for attempt in range(attempts):
        response = session.get(url, headers=headers, params=params, timeout=30)
        if response.status_code == 429:
            retry_after = response.headers.get('Retry-After')
            if retry_after and retry_after.isdigit():
                delay = int(retry_after)
            else:
                delay = max(1, 2 ** attempt)
            time.sleep(delay)
            continue
        if response.status_code in RETRYABLE:
            time.sleep(min(30, 2 ** attempt + random.random()))
            continue
        response.raise_for_status()
        return response.json(), response.headers
    raise RuntimeError('request failed after bounded retries')

Read the X-RateLimit-* headers on every response. OpenSea documents an example free-tier key with 600 read requests per hour and 30 write requests per hour, but those values can change and example keys expire after seven days. Do not hard-code them.

5. Batch, cache and resume safely

  • Cache collection metadata and traits because they change less often than listings.
  • Use batch identifier requests where the selected endpoint supports them. Batching reduces request count but produces larger payloads and can make one bad identifier harder to isolate.
  • Request only needed fields and apply collection, token or event filters server-side.
  • Checkpoint cursors and write each successful page atomically.
  • Deduplicate using a stable event, order or listing identifier plus an observed timestamp.
  • Keep raw JSON for auditability and normalized tables for analytics.

6. Monitor live events with Stream API

Polling is suitable for periodic snapshots. For low-latency monitoring of listings, sales, transfers, metadata updates or cancellations, use the Stream API over WebSocket. Streamed events do not count toward API rate limits. Persist event IDs or timestamps, reconnect with backoff, and make event handling idempotent because reconnects can overlap deliveries.

Approach Best for Trade-offs
REST polling Repeatable snapshots and backfills Consumes request quota and can miss short-lived changes between polls.
Stream WebSocket Near-real-time activity More connection and replay logic; maintain checkpoints and reconnect handling.

7. cURL, Python and Node.js examples

cURL metadata request

curl -G 'https://api.opensea.io/api/v2/metadata/ethereum/0x0000000000000000000000000000000000000000/1' \
  -H 'x-api-key: YOUR_API_KEY' \
  -H 'Accept: application/json'

Python metadata request

import requests

r = requests.get(
    'https://api.opensea.io/api/v2/metadata/ethereum/0x0000000000000000000000000000000000000000/1',
    headers={'x-api-key': 'YOUR_API_KEY', 'Accept': 'application/json'},
    timeout=30,
)
r.raise_for_status()
print(r.json())

Node.js metadata request

const url = 'https://api.opensea.io/api/v2/metadata/ethereum/0x0000000000000000000000000000000000000000/1';
const res = await fetch(url, {
  headers: { 'x-api-key': process.env.OPENSEA_API_KEY, 'Accept': 'application/json' }
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(await res.json());

For listings, keep the same authentication and retry logic, add the endpoint’s documented filters, and send the returned cursor on the next request.

8. Performance, reliability and cost considerations

  • Performance: smaller filtered pages, cached metadata and bounded concurrency reduce latency and quota use.
  • Reliability: use timeouts, retries only for 429 and transient 5xx responses, cursor checkpoints and idempotent writes.
  • Quota: response headers are authoritative. A larger page can reduce request count but increases payload size and failure impact.
  • Data freshness: listings can change quickly; record retrieval times and use Stream API when polling intervals are insufficient.
  • Compliance: check current OpenSea Terms and developer policies before collecting or redistributing a large dataset. API-data commercialization may require express written permission.

9. Troubleshooting checklist

  • 401: confirm the header is exactly x-api-key, remove accidental whitespace and verify the key has not expired.
  • 403: verify the endpoint and permissions. Do not evade access controls.
  • 404 for a token: check chain spelling, checksum/contract address and token ID; confirm the asset exists.
  • 429 loops: honor Retry-After, reduce concurrency, cache stable data and inspect X-RateLimit-Reset.
  • Missing listings: distinguish an empty valid result from an expired key, authorization error, rate limit or server failure.
  • Duplicate rows: checkpoint after a committed page and upsert by a stable listing or event identifier.
  • WebSocket reconnect duplicates: persist event IDs or timestamps and make consumers idempotent.
  • Malformed metadata: retain raw JSON, treat fields as nullable and handle traits as an array rather than a fixed schema.

Or skip the browser setup

If your workflow also needs screenshots of OpenSea pages or NFT assets, ScreenshotNeo provides a single-request screenshot API and MCP server. Cookie banners, newsletter popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed. AI agents can use its MCP tools, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000 shots.

See the ScreenshotNeo API documentation for all options.

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://opensea.io -o opensea.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://opensea.io'}, timeout=90)
open('opensea.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://opensea.io' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Create a free ScreenshotNeo account with 1,000 screenshots per month and no card.

FAQ

Can I scrape OpenSea without an API key?

The documented API requires an API key. Browser scraping may also conflict with OpenSea’s Terms, so obtain authorization and use the official API.

How do I know whether a listing is still active?

Fetch current listings through the documented endpoint and record retrieval time. For continuous changes, consume Stream API events and reconcile with periodic snapshots.

Should I store image files or metadata URLs?

Store the metadata response and URLs you are permitted to retain. Follow the applicable asset, attribution and redistribution terms.

How often should I poll?

Choose an interval based on freshness requirements and the limits shown in response headers. Use caching and Stream API channels when polling would be wasteful.