ScreenshotNeo

BlogHow-to

How to Scrape Instagram in 2026

Learn the compliant way to collect Instagram data in 2026 with Meta’s APIs, OAuth, pagination, privacy controls, and practical code examples.

By the ScreenshotNeo team1 October 20269 min read

Use Meta’s authenticated Instagram APIs with OAuth and approved permissions. In 2026, this is the supportable way to collect Instagram media, comments, mentions, hashtagged media, and approved professional-account metadata. The documented APIs target Instagram Professional accounts (Businesses and Creators); consumer-account access is not supported by the Facebook-Login API documentation.

Browser automation and HTML scraping can violate Meta’s terms, expose passwords and tokens, trigger enforcement, and produce unreliable data. Public visibility does not automatically grant permission to copy, sell, profile, or republish data.

What you can collect through the official API

Meta’s Instagram API documentation covers these broad operations. Exact endpoint names, permissions, fields, and limits change, so confirm them in the live documentation before deploying.

Use case Supported scope Important limitation
Media retrieval Media published by connected professional accounts Requires an authorized account and approved permissions
Publishing Create and publish media for eligible professional accounts Publishing permissions and review requirements apply
Comments Read, manage, and reply to comments where permitted Access is permissioned and account-specific
Mentions Discover mentions of connected accounts Use the documented mention scope and fields
Hashtagged media Find media associated with approved hashtag searches Do not assume unrestricted historical or consumer-feed access
Professional-account information Basic metadata and metrics for other Businesses and Creators Fields and visibility depend on current API rules

The Facebook-Login route generally requires the Instagram Professional account to be linked to a Facebook Page. AWS’s connector documentation independently describes the OAuth 2.0, Meta developer account, Business app, permission, and Page-linking prerequisites.

Set up an authorized collection workflow

  1. Define the purpose and minimum fields. Write down whether you need media IDs, timestamps, captions, comments, hashtag results, metrics, or another field. Request only what the use case needs.
  2. Create a Meta developer app. Use the Instagram API flow documented for your account type in the Instagram API documentation.
  3. Connect an eligible account. Confirm that the account is a Business or Creator account and complete any required Facebook Page connection.
  4. Implement OAuth. Redirect the account owner to Meta’s authorization page, exchange the authorization result for a user access token, and store it in a secrets manager.
  5. Request least-privilege permissions. Some permissions require App Review or advanced access. Do not request broad scopes “just in case.”
  6. Resolve the account identifier. Use the current documentation to obtain the professional Instagram account ID associated with the authorized user.
  7. Read data with cursor pagination. Persist the returned cursor and request the next page until no cursor remains. Do not rely on page numbers or stable ordering.
  8. Apply retention and deletion controls. Record why each field is stored, how long it is needed, and how a deletion request removes it from caches, queues, indexes, and backups.
  9. Monitor permissions and errors. Meta can change fields, scopes, limits, and review requirements. Revalidate your integration after platform changes.

OAuth and token handling

Keep access tokens server-side. Never ask a user to send an Instagram password to your application, and never place a token in browser JavaScript, a public repository, logs, screenshots, or URLs that third parties can see. Meta warns people not to provide Facebook or Instagram passwords outside official sites, apps, or authorized Login with Facebook flows.

Token checklist

  • Encrypt tokens at rest and restrict access by service role.
  • Use short-lived credentials where the chosen flow supports them, then exchange or refresh according to Meta’s current documentation.
  • Store token owner, granted scopes, issued time, expiry, and revocation state.
  • Redact authorization headers and query strings from logs.
  • Delete tokens and derived data when access is revoked or a valid deletion request arrives.
  • Run a scheduled permission check so failures are detected before a collection job silently returns partial data.

Runnable collection examples

The examples below show the request pattern for a connected professional account. Replace the placeholders with values from your app and verify the current endpoint, fields, permissions, and version in Meta’s documentation before production use.

cURL

curl --get "https://graph.facebook.com/vXX.X/IG_USER_ID/media" \
  --data-urlencode "fields=id,caption,media_type,media_url,permalink,timestamp" \
  --data-urlencode "access_token=USER_ACCESS_TOKEN"

Python

import os
import requests

API_VERSION = "vXX.X"
ig_user_id = os.environ["IG_USER_ID"]
access_token = os.environ["INSTAGRAM_ACCESS_TOKEN"]

url = f"https://graph.facebook.com/{API_VERSION}/{ig_user_id}/media"
params = {
    "fields": "id,caption,media_type,media_url,permalink,timestamp",
    "access_token": access_token,
    "limit": 25,
}

while True:
    response = requests.get(url, params=params, timeout=30)
    response.raise_for_status()
    payload = response.json()

    for media in payload.get("data", []):
        print(media)

    next_url = payload.get("paging", {}).get("next")
    if not next_url:
        break
    url = next_url
    params = {}

Node.js

const version = 'vXX.X';
const igUserId = process.env.IG_USER_ID;
const accessToken = process.env.INSTAGRAM_ACCESS_TOKEN;

let url = new URL(`https://graph.facebook.com/${version}/${igUserId}/media`);
url.searchParams.set('fields', 'id,caption,media_type,media_url,permalink,timestamp');
url.searchParams.set('access_token', accessToken);
url.searchParams.set('limit', '25');

while (url) {
  const response = await fetch(url);
  if (!response.ok) {
    throw new Error(`Meta API error: ${response.status} ${await response.text()}`);
  }

  const payload = await response.json();
  for (const media of payload.data ?? []) console.log(media);
  url = payload.paging?.next ? new URL(payload.paging.next) : null;
}

These snippets intentionally use vXX.X. Pin a version supported by your app, then update it on a planned schedule. Treat response fields as optional: a field can be absent because it is not permitted, unavailable for that object, or removed in a newer version.

Pagination, ordering, and incremental collection

Instagram API collections use cursor-based pagination. Save the cursor or next URL exactly as returned, and stop when the response has no next page. Do not synthesize cursors, assume a fixed page size, or use offsets for deduplication.

  • Use the media ID as the primary deduplication key.
  • Keep the source account ID and retrieval timestamp beside every record.
  • For incremental jobs, store the newest successfully processed timestamp and overlap the next query window to handle late-arriving results.
  • Do not promise a globally ordered feed. The documentation does not support arbitrary ordering guarantees.
  • User Insights uses time-based pagination; follow the endpoint-specific instructions.

Hashtags, comments, mentions, and metrics

Each data type has separate permissions and eligibility rules. Design separate jobs instead of assuming one token grants every object.

Hashtagged media

Use the documented hashtag search flow and request only the fields your approved use case needs. Expect filtering, unavailable fields, and changing historical coverage. A hashtag result is not permission to copy every author’s content into a public database.

Comments

Store the comment ID, parent relationship, author information allowed by the response, timestamps, and moderation state only when required. Build idempotent reply and deletion operations so retries do not duplicate actions.

Mentions

Use mention discovery for the connected professional account. Do not infer that every mention, tag, or private reference is available through the API.

Metrics and insights

Metrics vary by account type, media type, permission, and time window. Persist the metric name, period, API version, and retrieval time so later reports remain interpretable.

Meta defines scraping as automated collection of data from a website or related interface. It distinguishes authorized crawling from unauthorized scraping that violates its terms. The Meta Platform Terms, captured as updated February 3, 2026, require compliance with applicable terms, documentation, and law.

The terms prohibit activities including selling, licensing, or purchasing Platform Data; processing Platform Data without valid user consent to build or augment user profiles; and processing data outside permitted purposes. Meta can suspend or remove apps, revoke API access, require deletion of Platform Data, and take other enforcement action.

Before collecting data, document:

  • The lawful purpose and legal basis for processing.
  • Which account owner authorized access.
  • Which fields are collected and why.
  • Retention, deletion, and user-access procedures.
  • Security controls for tokens and stored data.
  • How you respond to Meta policy or permission changes.

If the requirement is consumer-account access, broad public-web collection, or data about people who have not authorized your app, pause for jurisdiction-specific privacy, contract, and copyright review. Do not assume that “public” means unrestricted commercial reuse.

Why browser scraping is fragile

Method Authorization Coverage Reliability and risk
Official Instagram API OAuth and approved permissions Professional accounts and documented objects Supportable, versioned, permission-limited
Browser automation Often unclear or dependent on logged-in sessions May appear broader, but access is unstable Higher enforcement, CAPTCHA, account-lock, and maintenance risk
HTML scraping Usually no platform authorization Whatever a page happens to expose Breaks when markup, consent flows, or access rules change

Do not build systems that bypass login, CAPTCHAs, rate limits, robots controls, or Meta enforcement. Do not disguise automation as ordinary use.

Performance, reliability, and cost planning

  • Request only needed fields. Smaller responses reduce transfer time and parsing work.
  • Use bounded concurrency. Parallelize independent account jobs carefully and back off on transient errors.
  • Retry safely. Retry network failures and documented transient errors with exponential backoff and jitter. Do not blindly retry permission or validation errors.
  • Make writes idempotent. Key records and actions by Meta IDs so a retry cannot duplicate data or replies.
  • Cache by object ID and version. Refresh mutable fields on a schedule and retain the retrieval timestamp.
  • Budget for review and maintenance. App Review, permission changes, version upgrades, storage, and monitoring are part of the operating cost.
  • Do not invent a rate limit. Meta limits vary by endpoint and app. Check the live endpoint documentation and response headers before setting worker concurrency.

Troubleshooting

Symptom Likely cause Fix
OAuth succeeds but the account is missing Account is consumer, not Professional, or Page linking is incomplete Confirm Business or Creator status and complete the documented Page connection.
Permission or OAuth error Scope was not granted, reviewed, or enabled for the app Remove unnecessary scopes, request the required review, and reauthorize after approval.
Field is absent Field is unsupported for that object, account, version, or permission Check the current field reference and handle missing fields as normal.
Only the first page is returned Pagination cursor was ignored Follow paging.next until it is absent; persist progress after each page.
Duplicate records appear Retries or overlapping windows were not deduplicated Use the media, comment, or mention ID as an idempotency key.
Requests suddenly fail Token revoked, expired, permission changed, or API version changed Inspect the structured error, reauthorize when appropriate, and verify the version and scopes.
Collection is slow or incomplete Too much concurrency, transient throttling, or large field sets Reduce concurrency, add jittered backoff, request fewer fields, and record partial-job state.
A browser scraper hits a login wall or CAPTCHA Automated access was challenged Stop attempting to bypass it. Use the authorized API or obtain legal and platform approval for another method.

Or skip the browser setup

If your task is to capture a visual record of a permitted Instagram page or report, ScreenshotNeo provides a single screenshot request without maintaining browser automation. It removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.

See the ScreenshotNeo API documentation for options such as full-page capture, element selectors, device presets, custom headers and cookies, waits, blocking rules, caching, PDFs, signed links, async jobs, and bulk capture.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.instagram.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.instagram.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.instagram.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo has 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I scrape Instagram without getting banned?

No method can guarantee that. The lowest-risk approach is the documented API with OAuth, approved permissions, least privilege, secure tokens, and compliance with Meta’s terms.

Can I collect consumer profiles?

The Facebook-Login API documentation does not support consumer-account access. Do not infer access from what a logged-out browser can display.

It depends on jurisdiction, purpose, data, consent, contracts, copyright, and Meta’s terms. Obtain legal review for consumer-account or broad public-web collection.

Does the API provide unlimited historical data?

No. Coverage, ordering, retention, fields, and pagination depend on the endpoint and current platform rules.

Should I give a scraper my Instagram password?

No. Use Meta’s official OAuth flow and keep access tokens in server-side secret storage.