The Best Instagram Scrapers: Tools to Collect Public Data
Compare Instagram scraper APIs, hosted Actors, and Meta’s official API for collecting public profiles, posts, comments, and metadata.
Short answer: Bright Data is the strongest documented fit for production-scale structured collection. Apify is a practical hosted workflow for public profiles and posts when you want a serverless Actor and quick exports. Meta’s official Instagram Graph API is the right first-party route for Instagram Professional accounts that you manage.
Choose based on access model, fields, delivery format, scale, and compliance. Public visibility does not automatically grant permission to collect or reuse data. Review Instagram and Meta terms, applicable privacy and data-protection law, purpose limitation, retention, and whether the account or data is yours to process.
What an Instagram scraper can collect
Hosted scrapers and Actors can expose different combinations of public profile, post, reel, comment, and engagement fields. Confirm the current schema before building a data pipeline.
| Data | Typical fields | Questions to verify |
|---|---|---|
| Profiles | Username, profile URL, user ID, biography, follower and following counts, post count, verification status, external link | Are private accounts excluded? Are counts captured at request time? |
| Posts | Post URL, author, description, hashtags, comment count, date posted, likes, photos or media URLs | Are carousels and deleted posts handled? |
| Reels | Reel URL, author, description, hashtags, engagement fields, media URLs | Does the tool expose reels separately from posts? |
| Comments | Comment text, author, timestamp, likes, replies | What pagination and depth limits apply? |
| Hashtags and discovery | Posts or profiles associated with a public hashtag | Are results ranked, sampled, or complete? |
Best Instagram scrapers compared
| Tool | Best for | Access and setup | Delivery and operations |
|---|---|---|---|
| ScreenshotNeo | Capturing clean visual evidence of public pages | One HTTP request; removes consent banners, newsletter popups, and chat widgets before capture | PNG, JPEG, WebP, or PDF; clean shots only are billed; MCP tools for AI agents |
| Bright Data Instagram Scraper API | Production-scale structured profile, post, reel, and comment collection | Hosted API with documented scraper types | JSON, NDJSON, CSV, recurring jobs, webhooks, cloud storage, and up to 5,000 target URLs per API call are documented on its product page |
| Apify Instagram Profile & Posts Scraper | Hosted or serverless exports of public profiles and posts | The listed Actor says no Instagram login, password, cookies, API key, app registration, or access token is required | JSON, CSV, or Excel exports; recheck current maintenance, limits, pricing, and terms |
| Meta Instagram Graph API | First-party access for Instagram Professional accounts you manage | Facebook App plus user access token; intended for Businesses and Creators | Platform-supported permissions and fields; not a general public-profile discovery scraper |
1. Bright Data Instagram Scraper API
Bright Data documents separate scrapers for accounts, posts, reels, and comments. Its profile fields include account name, profile URL, user ID, biography, follower and following counts, post count, verification status, and an external link. Its post scraper lists URL, author, description, hashtags, comment count, date posted, likes, and photos.
The documented workflow supports structured JSON, NDJSON, or CSV output, recurring jobs, webhooks, cloud-storage delivery, and Python or Node.js SDKs or HTTP clients. The product page also advertises 5,000 free records per month for new accounts and up to 5,000 target URLs in one API call. These allowances and capabilities can change, so verify them before publication or procurement.
When to choose it
- You need several scraper types rather than only profiles and posts.
- You need scheduled collection, webhooks, or cloud delivery.
- You have a large URL list and need batch submission.
- Your downstream system accepts JSON, NDJSON, or CSV.
Bright Data integration pattern
Use the provider’s current endpoint, authentication method, and dataset schema from its documentation. Keep provider-specific code behind an adapter so a schema or endpoint change does not affect the rest of your pipeline.
# Save the provider response as export.json, then normalize selected fields.
python - <<'PY'
import csv, json
from pathlib import Path
rows = json.loads(Path("export.json").read_text())
if isinstance(rows, dict):
rows = rows.get("data", [rows])
fields = ["username", "profile_url", "followers", "following", "post_count", "verified"]
with Path("profiles.csv").open("w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=fields, extrasaction="ignore")
writer.writeheader()
writer.writerows(rows)
PY
2. Apify Instagram Profile & Posts Scraper
Apify’s listed Instagram Profile & Posts Scraper is designed around public Instagram web-profile data. Its description says an Instagram account, password, cookies, API key, app registration, or access token is not required. Results can be fetched as JSON, CSV, or Excel.
When to choose it
- You want a hosted Actor or serverless workflow.
- You need a quick export for a finite list of public profiles.
- Your team prefers downloading JSON, CSV, or Excel before adding a permanent data warehouse.
Check the exact Actor’s current maintenance status, pricing, limits, output schema, and terms. Hosted Actors can change behavior as Instagram’s public web pages change.
Normalize an Apify export
import json
from pathlib import Path
items = json.loads(Path("apify-export.json").read_text())
for item in items:
print({
"username": item.get("username"),
"url": item.get("url") or item.get("profileUrl"),
"followers": item.get("followers") or item.get("followersCount"),
"posts": item.get("posts") or item.get("postsCount"),
})
3. Meta Instagram Graph API
Meta’s official Instagram API collection targets Instagram Professionals: Businesses and Creators. Setup requires a Facebook App and a user access token. This is the clearest choice when first-party permissions and platform-supported access matter more than broad public discovery.
Use the Graph API when
- You manage the Instagram account or have the required authorization.
- You need a supported integration for your own business or creator account.
- Your compliance review prefers a first-party permission model.
Do not treat the Graph API as a drop-in replacement for a public-profile scraper. Its account type, permissions, fields, and review requirements are different.
How to choose: a practical decision process
- Define ownership and purpose. Write down whose data you are collecting, why, how long you will retain it, and who can access it.
- List required fields. Separate profiles, posts, reels, comments, hashtags, media URLs, and engagement counts.
- Pick the access model. Use Meta for authorized Professional accounts; use a hosted API or Actor for public-web collection where its terms allow it.
- Choose delivery. JSON is convenient for applications, NDJSON for streaming pipelines, CSV for analysts, and webhooks or cloud storage for recurring jobs.
- Plan retries and deduplication. Key records by stable IDs or canonical URLs, store capture timestamps, and make imports idempotent.
- Test a small sample. Compare expected fields, missing values, pagination behavior, and media URL longevity before scheduling a large run.
Exporting profiles, posts, comments, hashtags, and follower counts
Profiles and follower counts
Capture the username, canonical profile URL, user ID when available, follower and following counts, post count, verification state, biography, and external link. Store the observation time because counts change.
Posts and reels
Store the post or reel URL, author, caption or description, hashtags, publication date, likes, comment count, and media URLs. Treat media URLs as potentially expiring and retain only what your purpose and legal basis support.
Comments
Comments can contain personal data and require stricter retention and access controls. Record pagination state and collection time. Avoid assuming that a first page represents all comments.
CSV design
Use one row per entity type or a normalized schema. For repeated hashtags or media URLs, either use a child table or a delimiter that cannot occur in the source data. Keep the raw export separately so you can reprocess it when your schema changes.
Compliance and responsible collection
Meta defines scraping as automated collection of data from a website or other interfaces built for people. Meta also distinguishes authorized from unauthorized scraping and says unauthorized collection can violate its terms. “Public” is therefore not a complete compliance test.
- Review Instagram and Meta terms before collecting.
- Identify a lawful purpose and, where required, a lawful basis.
- Limit collection to fields you need.
- Set retention and deletion rules.
- Document geography, data subjects, and access controls.
- Honor requests, restrictions, and platform signals that prohibit collection.
Reliability, performance, and cost planning
Reliability
Expect public pages and schemas to change. Record request status, provider job ID, response time, item count, and error reason. Retry transient failures with exponential backoff, but cap attempts and avoid creating duplicate jobs.
Performance
Batch URLs where the provider supports it, process exports as streams when possible, and parallelize within documented limits. Separate discovery, extraction, normalization, and storage so a failed export can be resumed without recollecting everything.
Cost
Compare per-record or per-request pricing, scheduled-job charges, storage, bandwidth, and retries. A cheap first run can become expensive if you repeatedly fetch unchanged profiles. Cache canonical URLs and schedule only the frequency your use case needs.
Common errors and fixes
| Error | Likely cause | Fix |
|---|---|---|
| Empty result | Private, deleted, invalid, or geo-restricted profile | Validate the canonical URL, record an explicit empty status, and do not treat it as zero followers. |
| Missing fields | Provider schema differences or a field unavailable for that page type | Version your mapper, preserve raw JSON, and use null instead of guessing. |
| Repeated records | Retry created a second job or pagination was replayed | Deduplicate by stable ID or canonical URL plus timestamp. |
| Rate-limit or block response | Request volume exceeded provider or platform limits | Reduce concurrency, add backoff, and follow the provider’s terms and limits. |
| Expired media URL | Media links are temporary | Download only when permitted and needed; otherwise store the source URL and capture time. |
| Graph API permission error | Wrong account type, missing app setup, or invalid user token | Confirm the account is Professional, review app configuration, and renew the required token. |
Or skip the browser setup
If your deliverable is a visual record of a public Instagram page rather than structured rows, ScreenshotNeo provides a website screenshot API and MCP server. Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result.
One request returns PNG, JPEG, WebP, or PDF. The API supports full-page capture, element selectors, custom CSS and JavaScript, waits, blocking rules, headers, cookies, user agents, geolocation, resizing, caching, signed links, asynchronous jobs, bulk capture, and usage reporting. An MCP server lets Claude, Cursor, and other MCP clients use take_screenshot, get_page_info, and capture_pdf.
See the ScreenshotNeo documentation for the complete parameter list.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.instagram.com/example/ -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.instagram.com/example/"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.instagram.com/example/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There are 1,000 free screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
FAQ
Can I scrape public Instagram profiles without logging in?
Apify’s listed Profile & Posts Actor says no Instagram login, password, cookies, API key, app registration, or access token is required. Whether collection is permitted still depends on applicable terms and law.
Which option supports scheduled runs?
Bright Data documents recurring jobs, webhooks, and cloud-storage delivery. Confirm current limits and pricing before production use.
Should I use Meta’s API or a scraper?
Use Meta’s Graph API for authorized Instagram Professional accounts you manage. Use a hosted scraper or Actor only when public-web collection is appropriate and permitted.
Is a follower count permanent?
No. Store it with a capture timestamp and treat it as an observation, not a timeless account attribute.
Can ScreenshotNeo export comments or follower counts?
No. ScreenshotNeo captures visual pages and PDFs. Use a structured API or Actor for records, and use ScreenshotNeo when you need a clean visual snapshot of the page.
