Build Credible Influencer Lists With Instagram Scraping
Build a dated, auditable Instagram influencer list with relevance checks, engagement-quality review, compliance controls, and reproducible evidence.
Short answer: Build the list as a dated research dataset, not a ranking of follower counts. Define the campaign brief, collect candidates through a permitted method, save comparable evidence, normalize and deduplicate accounts, score relevance before popularity, inspect engagement quality and audience fit, document disclosure risks, and retain an inclusion decision for every account.
A credible list lets another person answer five questions for every row: why this account fits, what evidence supports the decision, when the evidence was observed, how the data was collected, and how confident the reviewer is. Instagram metrics change constantly, so an undated list becomes difficult to audit as soon as it is published.
What a credible influencer list contains
Follower count is a discovery signal. It is not proof of influence, audience quality, or campaign fit. A useful record combines relevance, audience fit, genuine interaction, business suitability, and reproducible evidence.
| Field | What to record | Why it matters |
|---|---|---|
| Identity | Handle, canonical profile URL, account type, collection date | Prevents duplicate, renamed, or misclassified accounts. |
| Discovery provenance | Search query, source, referral, or permitted collection method | Allows the list to be reproduced and audited. |
| Audience fit | Stated location, language, visible audience clues, topical alignment | Separates topical popularity from campaign usefulness. |
| Content fit | Recent post types, format capability, posting consistency, quality notes | Confirms the creator can deliver the required campaign format. |
| Engagement evidence | Sampled posts, visible likes/comments where available, comment-quality notes | Shows whether interactions appear specific and conversational. |
| Risk review | Sudden spikes, repeated generic comments, copied comments, disclosure patterns | Flags signals that require manual review. |
| Decision trail | Include/exclude reason, reviewer, confidence, conflicts, next review date | Makes the final recommendation explainable. |
1. Write the campaign brief before collecting accounts
Write the constraints first. Otherwise, the collection process will optimize for whatever is easiest to count, usually follower totals.
- Niche: the subjects and subtopics that qualify.
- Target audience: customer type, interests, buying context, and seniority.
- Geography: countries, regions, or cities that matter.
- Language: required content and audience languages.
- Format: feed post, Reel, Story, live session, product review, or another deliverable.
- Objective: awareness, qualified traffic, app installs, sales, signups, or another measurable outcome.
- Budget range: expected spend and whether product, payment, or both are involved.
- Exclusions: restricted categories, conflicts, unsafe content, competitors, or unavailable regions.
Turn the brief into explicit inclusion rules. For example: “Include creators whose recent content is about independent coffee shops, whose audience appears primarily in the United Kingdom, who can publish a Reel, and who have no unresolved brand-safety conflict.”
2. Collect candidates through a permitted method
Before automating any collection, check the current Instagram and Meta terms, permissions, and authorization requirements for your use case. Public visibility does not automatically grant permission for unrestricted automated collection. Do not automate likes, follows, comments, or other artificial interactions.
Use the least invasive method that meets the brief. Possible sources include an authorized account connection, a platform export, a vendor with permission to provide the data, or a small manual research pass. Store the method with every candidate.
Candidate record example
{
"handle": "example_creator",
"profile_url": "https://www.instagram.com/example_creator/",
"collected_at": "2026-10-01T12:00:00Z",
"discovery_source": "permitted search export",
"collection_method": "authorized export",
"account_type": "creator",
"followers_visible": 24800,
"following_visible": 610,
"bio": "Coffee, travel and independent cafes",
"stated_location": "Manchester, UK",
"language_observed": "English",
"recent_posts_sampled": 12,
"include_decision": null,
"confidence": null,
"evidence": []
}
Python: parse an authorized JSON export
This script creates a normalized candidate file from data you are authorized to access. It does not log in, bypass controls, or create Instagram activity.
import json
import sys
from datetime import datetime, timezone
from urllib.parse import urlparse
def canonical_handle(value: str) -> str:
value = value.strip().lstrip("@").rstrip("/")
if "/" in value:
value = urlparse(value).path.strip("/").split("/")[0]
return value.lower()
def normalize(row: dict) -> dict:
handle = canonical_handle(row.get("handle", row.get("username", "")))
return {
"handle": handle,
"profile_url": f"https://www.instagram.com/{handle}/",
"collected_at": row.get("collected_at") or datetime.now(timezone.utc).isoformat(),
"discovery_source": row.get("discovery_source", "authorized export"),
"collection_method": row.get("collection_method", "authorized export"),
"followers_visible": row.get("followers_visible"),
"following_visible": row.get("following_visible"),
"bio": row.get("bio", ""),
"stated_location": row.get("stated_location", ""),
"language_observed": row.get("language_observed", ""),
"evidence": row.get("evidence", []),
}
if len(sys.argv) != 3:
raise SystemExit("Usage: python normalize_export.py input.json output.json")
with open(sys.argv[1], encoding="utf-8") as source:
rows = json.load(source)
seen = set()
normalized = []
for row in rows:
item = normalize(row)
if not item["handle"] or item["handle"] in seen:
continue
seen.add(item["handle"])
normalized.append(item)
with open(sys.argv[2], "w", encoding="utf-8") as target:
json.dump(normalized, target, indent=2, ensure_ascii=False)
print(f"Wrote {len(normalized)} unique candidates")
Node.js: normalize the same export
import fs from "node:fs";
const [input, output] = process.argv.slice(2);
if (!input || !output) {
console.error("Usage: node normalize-export.mjs input.json output.json");
process.exit(1);
}
const rows = JSON.parse(fs.readFileSync(input, "utf8"));
const seen = new Set();
const normalized = [];
for (const row of rows) {
let handle = String(row.handle ?? row.username ?? "")
.trim().replace(/^@/, "").replace(/\/$/, "");
if (handle.includes("/")) handle = handle.split("/").filter(Boolean)[0];
handle = handle.toLowerCase();
if (!handle || seen.has(handle)) continue;
seen.add(handle);
normalized.push({
handle,
profile_url: `https://www.instagram.com/${handle}/`,
collected_at: row.collected_at ?? new Date().toISOString(),
discovery_source: row.discovery_source ?? "authorized export",
collection_method: row.collection_method ?? "authorized export",
followers_visible: row.followers_visible ?? null,
following_visible: row.following_visible ?? null,
bio: row.bio ?? "",
stated_location: row.stated_location ?? "",
language_observed: row.language_observed ?? "",
evidence: row.evidence ?? []
});
}
fs.writeFileSync(output, JSON.stringify(normalized, null, 2));
console.log(`Wrote ${normalized.length} unique candidates`);
3. Normalize, deduplicate, and classify
Canonicalize handles to lowercase, remove leading @ symbols, and convert profile links to one canonical form. Keep a history field when a creator changes handles. Deduplicate before scoring; otherwise one creator can appear several times with different URLs.
Separate personal creators, brands, agencies, and other account types. An agency account may be relevant for sourcing but should not be scored as an individual creator. Mark private, unavailable, renamed, and deleted profiles instead of silently removing them.
4. Capture comparable evidence
Use the same observation window and sampling rules for every account. Save the collection date because follower counts, likes, comments, biographies, and recent posts change.
- Record the visible follower and following counts.
- Save the biography and stated location.
- Sample a fixed number of recent posts, such as the latest 10 or 12 available posts.
- Record post dates, formats, visible likes and comments where available, and whether the post is sponsored.
- Copy a small sample of captions and comments for qualitative review, subject to your permitted collection method.
- Save evidence links or screenshots with timestamps.
- Record the reviewer and any uncertainty.
Do not mix a current follower count with interactions from an older period without labeling the mismatch. A dated observation is more useful than a precise-looking number with no provenance.
5. Score relevance before popularity
Use a transparent rubric. The weights below are a starting point; adapt them to the campaign and keep the same rubric for all candidates.
| Dimension | Example weight | Review questions |
|---|---|---|
| Niche fit | 25% | Does recent content consistently cover the campaign subject? |
| Audience geography and language | 20% | Do visible signals match the target market and language? |
| Content and format fit | 15% | Can the creator deliver the required format at the needed quality? |
| Engagement quality | 20% | Are interactions relevant, specific, and conversational? |
| Consistency and reliability | 10% | Is posting regular enough for the campaign schedule? |
| Brand safety and disclosure history | 10% | Are there unresolved conflicts or unclear sponsorship practices? |
Keep raw observations beside the score. A score without evidence is only an opinion, and a high score should not override an exclusion rule.
6. Evaluate engagement quality and fake-follower signals
No universal authoritative Instagram engagement-rate cutoff is established by the cited sources. Compare like with like: the same niche, format, audience size range, geography, date window, and sampling method. Document the formula and sample instead of publishing a single “good rate” threshold.
A reproducible comparison formula
post engagement rate = (visible likes + visible comments) / visible followers × 100
sample rate = average post engagement rate across the defined sample
Label missing values as unavailable. Do not treat hidden, rounded, or estimated metrics as exact.
Inspect quality signals alongside the arithmetic:
- Comments should contain specific reactions to the post rather than repeated generic phrases.
- Look for sudden unexplained follower or interaction spikes.
- Check for copied comments appearing across unrelated posts.
- Compare the apparent audience geography and language with the campaign brief.
- Look for coordinated or bot-like activity and unusual mismatches between follower size and sustained interaction.
- Treat third-party authenticity scores as screening aids, not proof.
The FTC distinguishes genuine influence from fake indicators such as bot-generated or otherwise non-genuine accounts. Meta has described enforcement against services that artificially inflated Instagram likes and followers and said those activities violated Instagram terms and policies. Use the FTC guidance and Meta policy information when defining your review controls.
7. Check disclosure and endorsement risk
A material connection can include payment, employment, family relationships, or free or discounted products. The disclosure should appear with the endorsement and be clear to the audience. The FTC advises against vague labels such as “sp,” “spon,” or “collab” without a clear explanation.
When reviewing sponsored history, distinguish a disclosed paid endorsement from a claim supported by evidence. An influencer cannot describe personal experience with a product they have not tried. The FTC has also addressed fake reviews, virtual influencers, tags, disclosure adequacy, and potential liability for advertisers, endorsers, and intermediaries in its revised guidance.
| Observation | Record | Follow-up |
|---|---|---|
| Paid or gifted post appears disclosed | Exact disclosure, post URL, date | Check whether placement is clear and conspicuous. |
| Product claim is unusually strong | Claim text and evidence available | Ask whether the creator actually used the product and whether the advertiser can substantiate it. |
| Disclosure is vague or missing | Screenshot and context | Escalate for legal or compliance review before outreach. |
8. Keep a decision trail
For every account, store an inclusion or exclusion reason, evidence links or captures, reviewer, confidence level, conflicts, and next review date. A compact decision record might look like this:
{
"handle": "example_creator",
"decision": "include",
"reason": "Strong coffee niche fit, UK audience signals, consistent Reels, conversational comments",
"evidence_ids": ["obs-2026-10-01-001", "obs-2026-10-01-002"],
"confidence": "medium",
"conflicts": [],
"reviewer": "analyst-07",
"next_review_date": "2026-11-01"
}
9. Capture evidence without clutter
When you need a dated visual record of a public page, capture the specific profile or post after confirming that your method is permitted. Preserve the capture timestamp and source URL next to the dataset row.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed. Each step can be turned off.
Only clean shots are billed. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the result in X-Page-Verdict and X-Billed headers. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const body = Buffer.from(await res.arrayBuffer());
require('node:fs').writeFileSync('shot.webp', body);
See the ScreenshotNeo documentation for the full option set, including full-page capture, CSS selectors, dark mode, device presets, retina scale, custom CSS and JavaScript, wait conditions, blocked resources, cookies, headers, user agents, caching, signed links, asynchronous jobs, bulk capture, and usage reporting. Do not capture pages or data your collection method does not authorize you to access.
There are 1,000 screenshots per month on the free plan with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.
Performance, reliability, and cost controls
- Bound the sample: define the number of recent posts and candidates before collection.
- Cache evidence: retain a capture for the review period instead of repeatedly fetching the same page.
- Use backoff: schedule permitted requests with delays and stop when access controls or errors appear.
- Separate discovery from verification: collect a broad candidate set, then spend expensive review time on likely matches.
- Track freshness: assign a next review date and recapture records when campaign decisions depend on current metrics.
- Handle partial data: preserve unavailable fields and explain why they are missing.
- Store provenance: keep source, method, timestamp, reviewer, and evidence identifier with every observation.
A smaller dataset with dated evidence and transparent methodology is often more useful than a large directory with stale counts. Budget for rechecks because follower counts and content change.
Troubleshooting
| Problem | Likely cause | Fix |
|---|---|---|
| Duplicate creators appear under different rows | Handle changes or inconsistent URL formats | Canonicalize handles, retain an alias history, and deduplicate before scoring. |
| Follower count is missing | Metric is hidden, unavailable, or not included in the permitted export | Store null or unavailable; do not estimate silently. |
| Engagement rates look incomparable | Different niches, formats, dates, or sample sizes | Group like with like and document the formula and window. |
| Comments look suspiciously repetitive | Copied, generic, coordinated, or bot-like activity | Flag for manual review and reduce confidence; do not call it proof by itself. |
| Audience geography does not match the brief | Creator popularity is coming from another market | Prioritize audience-fit evidence over follower count or exclude the account. |
| Automation is blocked or returns an authentication challenge | Permission, rate, or platform-control issue | Stop automated collection, verify authorization and current terms, and use an approved export or manual workflow. |
| Evidence cannot be reproduced | No timestamp, source, query, or collection method | Add provenance fields and recapture under a documented method. |
| ScreenshotNeo returns an unexpected result | Page timeout, bot check, blank response, or unsupported access condition | Inspect X-Page-Verdict and X-Billed, then adjust wait, viewport, headers, or permitted access settings in the docs. |
Checklist before you publish or contact anyone
- Campaign niche, audience, geography, language, format, objective, budget, and exclusions are written down.
- Every account has a collection date, source, method, and canonical profile URL.
- Duplicates, renamed accounts, brands, agencies, and unavailable profiles are classified.
- Recent posts were sampled with the same rule across candidates.
- Relevance and audience fit were scored before popularity.
- Engagement was compared within a defined sample, not against an unsupported universal cutoff.
- Comment quality, sudden spikes, audience mismatch, and copied comments were reviewed.
- Disclosure and product-claim risks were recorded for sponsored history.
- Every inclusion and exclusion has evidence, a reviewer, confidence, and a next review date.
- Collection and automation methods were checked against current Instagram and Meta permissions.
FAQ
Is scraping Instagram legal?
The answer depends on the data, jurisdiction, purpose, and method. Public visibility does not automatically authorize unrestricted automated collection. Check current Instagram and Meta terms, permissions, privacy obligations, and any applicable laws before implementation.
What is a good Instagram engagement rate?
The cited authorities do not establish a universal cutoff. Compare similar creators using the same niche, format, geography, date window, and formula, then inspect interaction quality.
Can follower count prove an influencer is credible?
No. It helps discover accounts, but credibility requires relevance, audience fit, genuine interaction, business suitability, and dated evidence.
Should I trust an authenticity score from a third-party tool?
Use it as a screening aid. Validate the result with recent-post samples, comment quality, audience fit, follower changes, and your own documented review.
How often should an influencer list be refreshed?
Set the interval from campaign risk and timing. Assign a next review date to every record and refresh before major outreach or spend decisions.
What should I do when an influencer discloses a sponsorship?
Record the exact disclosure and context. A clear disclosure addresses the relationship; it does not by itself prove that every product claim is true.


