Using Web Scraping to Collect Social Media Evidence of War Crimes
A careful workflow for collecting, preserving, verifying and securing public social media material for war-crimes investigations.
Short answer: Web scraping can help identify and preserve publicly accessible social media material, but a scrape is only a collection step. It does not authenticate a post, prove that a war crime occurred, establish who is responsible, or replace a documented open-source investigation.
Use a defined purpose, collect only what is necessary and proportionate, preserve the original material and associated information promptly, hash captured files, and corroborate source, time, location, metadata and authenticity. The Berkeley Protocol on Digital Open Source Investigations is the main professional reference for this process.
1. Define the question before you write a scraper
Write down the incident or hypothesis being examined, the relevant dates and places, the platforms and public sources to review, and the intended use of the material. State what you will not collect. The Berkeley Protocol frames collection around an articulable purpose, necessity and proportionality.
- Define the incident, location, time window and relevant actors.
- List the kinds of material that could answer the question: posts, captions, comments, images, videos, profile information, timestamps and links.
- Set inclusion and exclusion rules before collection starts.
- Record who collected each item, when, with which tool and under which settings.
- Plan secure storage and access controls before visiting sources.
2. Legal and ethical boundaries
There is no universal answer to whether scraping is allowed. The applicable rules depend on jurisdiction, platform terms, access controls and the investigation context. Public visibility alone does not settle legality or evidentiary value.
Do not bypass authentication, CAPTCHAs, technical access controls or private groups. Do not use false pretenses to obtain restricted information. Respect platform terms, applicable privacy and data-protection law, and requests that limit automated access. Minimize irrelevant personal data, especially information about children, witnesses and survivors.
A technically successful collection can still be unsafe. Malicious pages may target investigators, and exposing names, faces, locations or contact details can create physical or psychosocial risks. Use isolated accounts or environments where appropriate, keep systems patched, and restrict access to collected material.
3. Choose manual, automated or mixed collection
| Method | Use it when | Main tradeoff |
|---|---|---|
| Manual, itemized capture | The set of relevant posts is small or highly sensitive | More selective and easier to explain, but slower |
| Automated collection | There is a defined, public corpus and content may disappear quickly | Faster and repeatable, but can collect excessive personal data and irrelevant material |
| Mixed workflow | Automation finds candidates and an investigator decides what to preserve | Balances scale and selectivity; requires clear handoff and review records |
The Protocol does not make automation universally good or bad. Match the method to purpose, proportionality, available time and risk.
4. A cautious Python collector for public pages
The example below uses Playwright to visit a list of public URLs, save rendered HTML, record the final URL and title, and compute SHA-256 hashes. It does not log in, defeat access controls or attempt to evade rate limits. Adapt selectors and delays only for sources you are authorized to collect.
from pathlib import Path
from datetime import datetime, timezone
import hashlib
import json
import time
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright
URLS = [
"https://example.org/public-post",
]
OUT = Path("evidence_capture")
OUT.mkdir(exist_ok=True)
def sha256(path):
h = hashlib.sha256()
with path.open("rb") as f:
for chunk in iter(lambda: f.read(1024 * 1024), b""):
h.update(chunk)
return h.hexdigest()
manifest = []
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
context = browser.new_context(
java_script_enabled=True,
user_agent="EvidenceResearchBot/1.0 (contact: investigator@example.org)",
)
page = context.new_page()
for index, url in enumerate(URLS, start=1):
started = datetime.now(timezone.utc).isoformat()
record = {"requested_url": url, "started_at": started}
try:
page.goto(url, wait_until="domcontentloaded", timeout=60000)
page.wait_for_timeout(2000)
html_path = OUT / f"item-{index:04d}.html"
html_path.write_text(page.content(), encoding="utf-8")
record.update({
"final_url": page.url,
"title": page.title(),
"html_file": str(html_path),
"sha256": sha256(html_path),
"status": "captured",
})
except Exception as exc:
record.update({"status": "error", "error": type(exc).__name__ + ": " + str(exc)})
manifest.append(record)
time.sleep(3)
browser.close()
(OUT / "manifest.json").write_text(json.dumps(manifest, indent=2), encoding="utf-8")
print(json.dumps(manifest, indent=2))
Install and run it with:
python -m pip install playwright
playwright install chromium
python collect.py
For potentially probative material, preserve the downloaded media and relevant associated information as well as the rendered page. A screenshot or PDF can be useful work product, but it may not contain everything needed to explain the item later.
5. Capture associated information, not just pixels
For every item, record:
- Original source URL and final URL after redirects.
- Collection date and time in UTC, including the tool and version.
- Displayed author or account name, post identifier and visible timestamp.
- Caption, surrounding text, replies or context needed to interpret the item.
- Downloaded media filenames, content types, dimensions and cryptographic hashes.
- Any warning, login wall, deletion notice, edit indicator or loading failure.
- Collector, case identifier, decision to include the item and chain-of-custody events.
Hash each preserved file with SHA-256 or an equivalent documented method. A hash helps detect later changes to the file; it does not prove who created or uploaded the content, and it does not prove that the claims in a post are true.
6. Preserve promptly and securely
Posts can be deleted or edited, and platforms can remove content or change retention and disclosure practices. Preserve relevant material as soon as practical, keep the original file immutable, and work from copies for analysis.
- Store originals in access-controlled, encrypted storage.
- Keep a write-once or otherwise tamper-evident evidence copy where feasible.
- Maintain a manifest linking each file to its URL, timestamp and hash.
- Log every transfer, transformation and reviewer decision.
- Separate identifying information from broad analyst access when possible.
- Back up securely and test restoration without altering originals.
7. Verify and corroborate before drawing conclusions
A capture records what was available at collection time. Evaluate authenticity and context separately:
- Compare the account, URL, post identifier and surrounding posts.
- Inspect metadata when available, while recognizing that platforms often strip or rewrite it.
- Check location using visible landmarks, shadows, language, weather, maps and independent imagery.
- Check time using platform timestamps, events, sun position, weather and other dated sources.
- Look for earlier uploads, cropped versions, translations or contradictory versions.
- Seek independent corroboration and document alternative explanations.
Separate what the image or video visibly shows from the uploader’s caption and from an investigator’s inference. Record uncertainty instead of presenting a lead as a conclusion about a war crime or criminal responsibility. The 2024 guide Evaluating Digital Open Source Imagery emphasizes authenticity, source, metadata, location and time in assessment.
8. Screenshots, PDFs and fuller downloads
| Output | Good for | Limitations |
|---|---|---|
| Screenshot | Readable visual record and quick review | May omit hidden context, metadata, replies or original media |
| Paginated review copy and sharing within a case team | Usually a derivative; preserve source files separately | |
| Rendered HTML | Page structure, visible text and collection context | Scripts, embeds and later network content may not be fully represented |
| Original media plus associated information | Material with potential probative value | Requires more storage, documentation and security controls |
ICC and Eurojust guidance recommends downloading or recording online content with relevant associated information and assigning a hash when it may have probative value. Choose the least burdensome method that still lets a later reviewer understand and assess the item.
9. Rate limits, reliability and performance
- Use conservative request spacing and honor published limits.
- Cache only material you are authorized to retain; record cache status and retrieval time.
- Retry transient network failures with bounded exponential backoff, never by rapidly repeating requests.
- Set explicit navigation and download timeouts and record timeout errors.
- Run small pilot collections first to measure irrelevant-data volume and storage needs.
- Prefer itemized queues so a failure does not lose the entire collection.
- Keep a deterministic manifest so the same URL can be compared across collection dates.
Automation reduces repetitive work but does not remove review. The limiting cost is often analyst time, verification and secure retention rather than network transfer.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 403, 429 or repeated challenge page | Access restriction or rate limit | Stop, review authorization and platform rules, slow down, or use manual collection. Do not bypass the control. |
| Blank HTML | Content rendered by JavaScript or blocked resources | Use a real browser context, wait for a specific selector, and record what failed to load. |
| Post disappears between review and capture | Deletion, moderation or editing | Preserve promptly, retain timestamps and hashes, and document the unavailable state. |
| Wrong language or misleading translation | Automatic translation or transcription error | Keep the original, identify the translator, and obtain an independent translation. |
| Hash mismatch | File changed, was recompressed or was replaced | Quarantine the derivative, retain the original if available, and investigate the chain of custody. |
| Too much personal data | Overbroad query or bulk collection | Stop, narrow scope, delete unnecessary copies under the case policy, and document the decision. |
11. Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server. It can remove cookie and consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP tools let Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf.
Use it for a review copy, then preserve the underlying source and associated information required by your investigation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for options such as full-page capture, PDF output, custom headers, cookies, waiting rules, blocked resources, caching, signed links, asynchronous jobs and bulk capture. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.
Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.
12. Frequently asked questions
Can a screenshot alone prove a war crime?
No. It is a record of displayed content. Authentication, context, corroboration and legal analysis are still required.
Should every public post be scraped?
No. Collect only material justified by the investigation’s purpose and proportionate to the need.
Does hashing make evidence authentic?
No. Hashing detects later file changes; it does not establish authorship, upload history or truth.
What if the platform deletes the post?
Preserve promptly, retain your collection record and document what was unavailable, when and how you observed it.
Can I scrape behind a login if I have credentials?
Only if the access and collection are lawful, authorized and consistent with the investigation’s rules. The general-purpose example here does not bypass restricted access.
Primary references
- Berkeley Protocol on Digital Open Source Investigations, UC Berkeley Human Rights Center and OHCHR (2020).
- Documenting International Crimes and Human Rights Violations for Accountability Purposes, ICC Office of the Prosecutor and Eurojust (2022).
- Evaluating Digital Open Source Imagery: A Guide for Judges and Fact-finders (2024).
- Digitally Disappeared, Lindsay Freeman, UC Berkeley Human Rights Center (2022).

