How to Scrape Crunchbase in 2026
Crunchbase’s terms prohibit scraping. Learn the authorized API and data options, licensing checks, and a compliant workflow for company research.

Direct answer: Crunchbase’s Terms of Service prohibit crawling, scraping, or spidering pages, data, or any portion of the service, whether the method is manual or automated. The same terms prohibit bypassing restrictions or using automated means to access, download, export, or exploit content. That means a browser script, scraper, proxy rotation system, or similar crawler is not an authorized way to collect Crunchbase data.
The compliant path is to request an official Crunchbase API or data product, review the license you are offered, and build only within that written scope. Crunchbase describes its API as a read-only REST service for approved developers, and its documentation says full API access requires an Enterprise or Applications license. Your order form and current data-access agreement control the fields, volume, retention, attribution, and redistribution rights you actually have.
This guide explains how to make that decision in 2026, how to prepare an access request, how to design a compliant ingestion pipeline, and what to check before production use.
What Crunchbase’s terms prohibit
The Terms of Service govern access to Crunchbase websites, applications, products, services, and content. The prohibited-use list includes: “Crawls,” “scrapes,” or “spiders” any page, data, or portion of or relating to the Service (through use of manual or automated means). The terms also address circumvention, including using automated means to access, download, export, or exploit content in ways that bypass restrictions. The page identifies November 26, 2024 as its last-updated date; check the live page before publication or implementation.
In practical terms, do not treat a public page, a browser session, or an exposed network request as permission to copy the underlying records. A low request rate does not change the nature of an unauthorized crawl. Neither does adding delays, rotating IP addresses, solving a challenge, or limiting the number of pages.
These are the relevant primary sources: Crunchbase Terms of Service and the Crunchbase Data Access Terms. The contract that applies to your account is the controlling document for licensed use.
Use the official API or data product
Crunchbase directs users who need uses beyond the website terms to its products, including API and data offerings. Its API documentation describes a read-only RESTful service for approved developers and states that full API access requires an Enterprise or Applications license. Contact Crunchbase for the package that matches your use case; public documentation does not establish a universal price or entitlement.
The API documentation lists a documented limit of 200 calls per minute. Treat that as a property of an approved service, not as a target for scraping the website. The same documentation and your order form determine which endpoints, entities, fields, and actions are available.
The Basic API is not a new-user workaround
Crunchbase’s Basic API documentation says: “We are no longer offering our Basic API.” Existing users may be able to view previously generated keys, but they cannot regenerate them. A new project should go through the current access process rather than relying on an old key or a tutorial written for the legacy product. See the Basic API documentation for that limitation.
Decide whether licensed access fits your project
Before contacting sales or engineering, write down the answers to these questions. They make the request specific and expose contract issues early.

| Question | Why it matters |
|---|---|
| Who will use the data? | Internal analysts, a customer-facing application, and an external data product may require different rights. |
| Which entities and fields? | Companies, people, funding rounds, investors, and events may not be included in every package. List only what you need. |
| How fresh must records be? | A daily report, a weekly research job, and real-time enrichment imply different update and volume requirements. |
| How many records and requests? | Estimate initial backfill, daily changes, retries, and concurrent workers. Ask which limits apply to your license. |
| Will data leave your organization? | Customer exports, search results, reports, and derived scores can raise redistribution and attribution questions. |
| How long will you retain it? | Confirm retention, deletion, backup, and deletion-on-termination requirements in writing. |
| Does the product contain personal data? | Document access controls, lawful purpose, regional handling, and deletion procedures with your legal and privacy teams. |
Build a compliant ingestion workflow
- Obtain written authorization. Keep the signed order form, agreement, API credentials, and any implementation instructions together.
- Define an allowlist. Record permitted endpoints, fields, environments, users, destinations, and purposes. Reject requests outside the allowlist.
- Separate credentials by environment. Store production keys in a secret manager. Use different keys for development and production if the contract permits it.
- Respect the documented limits. Implement a token bucket or equivalent limiter below the limit assigned to your account. Add bounded retries only for transient responses.
- Log provenance. For each batch, store retrieval time, endpoint, request identifier, schema version, and license context. Avoid logging full personal records.
- Apply retention and deletion. Automate deletion or reprocessing when your agreement requires it. Include backups and derived tables in the procedure.
- Review redistribution. Before exposing a field in an API, dashboard, export, or model feature, verify that your order form allows that destination.
Local processing example
The following example processes a file that your licensed export process has already produced. It does not connect to Crunchbase or collect website content. Adapt the field names to the schema in your agreement.
import json
from pathlib import Path
source = Path('licensed_companies.json')
records = json.loads(source.read_text())
# Keep only fields approved for this internal report.
allowed = {'uuid', 'name', 'short_description', 'country_code'}
clean = [{k: row.get(k) for k in allowed} for row in records]
Path('internal_company_index.json').write_text(
json.dumps(clean, ensure_ascii=False, indent=2)
)
print(f'Wrote {len(clean)} records')
For a production pipeline, add schema validation, an audit table, a dead-letter queue for malformed records, and an explicit deletion job. Do not add a web crawler around this code.
Rate limits, retries, and reliability
Crunchbase documents a 200-calls-per-minute API limit. Your licensed package may impose additional limits or a different allowance, so configure the value rather than hard-coding the public figure. A safe client should:
- Use exponential backoff with jitter for temporary 429 and 5xx responses.
- Honor a server-provided Retry-After value when present.
- Use idempotent job identifiers so a retry cannot duplicate a write.
- Cap retries and send exhausted requests to a review queue.
- Persist a cursor or checkpoint after each successful page or batch.
- Alert on sustained authorization failures instead of retrying them.
Do not assume that retrying faster will improve throughput. It can increase failures and may violate your agreement. Measure successful records per minute, error classes, and lag from the provider’s update time, while keeping only the telemetry your contract permits.
Cost and licensing questions
Crunchbase’s public materials reviewed for this article do not publish a universal price, field entitlement, or redistribution rule. Ask for a quote that states the billing unit, included entities and fields, refresh cadence, overage treatment, support level, and limits. Confirm whether a historical backfill is priced differently from ongoing updates.
Compare offers using the same workload: initial records, monthly changes, request volume, users, environments, retention, and external distribution. A lower headline price can be unsuitable if it excludes the fields or destinations your application needs. Keep the final order form with your architecture decision record.
Common mistakes and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| A script receives a block page or challenge | Automated website access is restricted. | Stop the crawler. Request an authorized API or data product. |
| An old API key cannot be regenerated | The legacy Basic API is no longer offered to new users. | Contact Crunchbase about current API access. |
| HTTP 401 or 403 from the API | Invalid credentials, an inactive license, or an endpoint outside your package. | Check the key, environment, endpoint entitlement, and account status with Crunchbase. |
| HTTP 429 responses | Your client exceeded the applicable rate limit. | Reduce concurrency, honor Retry-After, and confirm the licensed limit. |
| Fields are missing | The field is not included in your plan, is null for that entity, or the schema changed. | Check the schema and order form; do not infer permission from a missing field. |
| Customers can see records in an export | Redistribution rights were not confirmed. | Pause the export and obtain written clarification. |
| Deletes do not remove data from reports | Derived tables, caches, or backups were omitted. | Trace lineage and include every copy in the deletion workflow. |
Compliance checklist before launch
- Current Terms of Service and Data Access Terms reviewed.
- Signed order form identifies permitted purpose and destinations.
- API credentials are stored and rotated securely.
- Endpoints, fields, volume, and refresh schedule are allowlisted.
- Rate limiting, bounded retries, and checkpointing are implemented.
- Attribution requirements are visible wherever required.
- Retention, deletion, backup, and termination procedures are tested.
- External redistribution and customer exports have written approval.
- Personal-data handling has privacy and security owners.
- A process exists to re-check terms when the product or contract changes.
Or skip the browser setup
If your project needs screenshots of public company pages or research dashboards, you can use ScreenshotNeo for the visual capture step without maintaining browser automation. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options. A minimal request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture, CSS-selector element capture, device presets, custom viewports, retina scale, PDF settings, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, bulk capture, and a usage API. 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Can I scrape only a few Crunchbase pages for internal research?
The terms prohibit crawling and scraping through manual or automated means. Request an authorized product for the intended use.
Is 200 calls per minute permission to crawl?
No. It is a documented API limit for approved access, not authorization to scrape the website.
Can I redistribute API results to customers?
Only if your current order form and agreement expressly allow it. Confirm fields, destinations, attribution, and retention in writing.
What should I do if a tutorial uses the Basic API?
Treat it as legacy guidance. Crunchbase says it is no longer offering the Basic API to new users; follow the current access process.
How often should I re-check the rules?
Review the live terms and your contract before launch and whenever your use case, data fields, audience, or product distribution changes.
Final decision
Do not build a Crunchbase scraper in 2026. Build against an approved API or data product whose written license covers your exact purpose, fields, volume, retention, and distribution. If you cannot obtain those permissions, change the data source or the product requirement before collecting anything.


