How to Scrape Websites with CrewAI
Build a CrewAI scraper for one page, selected elements, or a bounded crawl. Choose the right tool, handle dynamic content, and validate results.

To scrape a page with CrewAI, install its tools package, create a ScrapeWebsiteTool with the page URL, and call run(). Add the tool to an agent when scraping is one step in a larger workflow. For selector-based extraction, browser interaction, or crawling multiple pages, use the matching tool instead of asking a basic fetch tool to do more than it supports.
This guide covers direct Python use and agent workflows, choosing among CrewAI’s documented tools, handling dynamic pages, validating output, and troubleshooting. CrewAI’s integration pages span several documentation versions, so confirm imports and options against the version installed in your project. [ScrapeWebsiteTool]
1. Install CrewAI’s tools and scrape one page
The official example installs the tools extra. In a virtual environment, run:

python -m pip install 'crewai[tools]'
Then fetch a page directly:
from crewai_tools import ScrapeWebsiteTool
scraper = ScrapeWebsiteTool(website_url="https://example.com")
text = scraper.run()
print(text)
This is the shortest path when you know the URL and want the page content. It is a page-fetch and HTML-parsing tool; do not assume it runs the page in a browser or waits for JavaScript to render. If the page builds its content client-side, use a browser-based approach or an extraction service whose documented behavior fits the page. [CrewAI ScrapeWebsiteTool documentation]
Pass the URL when the agent calls the tool
You can also construct the scraper without a fixed URL. This lets an agent choose a URL based on its task context. Keep the task bounded: specify which pages it may inspect and what fields it should return.
from crewai import Agent, Task, Crew
from crewai_tools import ScrapeWebsiteTool
scraper = ScrapeWebsiteTool()
researcher = Agent(
role="Page researcher",
goal="Extract only the requested facts from an allowed page",
backstory="You return concise, source-grounded findings.",
tools=[scraper],
verbose=True,
)
task = Task(
description=(
"Read https://example.com and return the page title, "
"the main heading, and the first three product names. "
"If a field is absent, say so. Do not infer missing values."
),
expected_output="A JSON object with title, heading, and product_names fields.",
agent=researcher,
)
result = Crew(agents=[researcher], tasks=[task]).kickoff()
print(result)
The exact task and agent APIs can change by CrewAI version. This pattern follows the documented scraper integration: put the tool in the agent’s tools list and describe a concrete extraction goal. When the only job is retrieving a known page, calling run() directly is simpler and avoids agent orchestration. [CrewAI tool example]
2. Pick the tool that matches the page
| Need | Start with | Tradeoff |
|---|---|---|
| Fetch one ordinary page | ScrapeWebsiteTool |
Simple URL-based HTTP fetch and HTML parsing. No documented browser rendering. |
| Extract a known page section | ScrapeElementFromWebsiteTool |
Requires a CSS selector that matches the current page markup. |
| Wait for or interact with browser-rendered content | SeleniumScrapingTool |
CrewAI labels it “currently in development”; verify status and behavior before relying on it. |
| Extract one page through a hosted extraction API | FirecrawlScrapeWebsiteTool |
Requires Firecrawl configuration and an API key. |
| Crawl a site from a starting URL | FirecrawlCrawlWebsiteTool |
Set inclusion and exclusion rules, depth, page limits, and timeout to bound the run. |
CrewAI’s overview also names Browserbase for cloud browser infrastructure and Stagehand for complex interaction, and recommends Firecrawl for scale and performance. Those are CrewAI’s selection suggestions, not comparative benchmark results. The reviewed documentation provides no comparable speed, accuracy, reliability, or price measurements. [CrewAI web-scraping overview]

Extract a section with a CSS selector
Use ScrapeElementFromWebsiteTool when you know the selector for the content of interest. The documented tool supports a URL, selector, and optional cookies; it returns matching text joined by newlines. The documentation describes a requests and Beautiful Soup approach, so this is still not a substitute for browser rendering on JavaScript-only pages. [CrewAI ScrapeElementFromWebsiteTool documentation]
from crewai_tools import ScrapeElementFromWebsiteTool
extract = ScrapeElementFromWebsiteTool(
website_url="https://example.com/products",
css_element="article.product h2",
)
text = extract.run()
print(text)
Inspect the target page’s actual HTML before settling on a selector. A selector that matches no nodes can produce an empty result; a selector that matches a broad container can capture navigation, recommendations, or repeated text as well as the intended field. Treat selectors as page-specific configuration and revisit them when the site changes.
Use a browser tool for interactive pages, with its maturity caveat
CrewAI documents SeleniumScrapingTool with URL and CSS selector inputs, optional cookies, a wait time, and text or HTML output. Its page lists Selenium, webdriver-manager, Chrome, and Chrome WebDriver requirements. The same page says the tool is currently in development and users may encounter unexpected behavior. Check the current documentation and test against the exact target before building a production dependency on it. [CrewAI SeleniumScrapingTool documentation]
For a client-rendered page, the key requirement is that content has rendered before extraction. A fixed delay may be too short on a slow response and wasteful on a fast one; when possible, wait for a meaningful selector or page state. If the documented integration cannot reliably express the interaction, use a browser automation system suited to that workflow or an external extraction service.
Scrape or crawl with Firecrawl
The Firecrawl integrations cover two different jobs. FirecrawlScrapeWebsiteTool extracts one URL and documents options for main-content filtering, raw HTML, or an LLM extraction prompt and schema. FirecrawlCrawlWebsiteTool starts from a URL and supports include or exclude patterns, main-content control, crawl depth, page limits, and timeout. Both require Firecrawl API configuration; its documentation shows the FIRECRAWL_API_KEY environment variable. [Firecrawl scrape tool, Firecrawl crawl tool]
import os
from crewai_tools import FirecrawlScrapeWebsiteTool
# Set FIRECRAWL_API_KEY in the environment before running this script.
if not os.environ.get("FIRECRAWL_API_KEY"):
raise RuntimeError("Set FIRECRAWL_API_KEY before running")
scraper = FirecrawlScrapeWebsiteTool(
url="https://example.com",
only main content=True,
)
print(scraper.run())
Option names can differ by installed version. The snippet above illustrates the configuration shape, but Python does not allow a keyword argument containing spaces. Check the current Firecrawl tool page for the exact spelling of the main-content option before using it; a minimal, version-safe starting point is:
from crewai_tools import FirecrawlScrapeWebsiteTool
scraper = FirecrawlScrapeWebsiteTool(url="https://example.com")
print(scraper.run())
For a crawl, set a page ceiling and scope rules rather than handing an agent a broad domain and an open-ended instruction. The integration’s documented controls are useful for limiting what it visits, but validate what was actually returned and deduplicate results before downstream use.
3. Make the extraction useful and safe
Give the agent a narrow output contract
“Scrape this site” leaves decisions about scope, fields, and format to the agent. A better task states:
- Scope: the permitted URL or a defined set of pages.
- Fields: the exact values to extract, such as title, date, and price.
- Format: JSON, a table, or one record per page.
- Missing-data rule: return null or an explicit missing marker; do not guess.
- Limits: maximum pages, depth, and acceptable runtime for a crawl.
Validate the returned data in ordinary Python before passing it to another agent or storing it. For example, reject records with missing required fields, normalize whitespace, check URL uniqueness, and verify that a numeric field parses as a number. An LLM-produced structure still needs validation: a requested JSON shape does not guarantee that every value is present or correctly typed.
Respect site controls and handle requests carefully
CrewAI’s scraping overview advises checking robots.txt and the site’s scraping policies, using delays or rate limits, setting an appropriate identifying user agent, handling network errors and blocked requests, and validating and cleaning results. Whether a particular scrape and use of its data are allowed depends on the target, terms, rights, and applicable law; the tool documentation does not settle that question. [CrewAI scraping guidance]
The current ScrapeWebsiteTool page describes an SSRF-safe HTTP helper: it checks the requested URL and redirects against private or reserved network ranges and pins the TCP connection to a checked IP. Treat this as a property documented for that tool specifically, not as a guarantee for every integration or custom tool. If your application accepts user-supplied URLs, apply equivalent destination controls to every fetch path and redirect. [CrewAI ScrapeWebsiteTool security notes]
4. Troubleshoot common failures
| Symptom | Likely cause | What to do |
|---|---|---|
Import error for crewai_tools |
The tools extra is missing, or the active Python environment differs from the one where it was installed. | Activate the project environment and install crewai[tools] there. Confirm the package and documentation version match. |
| Empty or incomplete content | The page may render content in JavaScript, the requested selector may not match, or the response may omit the expected section. | Inspect the returned text and source markup; verify selector matches. For rendered content, use a browser-based option or documented extraction service. |
| Selector returns unrelated text | The selector matches a broad or repeated container. | Narrow it to the content node, test against a representative page, and validate the number and shape of matches. |
| Selenium cannot start Chrome | Chrome, the driver, or Selenium dependencies are absent or incompatible; the CrewAI tool is also documented as in development. | Follow the current tool page’s dependency instructions, check browser and driver compatibility, and verify whether the integration remains appropriate for your version. |
| Firecrawl authentication failure | FIRECRAWL_API_KEY is unset, invalid, or unavailable to the running process. |
Set the key in the process environment or secret store. Do not put it in task text or commit it to source control. |
| Crawl runs too broadly or long | The starting scope, depth, or page ceiling is too permissive. | Use include and exclude patterns, reduce depth and page limit, and set a timeout appropriate to the job. |
| Intermittent failures or blocked requests | Network trouble, rate limits, site controls, or transient upstream errors. | Record the failing URL and error; use bounded retries with backoff, respect rate limits, and avoid retrying permanent errors indefinitely. |
| Agent gives prose instead of usable records | The task has no strict output contract, or the page lacks a requested field. | Specify fields and missing-value behavior, then parse and validate the output before using it. |
5. Performance, reliability, and cost
A direct page fetch has fewer moving parts than a browser workflow: it avoids launching a browser and is a sensible first choice for ordinary server-delivered HTML. Browser rendering adds setup and page execution; a crawl adds work proportional to the pages visited. These are architectural tradeoffs, not measured CrewAI benchmarks. The reviewed sources do not establish comparative throughput, accuracy, or success rates.
For reliability, make each unit of work small enough to retry, capture errors by URL, and validate the output before continuing. Use bounded retries for temporary network failures, while treating invalid URLs, access denials, and selector mismatches as cases that usually need configuration changes. For crawls, enforce page and time limits, retain the source URL with each record, and deduplicate before aggregation. Pin and document the CrewAI version used by your project; check the matching docs because the sources reviewed here cover versions v1.15.18, v1.15.22, and v1.15.23 as well as an unversioned overview.
Local scraping costs include the compute and operational work needed for your chosen approach; browser installations add dependencies to maintain. Hosted extraction services require their own API credentials and may have separate pricing. No cost or performance comparison among these options was established by the reviewed CrewAI sources, so check provider terms and current pricing for your actual workload. CrewAI’s tool and option names can evolve; install only the integrations needed and verify them against your project version.
6. Or skip the browser setup
If your CrewAI workflow needs a clean visual capture of a page rather than extracted text, [ScreenshotNeo] provides a website screenshot API and MCP server. A single request returns an image or PDF. For example, save a WebP screenshot with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" \
-d access_key=YOUR_API_KEY \
--data-urlencode url=https://example.com \
-o shot.webp
See the [ScreenshotNeo API documentation] for request configuration. The product also accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card required. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. [Sign up free for 1,000 screenshots a month, no card required.]
7. FAQ
Can CrewAI scrape a whole website with one call?
Use the documented Firecrawl crawl integration when the job spans pages. Set scope and page limits. A single-page scraper is not a site crawler.
Does ScrapeWebsiteTool execute JavaScript?
The reviewed page documents HTTP fetching and HTML parsing, not browser rendering. For JavaScript-generated content, choose a browser-based approach or a suitable extraction service.
Should I always involve an agent?
No. Call the tool directly for a known one-page fetch. Use an agent when the scrape is part of a workflow that needs task planning or coordinated outputs.
Can I rely on SeleniumScrapingTool in production?
CrewAI’s documentation labels it currently in development. Check its current status and test it against your target and installed version before relying on it.
How do I keep scraped results auditable?
Store the source URL and retrieval context with each record, and validate required fields before analysis or storage.


