ScreenshotNeo

BlogEngineering

How to Build a Web Crawler in Java: A Breadth-First Crawler with HttpClient and Jsoup

Build a bounded breadth-first Java crawler with HttpClient and Jsoup, including URL normalization, robots.txt, retries, limits, and troubleshooting.

By the ScreenshotNeo team29 September 202610 min read

How to Build a Web Crawler in Java: A Breadth-First Crawler with HttpClient and Jsoup

Direct answer: a breadth-first web crawler keeps a FIFO queue of URLs to visit and a set of canonical URLs already seen. It removes the next URL from the head of the queue, fetches it with one reusable Java HttpClient, parses the HTML with Jsoup, resolves eligible links against the page URL, and appends unseen links to the queue. Enqueuing at the tail and processing from the head is what creates breadth-first order; neither HttpClient nor Jsoup provides that algorithm automatically.

This tutorial targets Java 21 (the API is available from Java 11 onward) and a small, explicitly selected set of public HTTP(S) pages. It includes host and path boundaries, timeouts, status and content-type checks, URL normalization, a page limit, conservative delays, per-page error handling, and a starter robots.txt policy. A production crawler needs durable state, scheduling, retry budgets, metrics, and more complete robots parsing.

What you will build

  • A reusable HttpClient with an explicit redirect policy.
  • A FIFO ArrayDeque frontier and a visited set of canonical URLs.
  • Relative-link resolution and filtering to one host and path prefix.
  • HTML parsing and link extraction with Jsoup.
  • Per-request timeout, maximum pages, response-size guard, user-agent, and delay.
  • A conservative robots.txt check before fetching pages.

Java’s HttpClient is immutable after construction and intended to be reused for multiple requests. It supports synchronous send and asynchronous sendAsync; its default redirect policy is NEVER, so this example opts into normal redirects deliberately. See the Java 21 HttpClient API.

Project setup

Create a Java 21 project and add Jsoup. The official Jsoup site listed 1.23.2 on the research date; check the download page for the current release before pinning a new application.

A breadth-first crawler dequeues one URL, parses its HTML, and appends unseen links to the frontier.
A breadth-first crawler dequeues one URL, parses its HTML, and appends unseen links to the frontier.
<dependency>
  <groupId>org.jsoup</groupId>
  <artifactId>jsoup</artifactId>
  <version>1.23.2</version>
</dependency>

Jsoup can fetch and parse in one call through its Connection API. This article uses direct HttpClient requests and then passes the response body to Jsoup, because that makes queueing, status checks, content-type checks, and request policy visible. Jsoup’s cookbook documents both Jsoup.connect(url).get() and HTTP/HTTPS loading patterns.

Complete breadth-first crawler

Save this as BreadthFirstCrawler.java. It starts at one URL, stays on the same host and path prefix, visits at most 20 pages, and prints each discovered canonical URL. The robots implementation is intentionally small: it reads plain Disallow rules for the selected user-agent group. Replace it with a tested robots parser when you need wildcard, Allow, sitemap, caching, or multiple-origin support.

import java.io.IOException;
import java.net.URI;
import java.net.URISyntaxException;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.ArrayDeque;
import java.util.HashSet;
import java.util.Locale;
import java.util.Set;

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;
import org.jsoup.nodes.Element;
import org.jsoup.select.Elements;

public class BreadthFirstCrawler {
    private static final String USER_AGENT = "ExampleResearchCrawler/1.0 (+https://example.com/contact)";
    private static final int MAX_PAGES = 20;
    private static final int MAX_BODY_BYTES = 2_000_000;
    private static final Duration REQUEST_TIMEOUT = Duration.ofSeconds(15);
    private static final Duration DELAY_BETWEEN_REQUESTS = Duration.ofMillis(750);

    private final HttpClient client = HttpClient.newBuilder()
            .followRedirects(HttpClient.Redirect.NORMAL)
            .connectTimeout(Duration.ofSeconds(10))
            .build();
    private final ArrayDeque<URI> frontier = new ArrayDeque<>();
    private final Set<String> visited = new HashSet<>();
    private final Set<String> disallowedPaths = new HashSet<>();
    private final String allowedHost;
    private final String allowedPathPrefix;

    public BreadthFirstCrawler(URI start) {
        this.allowedHost = start.getHost();
        this.allowedPathPrefix = start.getPath().isBlank() ? "/" : start.getPath();
        frontier.add(canonicalize(start));
        loadRobots(start);
    }

    public void crawl() {
        int pages = 0;
        while (!frontier.isEmpty() && pages < MAX_PAGES) {
            URI url = frontier.removeFirst();
            String key = url.toString();
            if (!visited.add(key) || !allowed(url) || blockedByRobots(url)) continue;

            try {
                HttpRequest request = HttpRequest.newBuilder(url)
                        .timeout(REQUEST_TIMEOUT)
                        .header("User-Agent", USER_AGENT)
                        .header("Accept", "text/html,application/xhtml+xml")
                        .GET()
                        .build();
                HttpResponse<String> response = client.send(
                        request, HttpResponse.BodyHandlers.ofString());

                String contentType = response.headers().firstValue("content-type").orElse("");
                if (response.statusCode() / 100 != 2) {
                    System.err.println("Skip " + url + ": HTTP " + response.statusCode());
                    continue;
                }
                if (!contentType.toLowerCase(Locale.ROOT).contains("text/html")) {
                    System.err.println("Skip " + url + ": content type " + contentType);
                    continue;
                }
                String html = response.body();
                if (html.getBytes(java.nio.charset.StandardCharsets.UTF_8).length > MAX_BODY_BYTES) {
                    System.err.println("Skip " + url + ": response too large");
                    continue;
                }

                Document document = Jsoup.parse(html, url.toString());
                System.out.printf("%d %s%n", ++pages, url);
                discoverLinks(document, url);
            } catch (InterruptedException e) {
                Thread.currentThread().interrupt();
                return;
            } catch (IOException | RuntimeException e) {
                System.err.println("Failed " + url + ": " + e.getMessage());
            }

            sleepBetweenRequests();
        }
    }

    private void discoverLinks(Document document, URI pageUri) {
        Elements anchors = document.select("a[href]");
        for (Element anchor : anchors) {
            String raw = anchor.attr("href").trim();
            if (raw.isEmpty() || raw.startsWith("#") || raw.startsWith("mailto:")
                    || raw.startsWith("tel:") || raw.startsWith("javascript:")) continue;
            try {
                URI resolved = pageUri.resolve(raw);
                URI canonical = canonicalize(resolved);
                if (allowed(canonical) && !blockedByRobots(canonical)
                        && !visited.contains(canonical.toString())) {
                    frontier.addLast(canonical);
                }
            } catch (IllegalArgumentException e) {
                System.err.println("Bad link on " + pageUri + ": " + raw);
            }
        }
    }

    private boolean allowed(URI uri) {
        return ("http".equalsIgnoreCase(uri.getScheme())
                || "https".equalsIgnoreCase(uri.getScheme()))
                && allowedHost.equalsIgnoreCase(uri.getHost())
                && uri.getPath().startsWith(allowedPathPrefix);
    }

    private URI canonicalize(URI input) {
        try {
            String scheme = input.getScheme().toLowerCase(Locale.ROOT);
            String host = input.getHost().toLowerCase(Locale.ROOT);
            int port = input.getPort();
            if (("http".equals(scheme) && port == 80) || ("https".equals(scheme) && port == 443)) port = -1;
            String path = input.getPath().isEmpty() ? "/" : input.getPath();
            while (path.contains("//")) path = path.replace("//", "/");
            return new URI(scheme, null, host, port, path, input.getQuery(), null);
        } catch (URISyntaxException e) {
            throw new IllegalArgumentException("Cannot canonicalize " + input, e);
        }
    }

    private void loadRobots(URI start) {
        URI robots = URI.create(start.getScheme() + "://" + start.getAuthority() + "/robots.txt");
        try {
            HttpRequest request = HttpRequest.newBuilder(robots).timeout(REQUEST_TIMEOUT)
                    .header("User-Agent", USER_AGENT).GET().build();
            HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
            if (response.statusCode() / 100 != 2) return;
            boolean applies = false;
            for (String line : response.body().split("\\R")) {
                String clean = line.split("#", 2)[0].trim();
                if (clean.isEmpty()) continue;
                String lower = clean.toLowerCase(Locale.ROOT);
                if (lower.startsWith("user-agent:")) {
                    applies = lower.substring(11).trim().equals("*")
                            || lower.substring(11).trim().equalsIgnoreCase("exampleresearchcrawler");
                } else if (applies && lower.startsWith("disallow:")) {
                    String path = clean.substring(clean.indexOf(':') + 1).trim();
                    if (!path.isEmpty()) disallowedPaths.add(path);
                }
            }
        } catch (Exception e) {
            System.err.println("Could not read robots.txt: " + e.getMessage());
        }
    }

    private boolean blockedByRobots(URI uri) {
        return disallowedPaths.stream().anyMatch(uri.getPath()::startsWith);
    }

    private void sleepBetweenRequests() {
        try { Thread.sleep(DELAY_BETWEEN_REQUESTS.toMillis()); }
        catch (InterruptedException e) { Thread.currentThread().interrupt(); }
    }

    public static void main(String[] args) {
        URI start = URI.create(args.length == 0 ? "https://example.com/" : args[0]);
        new BreadthFirstCrawler(start).crawl();
    }
}

How the algorithm works

  1. Seed: the start URI is canonicalized and placed in frontier.
  2. Dequeue: removeFirst() takes the oldest pending URL.
  3. Guard: the visited set, scheme, host, path prefix, robots rules, page limit, and content checks decide whether to fetch.
  4. Fetch: one reusable HttpClient sends a bounded-time request with a descriptive user-agent.
  5. Parse: Jsoup builds a document from the returned HTML.
  6. Discover: anchor hrefs are resolved against the current page, normalized, filtered, and appended with addLast().

Suppose the seed links to A and B, and A links to C. The queue becomes A, B after the seed; C is added behind B, so the order is seed, A, B, C. A stack would produce depth-first behavior instead.

URL normalization and scope

Canonicalization removes fragments, lowercases scheme and host, removes default ports, supplies a root slash, and collapses repeated slashes. It does not prove that two URLs serve identical content. Query parameters can still create thousands of distinct URLs; add an allowlist, parameter policy, or per-path query budget when crawling a site with tracking links or calendars.

Apply host and path restrictions before queueing and again before fetching. A relative link can resolve outside the intended site, and a redirect can lead elsewhere. For stricter isolation, reject redirects whose final URI is outside the allowed origin, or use a custom redirect strategy that validates each hop.

Robots.txt and responsible pacing

RFC 9309 asks crawlers to follow parseable robots.txt rules, but states: “These rules are not a form of access authorization.” Robots.txt belongs at the origin’s top-level /robots.txt; match the applicable user-agent group after a successful retrieval. A failure to retrieve it should be an explicit operator decision, often fail-closed for a polite crawler. The sample uses a conservative sequential delay and a descriptive product token. Crawl-delay is prudent operator guidance where published, not a universal RFC 9309 directive.

Do not use robots.txt as permission to access private or restricted resources. Authentication, authorization, terms, and applicable law remain separate concerns.

Useful variations

Jsoup’s integrated connection

Document document = Jsoup.connect("https://example.com/docs")
        .userAgent("ExampleResearchCrawler/1.0 (+https://example.com/contact)")
        .timeout(15_000)
        .followRedirects(true)
        .get();

This is shorter and convenient for a small script. Direct HttpClient plus Jsoup gives you explicit status handling, headers, body limits, and one shared client.

Asynchronous fetching

sendAsync can increase throughput, but it does not create a crawler scheduler. Add a bounded executor, per-host concurrency limit, queue backpressure, timeout handling, retry budgets, and politeness delays. Launching one future per discovered URL can exhaust memory and overwhelm a host.

Persistent crawls

For a crawl that must resume, store frontier entries, visited keys, response metadata, retry count, and timestamps in a durable database or log. In-memory collections are appropriate for the bounded example only.

Performance, reliability, and cost notes

  • Connection reuse: build one HttpClient and reuse it; creating a client per URL commonly prevents connection reuse.
  • Latency: sequential crawling is easy to reason about but limited by request latency and your delay. Concurrency improves throughput only when host limits and backoff are explicit.
  • Memory: the queue and visited set grow with discovered URLs. Enforce page, URL, body, and query limits.
  • Reliability: classify timeouts, connection failures, DNS errors, non-2xx responses, redirects, oversized bodies, and parse errors separately. Retry transient failures with exponential backoff and a cap; do not retry every 4xx response.
  • Cost: a self-hosted crawler spends compute, bandwidth, storage, and operator time. Estimate from pages, average response size, retry rate, and concurrency before expanding scope.

Troubleshooting

Symptom Cause Fix
No links are discovered The response is not HTML, links are JavaScript-generated, or anchors lack hrefs. Check content-type and status. A static HttpClient fetch cannot execute JavaScript; use a browser renderer for that requirement.
Everything is skipped after redirects The redirect target leaves the host or path scope. Log the final URI and decide whether to allow that origin explicitly.
Repeated URLs consume the page limit Tracking queries, fragments, or inconsistent trailing slashes create distinct keys. Strip fragments, define query allowlists, and normalize trailing slashes for your site.
HTTP 403 or 429 The site blocks the user-agent or rate is too high. Identify the crawler, slow down, honor Retry-After where present, and stop when access is not permitted.
Socket timeout Slow server, network path, or too-short timeout. Use separate connect and request budgets, retry only transient failures, and keep a maximum retry count.
OutOfMemoryError Unbounded frontier or response bodies. Set page, URL, queue, and body limits; persist state and stream or reject oversized responses.
Malformed URL exception Broken href values or unsupported schemes. Resolve in a try/catch, accept only HTTP(S), and log the source page and raw value.

Or skip the browser setup

If your goal is collecting clean page images while crawling, ScreenshotNeo provides a single screenshot API request instead of maintaining a browser renderer. Cookie and consent banners are accepted and removed before capture, along with 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

ScreenshotNeo removes common consent banners, popups, and chat widgets before capture.
ScreenshotNeo removes common consent banners, popups, and chat widgets before capture.

See the ScreenshotNeo API documentation for all options, including full-page capture, CSS selectors, device presets, custom headers and cookies, JavaScript, waiting rules, blocking, caching, signed links, async jobs, bulk capture, and usage reporting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots a month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Why is this crawler breadth-first?

Because it removes from the queue head and appends discovered URLs at the tail. Change those data-structure operations and you change the traversal order.

Can HttpClient execute JavaScript?

No. It downloads HTTP responses. JavaScript-rendered links require a browser automation system or a service that renders pages.

Should I crawl multiple hosts with this class?

No. Add a per-origin scheduler, separate robots policies, host-specific delays, concurrency limits, and durable state before expanding beyond one explicit scope.

Does a successful robots.txt fetch grant access?

No. Robots rules are crawl instructions, not authorization. Authentication and access controls still apply.

When should I use Jsoup.connect instead?

Use it for a compact fetch-and-parse script. Keep direct HttpClient when you need detailed request policy, response validation, shared connection management, or custom scheduling.