ScreenshotNeo

BlogGuides

Kotlin Web Scraping: Learn to Extract Data Step by Step

Learn a reliable Kotlin scraping workflow: fetch HTML, parse with jsoup, validate fields, handle pagination, and store structured results.

By the ScreenshotNeo team1 October 202610 min read

To scrape a website with Kotlin, use an HTTP client to download the page, parse the returned HTML with an HTML parser, select the fields you need, normalize and validate them, then save explicit Kotlin data objects. On the JVM, a practical combination is Ktor Client for HTTP and jsoup for HTML parsing.

This guide starts with a complete working example, then explains selectors, pagination, JavaScript-rendered pages, reliability, storage, troubleshooting, and operating limits. The examples use a permitted public page and avoid personal or sensitive data. Check a site’s published API, terms, rate limits, and robots.txt before collecting anything.

1. Decide whether ordinary HTML scraping is enough

First request the page and inspect its response body. If the target values appear in that HTML, an HTTP client plus parser is usually simpler than a browser. If the values are inserted later by JavaScript, the initial response may contain only a shell. In that case, look for an official API or data export and assess whether a browser-based approach is permitted and necessary.

Situation Starting point
Data is present in the response HTML Ktor (or another HTTP client) plus jsoup
You need CSS/XPath selection and DOM traversal on the JVM jsoup
Data appears only after client-side JavaScript runs Investigate an official API; otherwise evaluate a permitted browser route separately
Browser or Node web application development Kotlin/JS, which targets browser or Node.js environments; this is distinct from a JVM scraper (Kotlin/JS overview)
WebAssembly application development Kotlin/Wasm web targets; it is not automatically the ordinary server-side scraping runtime

2. Create a Kotlin/JVM project

Ktor supports multiple client platforms, including JVM and JavaScript. jsoup is a Java library and is a direct fit for Kotlin/JVM. Confirm current dependency coordinates and engine support in the official documentation before pinning versions; versions change.

plugins {
    kotlin("jvm") version "YOUR_KOTLIN_VERSION"
    application
}

repositories {
    mavenCentral()
}

dependencies {
    implementation("io.ktor:ktor-client-core:YOUR_KTOR_VERSION")
    implementation("io.ktor:ktor-client-cio:YOUR_KTOR_VERSION")
    implementation("org.jsoup:jsoup:YOUR_JSOUP_VERSION")
    implementation("ch.qos.logback:logback-classic:YOUR_LOGBACK_VERSION")
}

application {
    mainClass.set("MainKt")
}

Ktor’s client documentation covers engines, requests, responses, plugins, and supported targets: ktor.io/docs/client.html. jsoup’s site and cookbook document URL loading, DOM traversal, CSS selectors, XPath, and attribute extraction: jsoup cookbook.

3. Fetch a page with Ktor

The fetcher below sets an honest identifying User-Agent, applies a request timeout, checks the status and content type, and returns the HTML string. It does not pretend that a successful HTTP response means the page contains the data you want.

import io.ktor.client.HttpClient
import io.ktor.client.engine.cio.CIO
import io.ktor.client.plugins.HttpTimeout
import io.ktor.client.request.get
import io.ktor.client.request.header
import io.ktor.client.statement.HttpResponse
import io.ktor.client.statement.bodyAsText
import io.ktor.http.HttpHeaders
import io.ktor.http.HttpStatusCode

suspend fun fetchHtml(url: String): String {
    val client = HttpClient(CIO) {
        install(HttpTimeout) {
            requestTimeoutMillis = 30_000
            connectTimeoutMillis = 10_000
            socketTimeoutMillis = 30_000
        }
    }

    try {
        val response: HttpResponse = client.get(url) {
            header(HttpHeaders.UserAgent, "ExampleKotlinScraper/1.0 (+https://example.com/contact)")
            header(HttpHeaders.Accept, "text/html,application/xhtml+xml")
        }

        if (response.status != HttpStatusCode.OK) {
            error("GET $url returned ${response.status}")
        }

        val contentType = response.headers[HttpHeaders.ContentType].orEmpty()
        if (!contentType.contains("text/html", ignoreCase = true) &&
            !contentType.contains("application/xhtml+xml", ignoreCase = true)) {
            error("Expected HTML from $url but received $contentType")
        }

        return response.bodyAsText()
    } finally {
        client.close()
    }
}

For a larger job, create one configured client and reuse it rather than opening one per URL. Close it when the job finishes. Add cookies, authorization, proxy settings, or other headers only when you are allowed to send them.

4. Parse HTML and extract structured records

Pass the source URL to jsoup so relative links can be resolved. Inspect the document structure before writing selectors. Prefer stable semantic classes, data attributes, or landmarks over deeply nested positional selectors.

import org.jsoup.Jsoup
import java.net.URI
import java.time.Instant

data class Article(
    val title: String,
    val url: String?,
    val summary: String?,
    val retrievedAt: Instant,
    val sourcePage: String
)

fun cleanWhitespace(value: String?): String? = value
    ?.replace(Regex("\\s+"), " ")
    ?.trim()
    ?.takeIf { it.isNotEmpty() }

fun parseArticles(html: String, sourceUrl: String): List
{ val document = Jsoup.parse(html, sourceUrl) return document.select("article, .article-card, [data-article]").mapNotNull { node -> val title = cleanWhitespace(node.selectFirst("h2, h3, [data-title]")?.text()) ?: return@mapNotNull null val link = node.selectFirst("a[href]")?.absUrl("href") ?.takeIf { it.isNotBlank() } val summary = cleanWhitespace( node.selectFirst("p, .summary, [data-summary]")?.text() ) Article( title = title, url = link, summary = summary, retrievedAt = Instant.now(), sourcePage = sourceUrl ) } } suspend fun main() { val pageUrl = "https://example.com/articles" val html = fetchHtml(pageUrl) val articles = parseArticles(html, pageUrl) require(articles.isNotEmpty()) { "No articles found; inspect the HTML and selectors" } articles.forEach(::println) }

jsoup does not execute page JavaScript. Its parser handles real-world HTML and exposes DOM, CSS selector, XPath, text, and attribute APIs. The absUrl("href") call resolves a relative link against the document base URL when that base is known.

Extracting common field types

val title = element.selectFirst("h1")?.text()?.trim()
val imageUrl = element.selectFirst("img")?.absUrl("src")
val priceText = element.selectFirst(".price")?.text()
val price = priceText
    ?.replace(Regex("[^0-9.,-]"), "")
    ?.replace(",", ".")
    ?.toBigDecimalOrNull()
val published = element.selectFirst("time")?.attr("datetime")?.trim()

Do not silently turn a missing required field into an empty string. Use nullable properties for optional values and reject or quarantine records that lack required identifiers.

5. A complete small scraper

This example fetches one page, extracts records, validates a required title, and writes JSON-like lines. Replace the selectors with selectors observed in the target page.

import io.ktor.client.*
import io.ktor.client.engine.cio.*
import io.ktor.client.plugins.*
import io.ktor.client.request.*
import io.ktor.client.statement.*
import io.ktor.http.*
import kotlinx.coroutines.runBlocking
import org.jsoup.Jsoup
import java.time.Instant

data class Record(val name: String, val href: String?, val fetchedAt: Instant)

fun main() = runBlocking {
    val url = "https://example.com/catalog"
    val client = HttpClient(CIO) {
        install(HttpTimeout) { requestTimeoutMillis = 30_000 }
    }

    try {
        val response = client.get(url) {
            header(HttpHeaders.UserAgent, "ExampleKotlinScraper/1.0 (+https://example.com/contact)")
        }
        check(response.status.isSuccess()) { "HTTP ${response.status}" }

        val html = response.bodyAsText()
        val doc = Jsoup.parse(html, url)
        val records = doc.select(".product-card").mapNotNull { card ->
            val name = card.selectFirst(".product-name")?.text()?.trim()
                ?.takeIf { it.isNotEmpty() } ?: return@mapNotNull null
            val href = card.selectFirst("a[href]")?.absUrl("href")
                ?.takeIf { it.isNotEmpty() }
            Record(name, href, Instant.now())
        }

        check(records.isNotEmpty()) { "No records matched; verify the selector" }
        records.forEach { record ->
            val safeName = record.name.replace("\\"", "\\\\\\"")
            println("{\\"name\\":\\"$safeName\\",\\"href\\":${record.href?.let { "\\"$it\\"" } ?: "null"}}")
        }
    } finally {
        client.close()
    }
}

6. Inspect the document before choosing selectors

  1. Save or print a small portion of the returned HTML.
  2. Search for a visible value you expect to extract.
  3. Confirm whether the value is in the response or added by a script.
  4. Identify a stable container and then select its child fields.
  5. Test missing fields and repeated elements, not only the ideal record.
println(html.take(5_000))
println(document.select("article").size)
println(document.select("article").firstOrNull()?.outerHtml())

Useful selector forms include .class-name, #id, [data-id], article h2, a[href], and attribute predicates such as a[href^=/products/]. Keep selectors readable and centralize them so a markup change has one obvious repair point.

7. Pagination, retries, and bounded concurrency

Make one page correct before adding pagination. Derive the next URL from a discovered link where possible instead of guessing page numbers.

suspend fun collectPages(firstUrl: String, maxPages: Int): List {
    val all = mutableListOf()
    var nextUrl: String? = firstUrl
    var page = 0

    while (nextUrl != null && page++ < maxPages) {
        val html = fetchHtml(nextUrl!!)
        val doc = Jsoup.parse(html, nextUrl)
        all += parseRecords(doc, nextUrl)
        nextUrl = doc.selectFirst("a[rel=next], a.next[href]")?.absUrl("href")
    }
    return all
}

For transient network failures, retry a small bounded number of times with exponential backoff and jitter. Do not retry every HTTP error: authentication failures, access denials, invalid URLs, and persistent 4xx responses usually require a fix or a stop. Use bounded concurrency, a site-supportable request rate, and caching where repeated retrieval is unnecessary. There is no universal safe requests-per-second value.

8. Validate and store the results

Give every output an explicit schema. Validate required fields, normalize whitespace and dates, preserve the source URL, and record retrieval time. Send invalid records to a review path instead of silently dropping them.

Output When it fits Care point
JSON Nested records or API exchange Escape strings and keep a schema version
CSV Flat tabular exports Quote commas, newlines, and quotes correctly
Database Incremental jobs and deduplication Use a stable key and upsert policy

A useful record often includes sourceUrl, retrievedAt, a stable source identifier, and the parser version. Metrics should show request failures, empty extraction results, validation failures, and record counts so a selector break does not look like a successful empty run.

9. When JavaScript changes the approach

If a browser visibly shows data but bodyAsText() does not contain it, inspect the page’s network activity and look for a documented API or export. Confirm that the API is permitted and that its authentication and rate limits fit your use. A parser cannot execute client-side JavaScript, and this guide does not establish a particular browser automation library or benchmark.

Do not try to evade bot checks, access denials, or rate limits. Stop, contact the site owner, or use an authorized data source. Keep Kotlin/JVM scraping separate from Kotlin/JS and Kotlin/Wasm web application development; those targets have different runtime and library constraints.

10. Respect robots.txt and site rules

RFC 9309 says a crawler that successfully retrieves robots.txt must follow its parseable rules. The same RFC states: “These rules are not a form of access authorization.” Read the full Robots Exclusion Protocol. Terms of service, privacy, copyright, contractual limits, and applicable law require context-specific review; robots.txt alone does not answer those questions.

11. Or skip the browser setup

If your goal is a clean image or PDF of a page rather than structured HTML fields, ScreenshotNeo provides a single screenshot API request. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));

ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.

There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.

12. Troubleshooting

Symptom Likely cause Fix
401, 403, or access denied Authentication, permission, bot protection, or site policy Check authorization and site rules; do not attempt to bypass a block
429 Too Many Requests Rate limit Reduce concurrency, honor retry guidance, add backoff, and cache results
Timeout Slow server, network, or oversized response Set connect/request/socket timeouts, retry transient failures, and record the URL
Parser returns zero elements Wrong selector or JavaScript-rendered content Print the response HTML and verify the selector against the returned document
Relative links are blank No base URL supplied Use Jsoup.parse(html, sourceUrl) and absUrl
Numbers parse incorrectly Currency symbols, locale separators, or hidden text Normalize deliberately and parse with an explicit locale policy
Records silently disappear mapNotNull hides missing required fields Count and log rejected records; send them to a review output
HTML is not actually HTML Redirect, JSON, PDF, or an error page Check status and Content-Type before parsing
Memory grows during a large run All pages retained in memory Process and persist page by page; keep only bounded batches

13. Performance, reliability, and cost notes

  • Reuse one configured Ktor client per job and close it cleanly.
  • Bound concurrency; more parallel requests can increase failures and violate site limits.
  • Cache responses when freshness requirements allow it.
  • Retry only transient failures with a maximum attempt count and backoff.
  • Measure request latency, status codes, response sizes, extraction counts, and validation failures.
  • Store checkpoints for pagination so a failed run can resume without repeating every page.
  • No performance benchmark is established by the source material; measure your own target, network, selectors, and storage path.
  • For ScreenshotNeo, cache hits are not billed, while clean screenshots are billed according to the selected plan; the response includes X-Page-Verdict and X-Billed headers.

14. FAQ

Can I use jsoup directly without Ktor?

Yes. jsoup can connect to and parse a URL for simple cases. Ktor is useful when you need explicit HTTP configuration, reusable clients, timeouts, headers, retries, or broader client-platform choices.

Does jsoup run JavaScript?

No. It parses the HTML it receives. JavaScript-created content requires an API investigation or a separately evaluated browser approach.

Is Kotlin/JS the normal choice for a scraper?

Not necessarily. Kotlin/JS targets browser or Node.js environments. For a JVM backend scraper, Ktor plus jsoup is a direct fit; choose a target and dependencies that match your deployment.

How fast should my scraper request pages?

There is no universal safe rate. Follow the site’s instructions, honor rate-limit responses, use bounded concurrency, and reduce load with caching.

How do I know a selector broke?

Track the number of matched containers and validated records, log rejected fields, and alert when counts fall outside expected bounds. Keep a saved fixture of representative HTML for parser tests.

When should I use ScreenshotNeo?

Use it when you need page screenshots or PDFs without maintaining browser setup, especially when consent banners, popups, chat widgets, lazy images, waits, or device settings matter. It does not replace HTML extraction for structured fields.