Kotlin Web Scraping: Learn to Extract Data Step by Step
Learn a reliable Kotlin scraping workflow: fetch HTML, parse with jsoup, validate fields, handle pagination, and store structured results.
To scrape a website with Kotlin, use an HTTP client to download the page, parse the returned HTML with an HTML parser, select the fields you need, normalize and validate them, then save explicit Kotlin data objects. On the JVM, a practical combination is Ktor Client for HTTP and jsoup for HTML parsing.
This guide starts with a complete working example, then explains selectors, pagination, JavaScript-rendered pages, reliability, storage, troubleshooting, and operating limits. The examples use a permitted public page and avoid personal or sensitive data. Check a site’s published API, terms, rate limits, and robots.txt before collecting anything.
1. Decide whether ordinary HTML scraping is enough
First request the page and inspect its response body. If the target values appear in that HTML, an HTTP client plus parser is usually simpler than a browser. If the values are inserted later by JavaScript, the initial response may contain only a shell. In that case, look for an official API or data export and assess whether a browser-based approach is permitted and necessary.
| Situation | Starting point |
|---|---|
| Data is present in the response HTML | Ktor (or another HTTP client) plus jsoup |
| You need CSS/XPath selection and DOM traversal on the JVM | jsoup |
| Data appears only after client-side JavaScript runs | Investigate an official API; otherwise evaluate a permitted browser route separately |
| Browser or Node web application development | Kotlin/JS, which targets browser or Node.js environments; this is distinct from a JVM scraper (Kotlin/JS overview) |
| WebAssembly application development | Kotlin/Wasm web targets; it is not automatically the ordinary server-side scraping runtime |
2. Create a Kotlin/JVM project
Ktor supports multiple client platforms, including JVM and JavaScript. jsoup is a Java library and is a direct fit for Kotlin/JVM. Confirm current dependency coordinates and engine support in the official documentation before pinning versions; versions change.
plugins {
kotlin("jvm") version "YOUR_KOTLIN_VERSION"
application
}
repositories {
mavenCentral()
}
dependencies {
implementation("io.ktor:ktor-client-core:YOUR_KTOR_VERSION")
implementation("io.ktor:ktor-client-cio:YOUR_KTOR_VERSION")
implementation("org.jsoup:jsoup:YOUR_JSOUP_VERSION")
implementation("ch.qos.logback:logback-classic:YOUR_LOGBACK_VERSION")
}
application {
mainClass.set("MainKt")
}
Ktor’s client documentation covers engines, requests, responses, plugins, and supported targets: ktor.io/docs/client.html. jsoup’s site and cookbook document URL loading, DOM traversal, CSS selectors, XPath, and attribute extraction: jsoup cookbook.
3. Fetch a page with Ktor
The fetcher below sets an honest identifying User-Agent, applies a request timeout, checks the status and content type, and returns the HTML string. It does not pretend that a successful HTTP response means the page contains the data you want.
import io.ktor.client.HttpClient
import io.ktor.client.engine.cio.CIO
import io.ktor.client.plugins.HttpTimeout
import io.ktor.client.request.get
import io.ktor.client.request.header
import io.ktor.client.statement.HttpResponse
import io.ktor.client.statement.bodyAsText
import io.ktor.http.HttpHeaders
import io.ktor.http.HttpStatusCode
suspend fun fetchHtml(url: String): String {
val client = HttpClient(CIO) {
install(HttpTimeout) {
requestTimeoutMillis = 30_000
connectTimeoutMillis = 10_000
socketTimeoutMillis = 30_000
}
}
try {
val response: HttpResponse = client.get(url) {
header(HttpHeaders.UserAgent, "ExampleKotlinScraper/1.0 (+https://example.com/contact)")
header(HttpHeaders.Accept, "text/html,application/xhtml+xml")
}
if (response.status != HttpStatusCode.OK) {
error("GET $url returned ${response.status}")
}
val contentType = response.headers[HttpHeaders.ContentType].orEmpty()
if (!contentType.contains("text/html", ignoreCase = true) &&
!contentType.contains("application/xhtml+xml", ignoreCase = true)) {
error("Expected HTML from $url but received $contentType")
}
return response.bodyAsText()
} finally {
client.close()
}
}
For a larger job, create one configured client and reuse it rather than opening one per URL. Close it when the job finishes. Add cookies, authorization, proxy settings, or other headers only when you are allowed to send them.
4. Parse HTML and extract structured records
Pass the source URL to jsoup so relative links can be resolved. Inspect the document structure before writing selectors. Prefer stable semantic classes, data attributes, or landmarks over deeply nested positional selectors.
import org.jsoup.Jsoup
import java.net.URI
import java.time.Instant
data class Article(
val title: String,
val url: String?,
val summary: String?,
val retrievedAt: Instant,
val sourcePage: String
)
fun cleanWhitespace(value: String?): String? = value
?.replace(Regex("\\s+"), " ")
?.trim()
?.takeIf { it.isNotEmpty() }
fun parseArticles(html: String, sourceUrl: String): List {
val document = Jsoup.parse(html, sourceUrl)
return document.select("article, .article-card, [data-article]").mapNotNull { node ->
val title = cleanWhitespace(node.selectFirst("h2, h3, [data-title]")?.text())
?: return@mapNotNull null
val link = node.selectFirst("a[href]")?.absUrl("href")
?.takeIf { it.isNotBlank() }
val summary = cleanWhitespace(
node.selectFirst("p, .summary, [data-summary]")?.text()
)
Article(
title = title,
url = link,
summary = summary,
retrievedAt = Instant.now(),
sourcePage = sourceUrl
)
}
}
suspend fun main() {
val pageUrl = "https://example.com/articles"
val html = fetchHtml(pageUrl)
val articles = parseArticles(html, pageUrl)
require(articles.isNotEmpty()) { "No articles found; inspect the HTML and selectors" }
articles.forEach(::println)
}
jsoup does not execute page JavaScript. Its parser handles real-world HTML and exposes DOM, CSS selector, XPath, text, and attribute APIs. The absUrl("href") call resolves a relative link against the document base URL when that base is known.
Extracting common field types
val title = element.selectFirst("h1")?.text()?.trim()
val imageUrl = element.selectFirst("img")?.absUrl("src")
val priceText = element.selectFirst(".price")?.text()
val price = priceText
?.replace(Regex("[^0-9.,-]"), "")
?.replace(",", ".")
?.toBigDecimalOrNull()
val published = element.selectFirst("time")?.attr("datetime")?.trim()
Do not silently turn a missing required field into an empty string. Use nullable properties for optional values and reject or quarantine records that lack required identifiers.
5. A complete small scraper
This example fetches one page, extracts records, validates a required title, and writes JSON-like lines. Replace the selectors with selectors observed in the target page.
import io.ktor.client.*
import io.ktor.client.engine.cio.*
import io.ktor.client.plugins.*
import io.ktor.client.request.*
import io.ktor.client.statement.*
import io.ktor.http.*
import kotlinx.coroutines.runBlocking
import org.jsoup.Jsoup
import java.time.Instant
data class Record(val name: String, val href: String?, val fetchedAt: Instant)
fun main() = runBlocking {
val url = "https://example.com/catalog"
val client = HttpClient(CIO) {
install(HttpTimeout) { requestTimeoutMillis = 30_000 }
}
try {
val response = client.get(url) {
header(HttpHeaders.UserAgent, "ExampleKotlinScraper/1.0 (+https://example.com/contact)")
}
check(response.status.isSuccess()) { "HTTP ${response.status}" }
val html = response.bodyAsText()
val doc = Jsoup.parse(html, url)
val records = doc.select(".product-card").mapNotNull { card ->
val name = card.selectFirst(".product-name")?.text()?.trim()
?.takeIf { it.isNotEmpty() } ?: return@mapNotNull null
val href = card.selectFirst("a[href]")?.absUrl("href")
?.takeIf { it.isNotEmpty() }
Record(name, href, Instant.now())
}
check(records.isNotEmpty()) { "No records matched; verify the selector" }
records.forEach { record ->
val safeName = record.name.replace("\\"", "\\\\\\"")
println("{\\"name\\":\\"$safeName\\",\\"href\\":${record.href?.let { "\\"$it\\"" } ?: "null"}}")
}
} finally {
client.close()
}
}
6. Inspect the document before choosing selectors
- Save or print a small portion of the returned HTML.
- Search for a visible value you expect to extract.
- Confirm whether the value is in the response or added by a script.
- Identify a stable container and then select its child fields.
- Test missing fields and repeated elements, not only the ideal record.
println(html.take(5_000))
println(document.select("article").size)
println(document.select("article").firstOrNull()?.outerHtml())
Useful selector forms include .class-name, #id, [data-id], article h2, a[href], and attribute predicates such as a[href^=/products/]. Keep selectors readable and centralize them so a markup change has one obvious repair point.
7. Pagination, retries, and bounded concurrency
Make one page correct before adding pagination. Derive the next URL from a discovered link where possible instead of guessing page numbers.
suspend fun collectPages(firstUrl: String, maxPages: Int): List {
val all = mutableListOf()
var nextUrl: String? = firstUrl
var page = 0
while (nextUrl != null && page++ < maxPages) {
val html = fetchHtml(nextUrl!!)
val doc = Jsoup.parse(html, nextUrl)
all += parseRecords(doc, nextUrl)
nextUrl = doc.selectFirst("a[rel=next], a.next[href]")?.absUrl("href")
}
return all
}
For transient network failures, retry a small bounded number of times with exponential backoff and jitter. Do not retry every HTTP error: authentication failures, access denials, invalid URLs, and persistent 4xx responses usually require a fix or a stop. Use bounded concurrency, a site-supportable request rate, and caching where repeated retrieval is unnecessary. There is no universal safe requests-per-second value.
8. Validate and store the results
Give every output an explicit schema. Validate required fields, normalize whitespace and dates, preserve the source URL, and record retrieval time. Send invalid records to a review path instead of silently dropping them.
| Output | When it fits | Care point |
|---|---|---|
| JSON | Nested records or API exchange | Escape strings and keep a schema version |
| CSV | Flat tabular exports | Quote commas, newlines, and quotes correctly |
| Database | Incremental jobs and deduplication | Use a stable key and upsert policy |
A useful record often includes sourceUrl, retrievedAt, a stable source identifier, and the parser version. Metrics should show request failures, empty extraction results, validation failures, and record counts so a selector break does not look like a successful empty run.
9. When JavaScript changes the approach
If a browser visibly shows data but bodyAsText() does not contain it, inspect the page’s network activity and look for a documented API or export. Confirm that the API is permitted and that its authentication and rate limits fit your use. A parser cannot execute client-side JavaScript, and this guide does not establish a particular browser automation library or benchmark.
Do not try to evade bot checks, access denials, or rate limits. Stop, contact the site owner, or use an authorized data source. Keep Kotlin/JVM scraping separate from Kotlin/JS and Kotlin/Wasm web application development; those targets have different runtime and library constraints.
10. Respect robots.txt and site rules
RFC 9309 says a crawler that successfully retrieves robots.txt must follow its parseable rules. The same RFC states: “These rules are not a form of access authorization.” Read the full Robots Exclusion Protocol. Terms of service, privacy, copyright, contractual limits, and applicable law require context-specific review; robots.txt alone does not answer those questions.
11. Or skip the browser setup
If your goal is a clean image or PDF of a page rather than structured HTML fields, ScreenshotNeo provides a single screenshot API request. Before capture it accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', data));
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, caching with a chosen TTL, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.
There is a free plan with 1,000 screenshots per month and no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.
12. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 401, 403, or access denied | Authentication, permission, bot protection, or site policy | Check authorization and site rules; do not attempt to bypass a block |
| 429 Too Many Requests | Rate limit | Reduce concurrency, honor retry guidance, add backoff, and cache results |
| Timeout | Slow server, network, or oversized response | Set connect/request/socket timeouts, retry transient failures, and record the URL |
| Parser returns zero elements | Wrong selector or JavaScript-rendered content | Print the response HTML and verify the selector against the returned document |
| Relative links are blank | No base URL supplied | Use Jsoup.parse(html, sourceUrl) and absUrl |
| Numbers parse incorrectly | Currency symbols, locale separators, or hidden text | Normalize deliberately and parse with an explicit locale policy |
| Records silently disappear | mapNotNull hides missing required fields |
Count and log rejected records; send them to a review output |
| HTML is not actually HTML | Redirect, JSON, PDF, or an error page | Check status and Content-Type before parsing |
| Memory grows during a large run | All pages retained in memory | Process and persist page by page; keep only bounded batches |
13. Performance, reliability, and cost notes
- Reuse one configured Ktor client per job and close it cleanly.
- Bound concurrency; more parallel requests can increase failures and violate site limits.
- Cache responses when freshness requirements allow it.
- Retry only transient failures with a maximum attempt count and backoff.
- Measure request latency, status codes, response sizes, extraction counts, and validation failures.
- Store checkpoints for pagination so a failed run can resume without repeating every page.
- No performance benchmark is established by the source material; measure your own target, network, selectors, and storage path.
- For ScreenshotNeo, cache hits are not billed, while clean screenshots are billed according to the selected plan; the response includes
X-Page-VerdictandX-Billedheaders.
14. FAQ
Can I use jsoup directly without Ktor?
Yes. jsoup can connect to and parse a URL for simple cases. Ktor is useful when you need explicit HTTP configuration, reusable clients, timeouts, headers, retries, or broader client-platform choices.
Does jsoup run JavaScript?
No. It parses the HTML it receives. JavaScript-created content requires an API investigation or a separately evaluated browser approach.
Is Kotlin/JS the normal choice for a scraper?
Not necessarily. Kotlin/JS targets browser or Node.js environments. For a JVM backend scraper, Ktor plus jsoup is a direct fit; choose a target and dependencies that match your deployment.
How fast should my scraper request pages?
There is no universal safe rate. Follow the site’s instructions, honor rate-limit responses, use bounded concurrency, and reduce load with caching.
How do I know a selector broke?
Track the number of matched containers and validated records, log rejected fields, and alert when counts fall outside expected bounds. Keep a saved fixture of representative HTML for parser tests.
When should I use ScreenshotNeo?
Use it when you need page screenshots or PDFs without maintaining browser setup, especially when consent banners, popups, chat widgets, lazy images, waits, or device settings matter. It does not replace HTML extraction for structured fields.


