Take screenshots of Indian job listing pages in bulk with Go chromedp
Capture a supplied list of Indian job listing pages with Go and chromedp. Learn setup, readiness waits, screenshot scopes, bulk processing, and recovery.
Use Go’s chromedp package to drive Chrome or Chromium through the Chrome DevTools Protocol (CDP), process an explicit list of job listing URLs, wait for a page-specific readiness condition, and save each screenshot. chromedp is a browser automation library, not a browser: your machine or container must have a compatible Chrome or Chromium binary, or a reachable DevTools endpoint. The examples below capture full pages sequentially, report per-page errors, and continue after failures.
This is a general workflow, not a guide for any one Indian job portal. The target sites, their page markup, access rules, login requirements, and automation behavior have not been verified here. Use URLs you are authorized to access, review each site’s current rules, and adapt readiness selectors for the pages you actually process.
1. Install Go, chromedp, and Chrome or Chromium
Create a module and add chromedp:
mkdir job-page-shots
cd job-page-shots
go mod init example.com/job-page-shots
go get github.com/chromedp/chromedp
Install Chrome or Chromium using the method supported by your operating system or container. The browser executable must be available to chromedp. If it is not on the usual path, set CHROME_BIN in your environment and configure the chromedp allocator with that executable path. For remote execution, use a DevTools WebSocket endpoint instead of launching a local browser.
Record the Go version, the chromedp version pinned in go.mod, and the browser version alongside your deployment configuration. Screenshot and protocol behavior can depend on the selected package and browser versions; check the package documentation and example for the version you pin.
2. Choose what each screenshot should contain
| Capture type | chromedp action | Use it when |
|---|---|---|
| Viewport | chromedp.CaptureScreenshot |
You need only the currently visible browser area. |
| Element | chromedp.Screenshot |
You need one listing card, results region, or other selected node. |
| Scaled element | chromedp.ScreenshotScale |
You need a selected element rendered at a chosen scale. |
| Full document | chromedp.FullScreenshot |
You need content beyond the viewport, such as the complete results page. |
The runnable program below uses full-page PNG screenshots. Full-page capture can override device emulation settings. If you need a fixed viewport, capture the viewport instead and set the viewport dimensions explicitly. Confirm exact action signatures and behavior against the chromedp version in your module.
3. Build a sequential bulk capture program
Save this as main.go. Supply listing URLs as command-line arguments. The program launches one browser, creates an isolated context for each URL, gives each page a deadline, waits for a visible page condition, saves a full-page PNG, and continues to the next URL if a capture fails.
package main
import (
"context"
"fmt"
"net/url"
"os"
"path/filepath"
"strings"
"time"
"github.com/chromedp/chromedp"
)
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: go run . URL [URL ...]")
os.Exit(2)
}
outputDir := "shots"
if err := os.MkdirAll(outputDir, 0o755); err != nil {
fmt.Fprintf(os.Stderr, "create output directory: %v\n", err)
os.Exit(1)
}
// One browser process is reused for this run. Each URL gets a fresh tab context.
browserCtx, cancelBrowser := chromedp.NewContext(context.Background())
defer cancelBrowser()
// The first run starts Chrome. A browser startup deadline helps fail promptly
// when the runtime lacks a compatible browser or required system libraries.
startupCtx, cancelStartup := context.WithTimeout(browserCtx, 30*time.Second)
defer cancelStartup()
if err := chromedp.Run(startupCtx); err != nil {
fmt.Fprintf(os.Stderr, "start browser: %v\n", err)
os.Exit(1)
}
failures := 0
for index, pageURL := range os.Args[1:] {
file := filepath.Join(outputDir, filename(index+1, pageURL)+".png")
if err := capture(browserCtx, pageURL, file); err != nil {
failures++
fmt.Fprintf(os.Stderr, "FAILED url=%q file=%q error=%v\n", pageURL, file, err)
continue
}
fmt.Printf("SAVED url=%q file=%q\n", pageURL, file)
}
if failures > 0 {
fmt.Fprintf(os.Stderr, "completed with %d failed page(s)\n", failures)
os.Exit(1)
}
}
func capture(browserCtx context.Context, pageURL, file string) error {
parsed, err := url.ParseRequestURI(pageURL)
if err != nil || (parsed.Scheme != "https" && parsed.Scheme != "http") || parsed.Host == "" {
return fmt.Errorf("expected an absolute http or https URL")
}
pageCtx, cancel := context.WithTimeout(browserCtx, 60*time.Second)
defer cancel()
var image []byte
err = chromedp.Run(pageCtx,
chromedp.Navigate(pageURL),
// Replace this with a selector that means the listing is ready on your site.
// If page templates vary, choose a wait condition per URL or portal.
chromedp.WaitReady("body", chromedp.ByQuery),
chromedp.FullScreenshot(&image, 90),
)
if err != nil {
return err
}
if len(image) == 0 {
return fmt.Errorf("browser returned an empty screenshot")
}
if err := os.WriteFile(file, image, 0o644); err != nil {
return fmt.Errorf("write screenshot: %w", err)
}
return nil
}
func filename(index int, pageURL string) string {
u, err := url.Parse(pageURL)
host := "page"
if err == nil && u.Hostname() != "" {
host = u.Hostname()
}
// Keep names filesystem-friendly and unique within this run.
host = strings.NewReplacer(".", "_", ":", "_", "/", "_").Replace(host)
return fmt.Sprintf("%03d_%s", index, host)
}
Run it with a small, explicit set of URLs first:
go run . 'https://example.com/jobs' 'https://example.org/careers'
Replace the example URLs with listing pages you are permitted to capture. The filenames are numbered and include the hostname. For repeatable archives, extend the program to write a manifest containing each original URL, timestamp, output path, browser version, and capture result. Avoid putting sensitive query parameters into public logs or manifests.
4. Wait for meaningful page readiness
WaitReady("body") confirms that the body node is ready, but it does not prove that job results have finished rendering. Prefer a selector for a visible heading or listing container that is meaningful on the target portal. Inspect the page markup for the URLs you use and verify the selector after layout changes.
For example, replace the body wait with a site-specific condition:
chromedp.WaitVisible("main .job-results", chromedp.ByQuery)
The selector above is illustrative; it is not verified for any particular Indian job site. If multiple portals are in the input, represent each URL with its own wait selector or site configuration. A missing selector should produce a clear per-page error and leave the rest of the batch eligible to continue.
A fixed sleep can be useful for a known animation or delayed component, but it is usually less reliable than waiting for a state that indicates the content you need is present. Pages may render listings asynchronously, defer images, show consent dialogs, or change markup. If a screenshot must include lazy-loaded content, determine whether scrolling through the page is needed before full-page capture; behavior is site and page dependent.
5. Adapt capture scope and browser state
Capture the current viewport
Replace the full-page action with chromedp.CaptureScreenshot(&image) to capture the visible viewport. Set the desired viewport dimensions with the relevant chromedp emulation action before navigation. This is useful when you need consistent-sized images rather than a tall full-page document.
Capture a listing element
Use chromedp.Screenshot(selector, &image, chromedp.NodeVisible) for a selected element. Ensure the selector uniquely identifies the intended region and wait for it to become visible before capture. If the selector matches multiple nodes, choose a stable, specific selector or target the desired node explicitly.
Authentication and cookies
The example does not sign in or manage credentials. If a page requires authentication, use only an account and method you are authorized to automate. Avoid hard-coding passwords or session cookies; inject secrets through a secure runtime mechanism, limit their access, and do not print them. Confirm the portal’s current access requirements before automating authenticated pages.
Browser and network configuration
For a remote Chrome instance, connect to a DevTools endpoint using chromedp’s remote allocator rather than starting a local binary. Treat the endpoint as a credential: anyone who can access it may control the browser session. Restrict network access and avoid logging its token or full URL.
Use a new page context per URL to isolate navigation state while reusing the browser process. If you need a specific locale, timezone, viewport, or user agent, configure it deliberately and record it in the capture manifest. Do not claim a page was captured under a particular device profile unless you set and preserve that configuration.
6. Run larger batches safely
The example processes URLs sequentially. This is a good baseline for debugging and limits simultaneous browser work and requests to third-party sites. Begin with a small batch and check output before processing a large list.
- Prepare an explicit input file or URL list. Do not implicitly crawl additional pages.
- Validate that each entry is an absolute HTTP or HTTPS URL and reject malformed rows.
- Set a per-page deadline and decide whether the batch should continue after an individual error.
- Use stable output naming and write a manifest linking each image to its source URL.
- Retry only errors that are plausibly transient, with a small retry limit and backoff. Do not retry indefinitely.
- Only add bounded concurrency after measuring a real bottleneck. Concurrent tabs increase CPU and memory use and send more simultaneous requests to the sites.
The supplied code returns a nonzero process status if any page failed, while continuing through the list. This makes it suitable for scheduled jobs that should preserve successful captures and still alert an operator to partial failure. For workflows where any failed page should stop the run, return immediately instead of continuing.
7. Troubleshooting
| Symptom | Likely cause | What to change |
|---|---|---|
| Chrome executable not found | No compatible Chrome or Chromium is installed, or it is outside the expected path. | Install a browser in the runtime or configure chromedp to use the actual executable path. |
| Browser exits during startup | Missing system libraries, incompatible browser build, or container restrictions. | Check the browser’s own startup error, install its runtime dependencies, and verify the browser can launch in that environment. |
| Navigation or context deadline exceeded | Slow page, stalled request, inaccessible URL, or deadline shorter than page load. | Check the URL and network access, set a realistic deadline, and wait for a meaningful selector rather than requiring unrelated resources to finish. |
| Screenshot is blank or shows an error page | The page did not render the expected content, navigation failed, or the site returned a challenge or error state. | Inspect the result and logs, verify the URL manually, and treat the page as a failed capture instead of assuming the PNG is valid. |
| Listing content is missing | The body loaded before client-rendered results, a selector changed, or lazy content was not triggered. | Wait for a listing-specific visible selector; verify it against current markup and scroll if the required content is lazy-loaded. |
| Element screenshot fails | The selector does not match, matches a hidden node, or is ambiguous. | Inspect the page, select a stable visible node, and wait for it before capturing. |
| Output file is empty or missing | Capture action failed, the returned byte slice was empty, or file writing failed. | Check the returned error, ensure the output directory is writable, and retain the empty-byte check. |
| Portal shows a consent prompt or bot check | The page requires user interaction or applies an access control. | Follow the portal’s current rules and supported access flow. Do not attempt to evade access controls; record the page as unavailable if it cannot be captured appropriately. |
8. Performance, reliability, and cost
For modest batches, one browser process with sequential tabs is simpler to operate and easier to diagnose. Full-page images may be large, and rendering many pages can consume substantial memory and CPU. Keep batches bounded, close per-page contexts promptly, and monitor the machine running Chrome. No performance benchmark or safe concurrency value has been established for these portals.
Set deadlines that match the workload, log failures with URL identifiers, and preserve successful files when another page fails. A browser process can crash or a remote endpoint can disconnect; for scheduled work, persist the input list and completion manifest so the run can resume deliberately. Avoid automatic retries for persistent access denials, bot checks, malformed URLs, or selector mismatches.
With local chromedp, costs depend on the compute and browser runtime you operate; this guide makes no estimate. You are responsible for staying within the target sites’ rules and avoiding excessive request volume.
Or skip the browser setup
ScreenshotNeo takes a screenshot from one GET request, with PNG, JPEG, WebP, and PDF output options. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every plan includes every feature; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000.
Install the Go HTTP client dependency if needed: go get github.com/go-resty/resty/v2. This complete Go example reads one URL per line from urls.txt, requests a WebP screenshot for each, and saves the response. Create an API key in your ScreenshotNeo account and provide it in the environment as SCREENSHOTNEO_API_KEY.
package main
import (
"bufio"
"fmt"
"net/http"
"net/url"
"os"
"path/filepath"
"strings"
"github.com/go-resty/resty/v2"
)
func main() {
key := os.Getenv("SCREENSHOTNEO_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "set SCREENSHOTNEO_API_KEY")
os.Exit(2)
}
input, err := os.Open("urls.txt")
if err != nil {
fmt.Fprintln(os.Stderr, "open urls.txt:", err)
os.Exit(1)
}
defer input.Close()
if err := os.MkdirAll("shots", 0o755); err != nil {
fmt.Fprintln(os.Stderr, "create shots directory:", err)
os.Exit(1)
}
client := resty.New().SetTimeout(90 * time.Second)
scanner := bufio.NewScanner(input)
line := 0
failures := 0
for scanner.Scan() {
pageURL := strings.TrimSpace(scanner.Text())
if pageURL == "" {
continue
}
line++
parsed, parseErr := url.ParseRequestURI(pageURL)
if parseErr != nil || (parsed.Scheme != "https" && parsed.Scheme != "http") || parsed.Host == "" {
fmt.Fprintf(os.Stderr, "line %d: invalid absolute HTTP(S) URL\n", line)
failures++
continue
}
resp, reqErr := client.R().
SetQueryParams(map[string]string{"access_key": key, "url": pageURL}).
Get("https://api.screenshotneo.com/v1/shot")
if reqErr != nil {
fmt.Fprintf(os.Stderr, "line %d request: %v\n", line, reqErr)
failures++
continue
}
if resp.StatusCode() < 200 || resp.StatusCode() >= 300 {
fmt.Fprintf(os.Stderr, "line %d: HTTP %d: %s\n", line, resp.StatusCode(), resp.String())
failures++
continue
}
name := fmt.Sprintf("%03d.webp", line)
if err := os.WriteFile(filepath.Join("shots", name), resp.Body(), 0o644); err != nil {
fmt.Fprintf(os.Stderr, "line %d write: %v\n", line, err)
failures++
continue
}
fmt.Printf("saved line=%d file=%s verdict=%s billed=%s\n", line, name, resp.Header().Get("X-Page-Verdict"), resp.Header().Get("X-Billed"))
}
if err := scanner.Err(); err != nil {
fmt.Fprintln(os.Stderr, "read urls.txt:", err)
failures++
}
if failures > 0 {
os.Exit(1)
}
}
Add "time" to the imports in that snippet for the client timeout. For the API parameters and additional capture options, see the ScreenshotNeo API documentation. The one-call forms are also available in cURL, Python, and Node.js:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo is made by Yorker Media. Use the ScreenshotNeo website to review the service. Start with 1,000 free screenshots a month, no card required: create a free ScreenshotNeo account.
Frequently asked questions
Does chromedp include Chrome?
No. chromedp drives a separately installed Chrome or Chromium browser, or connects to a remote DevTools endpoint.
Can this program discover job listings automatically?
No. It captures the URLs you pass in. Discovery and crawling require separate decisions about scope and each portal’s current rules.
Should I use full-page or viewport screenshots?
Use full-page capture when the whole document is the artifact you need; use viewport capture for fixed-size images of the visible area. Use element capture when only a particular listing region matters.
Can I run several captures at once?
Yes, but first establish that sequential processing is too slow. Bound concurrency, observe resource use, and consider the load sent to each site; there is no verified universal concurrency limit for this use case.
Will the same selector work across Indian job portals?
Usually, selectors depend on each site’s own markup. Configure and periodically verify readiness and element selectors per portal.


