ScreenshotNeo

BlogHow-to

How to Download All Audio Files from a Website

Learn how to collect linked audio with Wget or HTTrack, find script-loaded media in DevTools, and handle streams, permissions, filters, and failures.

By the ScreenshotNeo team1 October 20268 min read

Short answer: there is no tool that can guarantee every audio file on every website. For audio files exposed as ordinary links, crawl a limited site or directory with HTTrack or GNU Wget, then filter the results by extension. If a page creates audio requests with JavaScript or only after a click, use the browser’s Network panel to inspect those requests. For a supported media platform, use yt-dlp and its documented audio options.

Before downloading, confirm that you own the files, have permission to copy them, or that the site explicitly permits downloads. Keep the scope narrow, respect robots.txt and site terms, use reasonable request rates, and make sure you have enough local storage.

Choose the method that matches the site

Site or media layout Best starting point What it can miss
Audio URLs appear in crawlable HTML or CSS Wget or HTTrack with a host/path scope and audio filters Files created only after scripts run, authentication, temporary URLs
Audio appears after pressing Play or another control Chrome DevTools Network panel Requests that never occur during your session; segmented streams may not be one file
Supported video or media platform yt-dlp with an audio format selection Unsupported pages, changed extractors, formats the source does not expose

1. Define a safe collection scope

  1. Write down the starting URL and the intended host or directory, such as https://example.com/podcasts/.
  2. Decide whether you need files linked from one page, a directory, or a set of media pages. Do not begin with an unlimited whole-site crawl.
  3. Check the site’s download permission, terms, and robots.txt. A refusal or exclusion is a boundary to respect.
  4. Estimate storage. Audio files can be large, and recursive tools may follow more pages than expected.
  5. Run a small trial first. Save logs and inspect several downloaded files before expanding the scope.

2. Download linked audio with GNU Wget

Wget can follow links recursively, rebuild a local directory tree, and accept or reject file types. Its manual describes recursion, depth limits, and file-type filters in the recursive download documentation and file type documentation.

Start with a small, same-host crawl

wget --recursive --level=2 --no-parent --no-host-directories \
  --domains example.com \
  --accept=mp3,m4a,ogg,wav,flac,aac,opus \
  --directory-prefix=audio-download \
  https://example.com/podcasts/
  • --recursive follows links found in retrieved pages.
  • --level=2 limits link depth; increase it only after reviewing the trial.
  • --no-parent prevents moving above the starting directory.
  • --domains example.com keeps links on the intended host. Add an approved asset host only when necessary.
  • --accept limits saved files to common audio extensions. Add an extension only when you know the site uses it.
  • --directory-prefix keeps the result in a dedicated folder.

Wget obeys robots.txt by default. Keep that behavior. Avoid flags that disable robots rules or broaden the crawl without a clear reason.

Limit requests and resume a collection

wget --recursive --level=1 --no-parent --wait=2 --random-wait \
  --timeout=30 --tries=3 --continue \
  --accept=mp3,m4a,ogg,wav,flac,aac,opus \
  --directory-prefix=audio-download \
  https://example.com/audio/

--wait and --random-wait reduce burst traffic, --timeout and --tries handle transient failures, and --continue resumes partial files when the server supports range requests. Recursive retrieval can consume disk space and place load on a server, so monitor both.

3. Mirror a linked site area with HTTrack

HTTrack mirrors a website recursively, preserves a local browsing structure, and can resume or update a mirror. Its command guide documents scope, filters, logs, and defaults. A basic same-host mirror is:

httrack "https://example.com/podcasts/" \
  -O "./site-mirror" \
  "+example.com/podcasts/*" \
  "-*/?*" \
  -v

After the mirror finishes, find likely audio files:

find site-mirror -type f \( \
  -iname '*.mp3' -o -iname '*.m4a' -o -iname '*.ogg' -o \
  -iname '*.wav' -o -iname '*.flac' -o -iname '*.aac' -o -iname '*.opus' \) \
  -print

HTTrack’s filters are specific to HTTrack; do not copy them into a Wget command. Review its logs for URLs that were refused, redirected, or filtered. The tool identifies itself as HTTrack, obeys robots.txt, and its documented default throttle is about 100 KB/s; exact behavior can vary by release.

4. Find audio loaded by JavaScript in Chrome DevTools

A crawler only sees what it can reach and parse. A page may request an audio file after a button click, after a player initializes, or after an API response supplies a URL. Chrome’s Network overview explains recording requests and the Network features reference covers filters, headers, initiators, and timing.

  1. Open the page in Chrome and press F12 or Ctrl/Cmd + Shift + I.
  2. Choose Network, enable recording, and turn on Preserve log if navigation is involved.
  3. Reload the page, then use the page’s normal controls to reveal or play the audio.
  4. Filter by Media, or search for extensions such as .mp3, .m4a, .ogg, or terms such as audio and manifest.
  5. Select a request and inspect its URL, response headers, initiator, status, and timing. Use Copy link address or Copy as cURL only when access and site terms permit.
  6. Repeat the interaction for each page or control you are authorized to inspect. The panel records activity in that browser session; it is not a whole-site inventory.

If the request is a playlist or manifest such as HLS or DASH, the page may be delivering segments instead of one ordinary audio file. Record the request type and headers before deciding whether the site’s authorized export or API is the appropriate route.

5. Use yt-dlp for supported media platforms

yt-dlp is designed for supported extractors and documents audio extraction and audio-only format selection. It is not a universal website crawler.

yt-dlp -F "https://media.example.com/item/123"
yt-dlp -f bestaudio "https://media.example.com/item/123"

Use -F first to see formats the page exposes. Select an audio format that actually exists; do not assume a particular codec or container. For a supported collection URL, follow yt-dlp’s current documentation for playlist scope and output templates, and check the site’s permission before saving copies.

6. Verify that the collection is complete enough

  • Count files by extension and compare the count with the pages or entries you intended to process.
  • Open samples from the beginning, middle, and end of the collection.
  • Check file sizes and response status; an HTML error page can be saved with an audio-looking filename.
  • Review crawler logs for refused, redirected, filtered, or failed URLs.
  • Keep a list of pages that require login, a click, a consent action, or a separate authorized export.
  • Record the date and scope of the collection so a later update can be compared with the first run.

7. Troubleshooting missing or unusable files

Symptom Likely cause Fix
Only the starting page was downloaded The page has few crawlable links, depth is too low, or filters exclude linked pages. Read the log, confirm the start page links to the target pages, raise depth gradually, and keep the host/path scope explicit.
Expected pages are outside the mirror They are on another host, above the starting directory, or blocked by a scope rule. Confirm the relationship and add only the approved host or path. Do not widen the crawl indiscriminately.
JavaScript-heavy pages are incomplete The crawler does not execute the interactions that create media requests. HTTrack notes that intensive JavaScript sites may be incomplete. Use DevTools while loading and using the page, then follow an authorized direct URL, API, or export.
A downloaded file is an HTML error page The request returned a login page, denial page, redirect, or error document. Inspect status and Content-Type; authenticate through an allowed method or request an official download.
The player URL is a shortcut or stream The page exposes a manifest, playlist, or segmented requests rather than a standalone file. Inspect request type and headers. Do not assume every stream can be saved as one file; use the platform’s permitted export.
Access is denied or robots excludes the path The owner has restricted automated retrieval. Stop at that boundary and seek permission, an API, or an official download.
The server slows down or blocks requests The crawl is too broad or too fast. Reduce depth, add delays, narrow filters, and run during an appropriate maintenance window if the owner permits it.
The disk fills during a crawl Recursive retrieval followed more pages or larger files than expected. Stop the job, remove unrelated output, check free space, and restart with a narrower path and extension filter.
Wget resumes incorrectly The server does not support ranges or the partial file changed. Remove the partial file and retry without --continue, then compare checksums or sizes.

Or skip the browser setup

If your goal is to document which pages contain media, create page previews, or automate visual checks around an audio collection, ScreenshotNeo provides a website screenshot API and MCP server. It does not download audio; it captures the page so you can inspect or archive its visual state without maintaining a browser.

One GET request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/podcasts/ -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/podcasts/"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/podcasts/' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Performance, reliability, and cost planning

  • Scope controls: A directory crawl with a depth limit is easier to estimate than a whole-domain crawl.
  • Request load: Add delays and avoid parallel downloads unless the owner allows them. A retry can multiply traffic.
  • Reliability: Save logs, resume only when supported, and rerun a narrow failed scope instead of repeating everything.
  • Storage: Reserve space for the files plus temporary partial downloads and logs. Check free space during long jobs.
  • Dynamic media: Browser observation is more complete for session-triggered requests, but it requires repeating the relevant interactions.
  • Cost: Wget, HTTrack, DevTools, and yt-dlp are software workflows; your main operational costs are bandwidth, storage, and any authorized platform or API fees. ScreenshotNeo’s free tier covers 1,000 page shots monthly, with paid plans beginning at $5 for 3,000.

FAQ

Can I guarantee that I downloaded every audio file?

No. A crawler can only follow exposed links, and browser observation only records requests made during the session. Authentication, JavaScript, temporary URLs, and segmented streams can hide or change resources.

Should I use Wget or HTTrack?

Use Wget when you want explicit recursion, depth, and file-type controls. Use HTTrack when you want a resumable local mirror and its site-oriented filtering and logs.

Why did filtering by .mp3 miss files?

The site may use another extension, extensionless URLs, a manifest, or a response whose URL does not reveal its media type. Inspect response headers and browser requests before changing filters.

Can yt-dlp download any website’s audio?

No. It works through supported extractors, and extractor behavior and available formats can change. Treat it as a platform-aware tool, not a universal crawler.

No. Permission depends on ownership, license, terms, and applicable law. Use an official download, API, or written permission when access is restricted.