ScreenshotNeo

BlogHow-to

How to Recover a Corrupted SingleFile HTML Archive

Diagnose a damaged SingleFile archive safely, identify its format, and recover readable content where possible without risking the original.

By the ScreenshotNeo team4 October 20267 min read

If a SingleFile archive will not open or looks incomplete, first make an untouched backup and diagnose what the file actually contains. There is no documented universal repair procedure, and missing or overwritten bytes may be impossible to recover. You may still be able to extract readable text, recover embedded resources, or discover that the symptom is caused by a file format, browser permission, or expected offline behavior.

1. Protect the original and identify the symptom

  1. Keep the original file unchanged. Make a duplicate and do every inspection, conversion, or extraction attempt on that copy.
  2. Record the file size and extension. If the file came from a download, check whether the download was interrupted and whether another copy exists.
  3. Write down what happens: does the browser show an error, a blank page, partial content, or text without images? Does an archive utility recognize it? These symptoms point to different causes.
  4. If any text is visible in a browser, copy it into a separate plain-text file before experimenting further.

SingleFile supports several output forms, including HTML, a self-extracting ZIP, MHTML, Safari Webarchive, and HTML with a resource folder. An extension is only a clue: a renamed file may not have the structure its name suggests. Identify the format before choosing a recovery path. See the SingleFile project documentation for its documented formats. [c001]

2. Check whether the file is actually corrupted

Try opening a duplicate in a browser

Open the copy locally in a browser. If it displays partially, save the visible text and note which assets or sections are absent. A browser’s local-file access restrictions can interfere with SingleFile interacting with files on disk; that limitation does not by itself prove the archive is damaged. [c003]

Account for expected offline behavior

SingleFile removes scripts by default because scripts can change page rendering and are not guaranteed to work offline. Missing interactive menus, maps, or carousels can therefore be expected behavior rather than corruption. Retaining scripts is an option when creating a new capture, but changing that setting cannot put scripts back into an existing archive that never contained them. [c003]

3. Try non-destructive recovery based on the format

HTML file

On a duplicate, open the file in a plain-text editor. HTML is text-based, so readable text or markup may survive even if the page no longer renders correctly. Copy intact portions into a new file and preserve the original markup where possible. This is a cautious recovery tactic, not a SingleFile-supported repair procedure. Incomplete tags may render partially, and absent embedded image or style data cannot be reconstructed from the remaining text alone.

SingleFile self-extracting ZIP

If the file may be SingleFile’s self-extracting ZIP format, try opening or extracting the duplicate with an archive utility. SingleFile explains that these files are essentially regular ZIP files with extra data allowed before and after the ZIP payload, so an archive utility may be able to recover the page and resources even when a browser cannot display the outer file correctly. Extraction is only an attempt; it is not guaranteed to work on damaged bytes. [c002]

MHTML

If the file is MHTML, use a format conversion path rather than assuming it is ordinary HTML. The SingleFile FAQ points to the mhtml-to-html project for converting MHTML into a single HTML file. That is a conversion pointer; the FAQ does not say it repairs arbitrary byte damage. Work from a duplicate. [c005]

Safari Webarchive or HTML with a resource folder

Keep the archive and any companion resource folder together while testing. If a page expects resources in a separate folder, moving or renaming only the HTML file can make images and styles appear missing. Use a browser or a tool that understands the identified format. The reviewed SingleFile guidance documents these output forms but does not prescribe a universal repair utility for them. [c001]

4. If the problem is the SingleFile extension

If saving new pages is failing, rather than an already-created archive being unreadable, follow the project’s basic troubleshooting sequence. On a test page, try saving in an incognito window, reset SingleFile’s options, restart the browser, and then disable other extensions to check for a conflict. These steps investigate extension behavior; they do not restore corrupted archive bytes. [c004]

5. When the archive cannot be recovered

  • Look for another copy in the original download location, browser downloads, backups, or synced storage.
  • If you know the page URL, check whether the original page is still available or whether a web archive has a copy.
  • For future captures, save a second copy in a separate location. A second independent copy helps if one file is truncated or overwritten.
  • Do not assume a repair utility can recreate data that is absent. The reviewed official documentation does not establish a way to reconstruct missing bytes.

6. Capture a fresh screenshot of the source page

If the source URL still loads, a screenshot can preserve the page’s visible appearance even when the SingleFile archive cannot be repaired. It will not restore the original HTML, links, or interactive behavior. You can capture it yourself with a browser automation tool, or use a screenshot API.

DIY browser capture with Playwright

For a new capture, install Playwright and its Chromium browser, then save this as capture.mjs. Replace the URL with the original page. This creates a full-page PNG; it does not repair or inspect the damaged archive.

npm install playwright
npx playwright install chromium
// capture.mjs
import { chromium } from 'playwright';

const url = 'https://example.com';
const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage({ viewport: { width: 1440, height: 900 } });
  await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
  await page.screenshot({ path: 'page.png', fullPage: true });
} finally {
  await browser.close();
}

For sites with long-lived network requests, networkidle may never arrive. Use domcontentloaded or load and wait for a specific selector or a short delay instead. Pages behind authentication may need a browser context with the required cookies or headers. Only capture pages you are authorized to access.

Or skip the browser setup

Make a one-call screenshot request with ScreenshotNeo. See the API documentation for the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, and failed loads are not billed; response headers indicate the page verdict and billing status. Its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

7. Troubleshooting

Symptom Likely cause What to try
The browser says the file is invalid or will not open The extension may not match the actual format, or the file may be truncated. Work on a duplicate, check its size and origin, and identify whether it is HTML, ZIP-based, MHTML, Webarchive, or HTML plus resources.
The page opens but images or styles are missing A companion resource folder may be separated, embedded data may be damaged, or the file may use another format. Keep companion files together; test ZIP extraction if appropriate; inspect readable HTML for surviving content.
Menus or other interactions do not work offline SingleFile removes scripts by default, and scripts may depend on the network. Treat this as expected capture behavior unless other evidence indicates corruption. Enable script saving for future captures if needed.
SingleFile cannot access a local file Browser local-file permissions may limit extension access. Review the browser’s local-file access setting. This alone does not show that the archive is corrupted.
An archive utility cannot extract the self-extracting ZIP The ZIP payload may be damaged, or the file may not be that format. Confirm the format and retry only on a fresh duplicate. There is no guarantee an archive tool can recover damaged payload data.
MHTML conversion produces no usable page Conversion does not repair arbitrary corruption; source bytes or MIME boundaries may be missing. Try another intact copy or recover any readable content from the original. Do not treat conversion as byte repair.
New saves fail across pages An extension setting, browser state, or another extension may be interfering. Try incognito, reset options, restart the browser, and disable other extensions in turn.

8. Performance, reliability, and cost

Recovery attempts are safest when each starts from a fresh duplicate, so an unsuccessful conversion or extraction does not compound damage. Text inspection is quick and low risk; extraction and conversion take longer and depend on the format and file condition. There is no documented recovery-rate benchmark in the reviewed sources, so a success probability cannot be stated.

A new screenshot is a visual record, not an archive repair. Browser automation requires a compatible runtime and browser installation, and page readiness can vary with network conditions and dynamic content. A screenshot API removes that setup work but requires an API key and a reachable source page. ScreenshotNeo’s free allowance and plan prices are listed above; do not incur repeated captures when a single visual record is enough.

FAQ

Can a corrupted SingleFile archive always be repaired?

No. The official materials reviewed do not describe a universal repair tool. Recovery depends on what data remains in the file.

Can I recover text from a broken HTML file?

Often you can at least inspect the duplicate in a text editor and copy surviving text or markup. That does not guarantee the page or its embedded resources can be reconstructed.

Does converting MHTML fix corruption?

No such guarantee is documented. The cited converter is for changing MHTML into single-file HTML, not repairing arbitrary missing or overwritten bytes.

No. It records visible pixels. Keep trying to recover the archive if you need the underlying HTML, links, or content structure.

Sources

  • SingleFile README: supported formats and extension troubleshooting. [c001] [c004]
  • SingleFile FAQ: self-extracting ZIP structure, local file permissions, script behavior, and MHTML conversion pointer. [c002] [c003] [c005]