ScreenshotNeo

BlogHow-to

How to Stop People from Stealing Your Website Content

You cannot make public website content impossible to copy, but you can deter casual scraping, limit specific forms of misuse, and respond when copying occurs.

By the ScreenshotNeo team4 October 20268 min read

You cannot guarantee that content visible on a public website will never be copied. A browser has to receive the text or image to display it, and a determined person can save, photograph, or reproduce what they can see. You can make casual copying and some forms of automated collection harder, state crawler preferences, preserve evidence, and pursue removal through the relevant site or service provider.

Choose a response based on what is happening: crawler instructions communicate preferences; access and bot controls enforce restrictions; hotlink controls target certain image requests; and copyright notices address specific material after you find a suspected copy. None is a universal copy-proof switch.

1. Identify what was copied and keep a record

Before changing settings or contacting anyone, record the original and copied URLs, the date you found the copy, and what material appears to match. Save screenshots or copies of the relevant pages where lawful and practical. This helps you describe both the original work and the allegedly infringing material clearly if you later contact a service provider.

Be precise about what you own. Copyright generally concerns qualifying original expression, such as your article text or an original photograph; a name or an idea by itself is not protected in the same way. The U.S. Copyright Office explains website and website-content registration in Circular 66 and its website registration guidance.

2. Publish crawler preferences in robots.txt

A robots.txt file asks cooperating crawlers to follow your rules. For example, to ask crawlers not to fetch a particular directory, place a file at the site root with content such as:

User-agent: *
Disallow: /private-content/

Replace /private-content/ with the path you mean to describe. A directive can also target a named crawler, subject to that crawler’s documented behavior. Check your file at the site root, ensure it is served successfully, and avoid listing sensitive URLs: robots.txt is public and does not protect confidential information.

Do not treat this as access control. Cloudflare states that robots.txt “expresses your preferences” but “does not prevent crawlers from accessing your content at a technical level.” Some operators disregard its directives. Use authentication or enforced access controls for content that must not be public; use bot or firewall rules when you need to restrict traffic. See Cloudflare’s robots.txt documentation.

3. Enforce access rules for the traffic you want to restrict

If automated clients are repeatedly collecting pages, use controls at your application, hosting provider, firewall, or CDN that can actually allow or deny requests. Start with a narrow rule: for example, protect a path that should require a login, rate-limit unusually frequent requests, or challenge traffic that matches a documented abuse pattern. Keep a way for legitimate users and approved integrations to access the site.

Review logs and false positives after a rule change. A broad block can affect search crawlers, accessibility tools, feed readers, monitoring, or users behind shared networks. Keep private material behind authentication rather than relying on a crawler name, user-agent string, or an undisclosed URL as a secret.

Cloudflare documents both crawler preference and enforcement controls, but choose settings for your own traffic and verify their current scope in its bot documentation. No setting can prevent all copying of content that remains publicly viewable.

Hotlink protection addresses a narrower problem: another page embedding an image from your server so that visitors’ image requests consume your bandwidth. Cloudflare’s feature checks the HTTP Referer on image requests and supports GIF, ICO, JPG, JPEG, and PNG. It can reduce some origin bandwidth use, but Cloudflare says it does not affect crawling.

There are tradeoffs. Blocking external image requests can also prevent wanted images from displaying in search results, social previews, feeds, or other legitimate embeds. Decide which uses you want to keep, then test the rule with representative pages. Cloudflare documents the feature and its exceptions in Hotlink Protection.

5. Find the right recipient when a copy appears

If you find your work on another site, identify who operates or hosts that site and check the recipient’s current copyright-reporting process. A CDN or other intermediary may not host the page or have direct control over the origin content. Cloudflare says it forwards complaints to website operators and hosting providers when it does not host the content itself; contacting the wrong party may only result in a referral.

For a U.S. DMCA notice, identify the copyrighted work and the specific material you believe is infringing, give the required contact details and statements, and submit it to the service provider’s designated agent according to that provider’s instructions. Cloudflare’s published complaint requirements include a signature, identification of the original and allegedly infringing material, contact information, a good-faith belief statement, and statements of accuracy and authority. Those are Cloudflare’s instructions, not universal legal advice for every recipient or jurisdiction. See the Cloudflare abuse reporting process and the U.S. Copyright Office’s Section 512 resources.

A notice is a process, not an automatic decision that infringement occurred. A counter-notice may lead to restoration under the statutory process unless the rightsholder takes the specified court action within the relevant window. The Copyright Office describes federal court and the voluntary Copyright Claims Board as possible routes for certain U.S. disputes; the Board’s stated total claim limit is $30,000. Whether either route fits depends on the facts and current rules. For legal advice, consult a qualified lawyer in the relevant jurisdiction.

6. Consider registration and clear ownership records

Registration does not stop someone from copying a page, but it provides an official route for registering eligible website content in the United States and can matter to later enforcement options. The Copyright Office directs website owners to Circular 66 for procedures. Confirm current eligibility and filing requirements directly with the Office; copyright rules and remedies differ by jurisdiction.

For AI-related scraping, Cloudflare offers illustrative sample terms language and points to robots.txt permissions. Treat that language as an example to discuss with counsel, not a universal legal template or a technical block. Cloudflare’s guidance describes the approach.

7. A practical response checklist

  1. Record the original URL, copied URL, discovery date, and the matching passages or images.
  2. Decide whether the problem is crawler access, image hotlinking, or a specific copy already published.
  3. Use robots.txt only to express preferences to cooperative crawlers; use access controls for actual restrictions.
  4. For image hotlinking, check whether blocking external requests would break sharing, search, feeds, or wanted embeds.
  5. Identify the host or service provider that can act on the reported material, then follow its current reporting instructions.
  6. Keep a copy of what you submitted and any response. Seek jurisdiction-specific legal advice if the dispute needs escalation.

8. Troubleshooting common approaches

Symptom Likely cause What to do
A crawler still fetches pages after a robots.txt change Robots.txt is a preference, not an enforced block; the crawler may ignore it. Use authentication for private material or a narrow server, firewall, or bot rule for access enforcement.
A robots.txt rule has no visible effect The file may be at the wrong path, return an error, or target a different path than intended. Check that https://your-domain.example/robots.txt is served from the site root and inspect the exact path and crawler group.
An external site still copies page text Hotlink protection applies to supported image requests; it does not stop page crawling or copying text. Use appropriate bot or access controls for traffic, and the host or service-provider reporting process for a published copy.
Images disappear from social or search previews after enabling hotlink protection The rule may block legitimate external image requests along with unwanted embeds. Review exceptions and test the sharing and discovery behavior you want to preserve.
A provider says it cannot remove the page It may be a pass-through service and not the site’s host or operator. Ask where to direct the report and identify the hosting provider or site operator able to control the origin content.
A takedown notice gets challenged or the content returns The recipient may receive a counter-notice or follow its statutory process; a notice is not a final adjudication. Review the response and applicable deadlines with qualified counsel before taking further action.

9. Monitor copied pages without confusing monitoring with protection

Periodic searches for distinctive phrases from your work, reverse image searches, and review of referral or bot logs can help you notice copies or unusual collection. These are discovery methods, not proof that every copy will be found and not a way to stop access. Preserve the relevant page and URL when you find a match, then assess the facts before reporting it.

10. ScreenshotNeo: document a page as it appears

A screenshot can help keep a visual record of a page you are reviewing, including a suspected copy. It does not establish ownership, prove when a page first appeared, or replace preserving source records and following the reporting process. ScreenshotNeo is a website screenshot API and MCP server for developers. You can capture a public URL with one GET request; its options include full-page capture and formats such as PNG, JPEG, WebP, or PDF. The API documentation lists request parameters.

Or skip the browser setup: make a single request to capture a page you are documenting.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo removes known cookie and consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers say which page verdict and billing status applied. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. These captures can help with visual documentation, but they do not prevent content theft or replace your own evidence process. Sign up for 1,000 free screenshots a month, with no card required.

Frequently asked questions

Can I completely prevent people from copying a public page?

No. If a browser can display public content, a person can reproduce it by some means. Controls can discourage casual copying or restrict certain traffic, but they cannot guarantee that visible material will never be copied.

Does robots.txt protect private pages?

No. It is public crawler guidance, not authentication or an access-control mechanism. Put private content behind a login or another enforced control.

Does a DMCA notice guarantee that a copy will be removed?

No. The recipient follows its process, and a counter-notice or other dispute may arise. U.S. procedures are jurisdiction-specific, and a notice does not itself decide whether infringement occurred.

No. It targets certain externally referred image requests to your server. A visitor who can view an image may still save or reproduce it.