ScreenshotNeo

BlogHow-to

How to Scrape Website Data into Excel in 2026

Import website tables into Excel with Power Query, handle pages without tables, and troubleshoot refreshes when the data loads dynamically.

By the ScreenshotNeo team30 September 202610 min read

How to Scrape Website Data into Excel in 2026

To import data from a website into Excel, use the built-in Power Query web connector: in desktop Excel, choose Data → From Web, enter the page URL, select a detected table in Navigator, then choose Load or Transform Data. You can refresh the resulting query later, provided the page and your Excel setup continue to expose the data in a way the connector can read. Microsoft documents the workflow and its table-detection options in its web connector guide.

This guide walks through ordinary HTML tables, pages without detected tables, APIs and downloadable files, dynamic content, authentication, refresh, and common failures. “Scraping” here means collecting information a site makes available; check the site’s terms and access rules before collecting it. A successful request alone does not establish permission.

1. Choose the right way to bring the data in

Start by identifying what the URL serves. A page with a visible table is a good candidate for From Web. A published CSV or JSON endpoint may be simpler and more stable than extracting the same values from a rendered webpage. If the page fills its content with JavaScript after it opens, the ordinary request may not see the final data. Power Query supports web pages and web-hosted formats including CSV, JSON, XML, Excel files, and PDFs; the Web connector can retrieve these sources too. See Microsoft’s Power Query Web Connector documentation for supported source types and differences among functions and host products.

Source Good starting point What to check
HTML table Data → From Web Right table, headings, rows, and refresh
Irregular page or list Navigator → Add table using examples Whether examples produce the intended fields
CSV, JSON, XML, PDF, or Excel file hosted online From Web; use a file-specific connector when available for your workflow File structure, authentication, and updated URL
API endpoint Web connector or an appropriate API workflow Documented endpoint, parameters, response format, and credentials
JavaScript-rendered page Try the connector preview; investigate Web.BrowserContents if scripts are needed Wait condition, runtime requirements, and refresh environment

2. Import a detected webpage table

  1. Open the workbook in desktop Excel and select Data → From Web. Some versions show this under Get Data → From Other Sources → From Web. Menus and capabilities vary by Excel edition.
  2. Enter the complete page URL, including its scheme, such as https://example.com/catalog, and continue.
  3. In Navigator, inspect Suggested Tables and select the candidate matching the visible page. Use the preview and, where available, Web View to compare the extracted table with the source.
  4. Choose Transform Data if the result needs cleanup, or Load to add it to the workbook.
  5. Check the loaded sheet: compare column names, row count, and sample values with the source. Save the workbook.

Power Query creates a query-backed result, so you can refresh its connection to request updated data. Refresh does not guarantee the site still has the same markup, content, access rules, or availability. Confirm the refreshed result before relying on it for a report.

Power Query can select, shape, and load a detected webpage table into a workbook.
Power Query can select, shape, and load a detected webpage table into a workbook.

Clean and shape the result before loading

In Power Query Editor, remove columns you do not need, rename headers, filter rows, remove duplicates, and set data types deliberately. Dates, prices, and identifiers deserve particular attention: automatic type detection may interpret a date according to a locale or strip leading zeros from an identifier. Keep codes such as postal codes and product IDs as text if their digits are not quantities. If the source uses multiple tables, verify that you are combining rows with matching meanings and compatible columns.

3. Extract data when Navigator finds no table

When the information appears on the page but Navigator does not offer a useful table, try Add table using examples. The feature lets you preview the page and enter sample values for the fields you want. Power Query uses those examples to infer matching values, including content that is not arranged in a conventional HTML table. This is an extraction aid, not a guarantee that every visual element or page layout can be interpreted accurately. Microsoft documents the workflow in Get webpage data by providing examples.

  1. Open the page with From Web and select Add table using examples in Navigator.
  2. For each field, enter sample values copied from the page, ideally examples that distinguish the field from nearby text.
  3. Review the suggested rows and columns. Add more examples when values are ambiguous.
  4. Accept the extracted table, then use Transform Data to remove mistakes, fix headers, and set types.
  5. Load it and compare several rows against the source page. Refresh once and repeat the comparison.

Example suggestions have a documented limit: suggested values are no longer than 128 characters. Long descriptions may therefore be unsuitable as examples. If the page is frequently redesigned, a generated extraction can become inaccurate without an obvious error; validate both the values and the row count after refresh.

4. Use the Web connector for files and API responses

If the address points directly to a web-hosted CSV, JSON, XML, PDF, or Excel file, From Web can retrieve it. The connector documentation says that Power Query Desktop uses the Web connector for files hosted on the web; a local file can instead use its specific connector. For an API, prefer the API’s documented data endpoint over extracting values from a decorative page when that endpoint is available and permitted. APIs may paginate results, require query parameters, or return nested JSON that needs expansion in Power Query.

The connector’s advanced URL experience can combine URL parts, and Web.Contents supports HTTP request headers. Microsoft lists Anonymous, Windows, Basic, Web API, and Organizational Account authentication for Web.Contents, with differences by function and host product. POST requests through Web.Contents can be made only anonymously. Do not paste secrets into a workbook query that will be shared; choose an authentication method supported by the endpoint and the Excel environment, and restrict who can access the workbook.

Here is a Power Query M example for reading a public JSON endpoint and turning a list of records into a table. Replace the illustrative address and field names with the endpoint’s documented values. This pattern assumes a JSON array of records; other response shapes need different navigation.

let
    Source = Json.Document(
        Web.Contents("https://example.com/api/items")
    ),
    Rows = Table.FromRecords(Source),
    Typed = Table.TransformColumnTypes(
        Rows,
        {{"name", type text}, {"price", type number}}
    )
in
    Typed

Create or edit a blank query in Power Query Editor, open Advanced Editor, and adapt the M expression. The endpoint above is a placeholder, not a real data source. For authentication, request headers, URL parameters, and supported options, consult Microsoft’s Web.Contents reference and the service’s own API documentation.

5. Handle JavaScript-rendered pages

Some pages show a shell first and populate the table after browser-side scripts run. If a normal From Web preview is empty or incomplete, determine whether the data is available from a documented API or downloadable file. If the page itself must be rendered, Microsoft’s troubleshooting guidance describes using Web.BrowserContents with a WaitFor selector or a wait time before retrieving HTML. Waiting for a meaningful element is generally more robust than guessing a delay, but it still depends on the page continuing to expose that element.

Dynamic pages may need a browser-aware wait before their content can be extracted.
Dynamic pages may need a browser-aware wait before their content can be extracted.
let
    Page = Web.BrowserContents(
        "https://example.com/catalog",
        [WaitFor = [Selector = "table"]]
    ),
    Tables = Html.Table(
        Page,
        {
            {"Name", "table tr td:nth-child(1)"},
            {"Price", "table tr td:nth-child(2)"}
        },
        [RowSelector = "table tr"]
    )
in
    Tables

This M expression is a template for a page that renders a table matching those CSS selectors; replace the URL and selectors to fit the page. It is not guaranteed to work against every site. Web.BrowserContents requires Microsoft Edge WebView2 runtime in the documented scenarios. Microsoft also lists a legacy Internet Explorer 10 prerequisite for Web.Page; check your host and connector version before adopting either function. Cloud refresh and gateway requirements vary by connector and host, so test in the environment where scheduled or shared refresh will actually run. See Microsoft’s web connector troubleshooting guide.

6. Refresh and validate the workbook

After loading, use the query or workbook refresh controls to request the source again. Before distributing the workbook, check these items:

  • Columns: Are names and order still right? Did a source redesign introduce a new header?
  • Rows: Does the count match the expected page or API result? Could pagination or filters omit records?
  • Values: Spot-check records against the source, including blanks, dates, decimal separators, and identifiers.
  • Refresh: Does a new refresh complete using the credentials and machine or service that will own future refreshes?
  • Load destination: Does the query still load to the expected worksheet or data model without overwriting unrelated work?

Authentication is scoped to a URL level in the connector, and saved credentials may need to be changed if the site’s access method changes. Excel editions do not all expose identical controls, and cloud publishing can have additional gateway or credential requirements. Treat desktop success as a useful check, not proof that every other refresh environment is configured.

7. Common errors and fixes

Symptom Likely cause What to try
No table appears Page content is irregular or rendered after load Try Add table using examples; inspect the preview; investigate an API or browser-rendered extraction if appropriate.
Table is empty or missing recent values Data loads dynamically, or the source changed Compare the visible page with the preview. Consider Web.BrowserContents with WaitFor, then validate the extracted table.
“Column not found” after refresh Header changed, extraction is inconsistent, or dynamic content was not ready Inspect the refreshed preview and query steps. Use a suitable WaitFor condition for dynamic pages and update transformations only after confirming the new structure.
Access denied, sign-in prompt, or credential error Wrong authentication method or unsupported service login Choose a connector-supported method at the appropriate URL scope; verify permissions with the site owner. Some arbitrary OAuth flows are not supported by default.
WebView2 runtime initialization error Required browser runtime is missing in the host environment Install or repair Microsoft Edge WebView2 runtime where permitted, then retry in that environment.
Works on desktop, fails after publishing Different host capabilities, credentials, or gateway setup Check the connector’s requirements for the target host and configure its credentials or gateway. Do not assume desktop and cloud refresh behave identically.
Values shift or types look wrong Locale differences, changed markup, or automatic type inference Set types explicitly, keep identifiers as text, and compare sample values with the source.

8. Performance, reliability, and cost

For a stable source, importing only the needed columns and rows reduces unnecessary transformation and keeps the workbook easier to audit. Avoid refreshing more often than the source and your use case require. Each refresh depends on network access, source availability, credentials, and the site continuing to serve compatible content. A website can change its HTML without changing its URL, so a successful refresh may still need a data-quality check.

Power Query is the ordinary built-in starting point for a page that exposes usable content. Costs depend on your Excel licensing and the source; Microsoft says the new Web connector in Excel is available with an Office 365 subscription. The connector itself does not make a site’s data free to collect, and the dossier does not establish a universal legal permission rule. Check the site’s terms and access rules. For recurring or business-critical imports, prefer an official API or file export when one is available and permitted, and keep a record of the expected schema and refresh owner.

9. Or skip the browser setup

If your task is to preserve how a page looks, rather than turn its rows into spreadsheet cells, ScreenshotNeo is a website screenshot API and MCP server. A screenshot is a visual record, not structured Excel data: use Power Query or a permitted data endpoint when you need rows and columns. ScreenshotNeo can help when you need page captures alongside your workflow, without configuring a browser capture stack. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Sign up for 1,000 free screenshots a month, with no card required.

10. FAQ

Can Excel automatically keep website data current?

You can refresh a Power Query connection. Whether it succeeds depends on the source, credentials, Excel edition, and refresh environment.

Can I use Excel for the web to create every web query?

Feature availability differs by Excel product and host. Build and verify the connection in a supported environment, then check whether your target web workflow supports refresh.

Why did the imported table change after I refreshed?

The source may have changed, the page may have rendered differently, or the query may have selected a different structure. Compare the current preview and loaded result with the source before using the data.

Does importing a public page mean I have permission to reuse its data?

No. Technical access does not decide permission. Review the target site’s terms and access rules for your intended collection and use.

Can ScreenshotNeo put webpage data into Excel columns?

No. ScreenshotNeo returns a screenshot or PDF. For structured rows and columns, use Power Query or a permitted structured data source.