Python Syntax Errors: Common Mistakes and How to Fix Them in Scraping Code
Fix Python syntax, indentation, and parser errors in scraping scripts with runnable examples, debugging steps, and Beautiful Soup troubleshooting.

A Python scraper that stops with SyntaxError: invalid syntax has not reached the network or HTML parsing stage. Python could not parse the source file, so fix the grammar first. Start at the reported line, then inspect the line immediately before it for a missing colon, quote, comma, or closing delimiter. The caret marks where the parser noticed the problem, which is often after the character that caused it.
Once the file parses, separate syntax problems from runtime failures such as NameError, TypeError, HTTP errors, timeouts, and Beautiful Soup issues. This guide covers the mistakes most often found in scraping code, a repeatable debugging workflow, complete working examples, and fixes for errors that look like syntax failures but occur later.
What a Python syntax error means
Syntax errors are parsing errors. Python reads the source before executing the scraper. If the grammar is invalid, no request is sent and no selector is evaluated. The standard traceback includes a filename, line number, source text, and an arrow near the earliest token where parsing failed. See the Python tutorial’s errors and exceptions chapter.

File "scrape.py", line 8
for link in links
^
SyntaxError: invalid syntax
The missing character is the colon on the preceding header, not the location under the caret:
for link in links:
print(link)
CPython also records filename, lineno, offset, text, end_lineno, and end_offset on a SyntaxError. IndentationError covers incorrect block indentation, while TabError identifies inconsistent tabs and spaces. These details are documented in the exception reference.
First response: classify the failure
- Read the final exception name, not only the first traceback line.
- If it is
SyntaxError,IndentationError, orTabError, inspect source grammar and indentation. - If the program starts and then fails, it is a runtime or library error. Inspect the exception type and the values involved.
- Run
python -m py_compile scrape.pybefore making any HTTP request. This checks syntax without executing the file.
Use the same interpreter for compilation and execution. python --version and python -m pip --version should point to the environment where your scraper’s packages are installed.
Common syntax mistakes in scraping code
1. Missing colons after block headers
Every if, for, while, def, class, try, except, else, and finally header that introduces a block ends in :.
# Invalid
if response.status_code == 200
html = response.text
# Valid
if response.status_code == 200:
html = response.text
Scraping scripts commonly hide this mistake in long loops or exception handlers:
try:
response = requests.get(url, timeout=20)
except requests.RequestException as exc:
print(exc)
else:
soup = BeautifulSoup(response.text, "html.parser")
2. Unmatched parentheses, brackets, and braces
Nested request parameters and CSS selectors make delimiters easy to lose. Format dictionaries and function calls vertically so each opening delimiter has a visible partner.
# Invalid: the params dictionary is not closed
response = requests.get(
url,
params={"page": 2, "category": "books",
timeout=20,
)
# Valid
response = requests.get(
url,
params={"page": 2, "category": "books"},
timeout=20,
)
Most editors can highlight matching delimiters. If the caret appears at a later closing parenthesis, inspect the whole expression above it.
3. Unterminated strings and quotes inside selectors
URLs, XPath expressions, JavaScript snippets, and CSS selectors are all string literals. Close the string and choose quote styles that do not conflict with the content.
# Invalid
selector = "div[data-role="product"]
# Valid alternatives
selector = 'div[data-role="product"]'
selector = "div[data-role=\"product\"]"
For multiline HTML or JavaScript, use triple-quoted strings and keep indentation deliberate:
script = """
const next = document.querySelector('a.next');
return next ? next.href : null;
"""
4. Malformed f-strings
Expressions inside an f-string’s braces must be valid Python expressions. Keep the outer quotes and inner quotes distinct, and do not place an unmatched brace in a CSS or JavaScript fragment.
# Invalid
url = f"https://example.test/items/{item['id']}"
# Valid
url = f"https://example.test/items/{item['id']}"
label = f"page-{page_number}"
The first example becomes valid when the source actually contains balanced quotes; when editing, a common failure is using the same quote character around the f-string and inside the dictionary key. Use item["id"] with a single-quoted f-string, or assign the value before formatting:
item_id = item["id"]
url = f"https://example.test/items/{item_id}"
Recent CPython versions identify field problems with an f-string: prefix. Read the field expression separately from the surrounding string.
5. Indentation drift, tabs, and spaces
Indentation is syntax in Python. A loop body, conditional body, function, and exception handler must be indented consistently. Configure your editor to insert four spaces and convert existing tabs.
# Invalid: the print statement is not aligned with the loop body
for url in urls:
response = requests.get(url)
print(response.status_code)
IndentationError usually means a block is missing or has unexpected indentation. TabError means tabs and spaces were mixed in a way Python cannot resolve. Show invisible characters in your editor, then reindent the complete block instead of adding spaces to one line.
6. Python 2 and Python 3 code mixed together
A tutorial written for Python 2 can fail immediately under Python 3. Common signs include print "text", urllib2, or old exception syntax. Beautiful Soup documents an invalid-syntax failure caused by running an old Python 2 version of the library under Python 3. Check the interpreter and package versions before changing otherwise-correct code.
python --version
python -m pip show beautifulsoup4 requests
Use a current Python 3 compatible release and update copied examples to Python 3 syntax.
7. Pasted markup, prompts, or notebook artifacts
Copying from a web page can add backticks, ellipses, line numbers, or prose to a .py file. Remove Markdown fences such as ```python, notebook prompts such as >>>, and any explanatory text before compiling.
A minimal, valid scraper to compare against
Use this small fixture to isolate grammar from networking and selectors. Install dependencies with python -m pip install requests beautifulsoup4.
from __future__ import annotations
import requests
from bs4 import BeautifulSoup
def fetch_titles(url: str) -> list[str]:
response = requests.get(url, timeout=20)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
return [tag.get_text(" ", strip=True) for tag in soup.select("h2")]
if __name__ == "__main__":
for title in fetch_titles("https://example.com"):
print(title)
Save it as scrape.py, run python -m py_compile scrape.py, then run python scrape.py. If compilation succeeds but execution fails, the grammar is no longer the problem.
Separate syntax errors from Beautiful Soup and requests errors
Beautiful Soup can parse syntactically valid Python and still fail because of the selected parser or because the returned object is not what the code expects. Its documentation recommends trying another parser when parser-related crashes occur.
ResultSet versus one Tag
find_all() returns a ResultSet. Treating it as one tag causes a runtime AttributeError, not a syntax error.
# Runtime error: ResultSet has no .get_text() method
cards = soup.find_all("article")
print(cards.get_text(strip=True))
# Correct: iterate
for card in cards:
print(card.get_text(" ", strip=True))
# Correct when one result is expected
card = soup.find("article")
if card is not None:
print(card.get_text(" ", strip=True))
Requests exceptions
After the file parses, handle expected network failures specifically. A timeout, DNS failure, refused connection, or non-2xx response is separate from Python grammar.
import requests
try:
response = requests.get(url, timeout=20)
response.raise_for_status()
except requests.Timeout:
print("The server took too long to respond")
except requests.HTTPError as exc:
print(f"HTTP failure: {exc}")
except requests.RequestException as exc:
print(f"Request failed: {exc}")
else:
print(response.text[:200])
finally:
print("Request attempt finished")
Python’s tutorial explains specific handlers, else for code that runs only after success, and finally for cleanup. Avoid catching every exception while debugging; a broad handler can hide the actual defect.
Reliable debugging workflow
- Copy the complete traceback. Record the filename, line, offset, and exception class.
- Inspect the previous line. Look for a missing colon, quote, comma, or closing delimiter before the caret.
- Compile only. Run
python -m py_compile file.pyorpython -m compileall .. - Reduce the script. Comment out network calls and keep one function, one URL, and one selector.
- Run request and parsing stages separately. Save a response to an HTML fixture, then parse that file.
- Restore inputs gradually. Start with one known URL, then add pagination, concurrency, retries, and additional selectors.
# Save a fixture after the request succeeds
from pathlib import Path
Path("page.html").write_text(response.text, encoding="utf-8")
# Parse without making another request
html = Path("page.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
When the page needs a browser
Requests and Beautiful Soup process the HTML returned by the server. They do not execute JavaScript, accept consent dialogs, or wait for client-rendered content. If the target page is generated in a browser, use a browser automation tool or a screenshot service and keep that capture step separate from your Python parser. Syntax debugging still happens first: browser setup cannot repair an invalid Python file.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API when you need a rendered page image or PDF instead of maintaining browser infrastructure. Cookie banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', buffer);
ScreenshotNeo supports full-page and element captures, dark mode, device presets, custom viewports, retina scale, PDF output, custom CSS and JavaScript, click and wait actions, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, and a usage API. You can choose a cache TTL and inspect X-Page-Verdict and X-Billed headers when diagnosing a response.
The Free plan includes 1,000 screenshots per month without a card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
Troubleshooting table
| Message or symptom | Likely cause | Fix |
|---|---|---|
SyntaxError at a closing parenthesis |
Missing quote, comma, or delimiter earlier | Inspect the preceding expression and reformat nested calls. |
IndentationError: expected an indented block |
Block header has no body | Add an indented statement or use pass temporarily. |
TabError |
Tabs and spaces are mixed | Convert all indentation to four spaces. |
NameError: name 'x' is not defined |
Code parsed but a variable is missing or misspelled | Check assignment order and spelling. |
AttributeError: 'ResultSet'... |
find_all() returned many tags |
Iterate the result or use find(). |
| Parser crash in Beautiful Soup | External parser or malformed input | Try html.parser, lxml, or html5lib as appropriate. |
| Works in a tutorial, fails locally | Python 2/3 or package-version mismatch | Check interpreter and package versions in the active environment. |
| HTML is empty or incomplete | Timeout, bot check, JavaScript rendering, or server response issue | Inspect status and response text; use browser rendering when required. |
Performance, reliability, and cost notes
Compile before crawling so a typo cannot waste requests. Use one persistent requests.Session, explicit timeouts, bounded retries, and a small concurrency limit. Cache responses or parsed fixtures while developing selectors. Separate fetch, parse, and storage functions so a parser change does not repeat network work.
For browser-rendered captures, wait only for a meaningful selector or network-idle condition instead of using a large fixed delay. Block unnecessary ads, trackers, and resource types when the page does not need them. Cache stable screenshots with a TTL and use asynchronous jobs for large batches. ScreenshotNeo supports bulk capture for up to 100 URLs per call; failed loads and cache hits are not billed, which makes retries and cache use easier to reason about. Review usage through its usage API and response headers.
Pre-run checklist
- Run the intended Python interpreter and verify package versions.
- Compile the file before executing it.
- Check the line before the caret for punctuation.
- Use four spaces consistently.
- Confirm every string and delimiter closes.
- Test one known URL and one saved HTML fixture.
- Distinguish parser, HTTP, and selector exceptions.
- Use a browser capture only when the page requires JavaScript or interaction.
FAQ
Does the caret identify the exact bad character?
No. It identifies where parsing became impossible. The missing character is often immediately before the marked token or on the previous line.
Can a syntax error come from a website’s HTML?
Not directly. Website HTML is runtime input. A syntax error comes from your Python source, although malformed HTML can later cause parser or selector problems.
Should I catch SyntaxError in the scraper?
Usually no. Fix source grammar before running the program. Catch specific runtime exceptions around requests, parsing, and storage operations.
Why does a scraper work with html.parser but not another parser?
Different parsers have different dependencies and recovery behavior. Install the chosen parser and test it against the saved HTML fixture.
How do I know whether a failed page capture was billed?
For ScreenshotNeo responses, inspect the X-Page-Verdict and X-Billed headers. The service reports whether the result was a clean capture and whether it was billed.


