How to Scrape Transfermarkt Data with an API
Transfermarkt has no clearly documented public API. Learn the permission-first architecture, authorized collection pipeline, FastAPI wrapper, retries, schema design, and safer alternatives.
Short answer: Transfermarkt has no clearly documented public API in the cited FAQ. A forum answer dated May 14, 2020 said the site did not have an API that was publicly available. Community projects therefore wrap web pages or undocumented endpoints, but Transfermarkt’s terms prohibit automated copying with bots, spiders, screen scraping, or other automated processes. Before writing a collector, obtain written permission or use a licensed data source.
This guide shows the engineering design for an authorized integration: a narrow acquisition worker, conservative pacing, retries, validation, a normalized store, and a FastAPI layer. It also explains the data surfaces commonly used by community projects, why undocumented endpoints are fragile, and how to operate the pipeline without presenting proxying or anti-bot techniques as permission.
1. Check permission before automating
Transfermarkt’s terms state: The User is not permitted to access or copy the Digital Content using bots, spiders, screen scraping or other automated processes.
Read the official terms and obtain written authorization or a licensed feed before collecting data. The 2020 forum answer is historical and does not prove that policy cannot change.
- Document the legal basis, allowed geography, seasons, fields, refresh rate, retention period, and redistribution rights.
- Ask for a sanctioned endpoint or export if your use is commercial, high volume, or involves redistribution.
- If you do not have permission, stop at this checkpoint. Do not use proxies, CAPTCHA services, fingerprint changes, or other access-control workarounds.
2. Choose an acquisition design
A production system should separate acquisition from serving:
- Acquisition: fetch only approved URLs and fields.
- Validation and retry: identify transient failures, changed HTML, blocked responses, and null records.
- Raw store: retain the response, source URL, entity ID, retrieval time, and parser version.
- Normalized store: load stable tables for clubs, leagues, players, competitions, games, lineups, appearances, and transfers.
- API layer: expose only the fields your application needs.
The transfermarkt-datasets project demonstrates a raw-to-prepared workflow using raw assets, competition IDs, dbt, and DuckDB. Keeping raw inputs makes parser updates and audits possible.
3. Know the data surfaces
| Entity | Typical data | Operational caution |
|---|---|---|
| Club | Profile, search, leagues, transfer history | Names and page structure can change |
| League | Profile, search, clubs | Season and competition IDs must be stored |
| Player | Profile, search, transfer history | Use stable IDs rather than names as keys |
| Market value | Historical value graph | Community scripts call https://www.transfermarkt.com/ceapi/marketValueDevelopment/graph/{player_id}; this is undocumented |
| Transfers | Transfer history list | Community scripts call https://www.transfermarkt.co.uk/ceapi/transferHistory/list/{player_id}; this is undocumented |
| Competition hierarchy | Confederations, competitions, editions, games, lineups, appearances | Recursive crawlers have broad request footprints |
Undocumented CE endpoints are implementation details from community projects, not an API contract. Expect authentication changes, HTML changes, bot protection, missing fields, and endpoint removal.
4. Build a conservative authorized collector
The following example is a FastAPI service that fetches an authorized page, extracts a small record set, and returns JSON. Replace the URL and selectors only after permission is established. It deliberately has no proxy rotation, CAPTCHA bypass, or stealth behavior.
Install
python -m venv .venv
. .venv/bin/activate
pip install fastapi uvicorn requests beautifulsoup4
collector_api.py
import os
import time
from typing import Any
import requests
from bs4 import BeautifulSoup
from fastapi import FastAPI, HTTPException, Query
app = FastAPI(title="Authorized football data collector")
USER_AGENT = os.getenv("COLLECTOR_USER_AGENT", "authorized-data-client/1.0 contact@example.com")
TIMEOUT = 30
MAX_RETRIES = 3
MIN_INTERVAL = 1.5 # example pacing; agree a limit with the data owner
_last_request = 0.0
def fetch(url: str) -> str:
global _last_request
for attempt in range(MAX_RETRIES):
wait = MIN_INTERVAL - (time.time() - _last_request)
if wait > 0:
time.sleep(wait)
try:
response = requests.get(
url,
headers={"User-Agent": USER_AGENT, "Accept": "text/html"},
timeout=TIMEOUT,
)
_last_request = time.time()
if response.status_code in (429, 500, 502, 503, 504):
if attempt + 1 == MAX_RETRIES:
response.raise_for_status()
time.sleep(2 ** attempt)
continue
response.raise_for_status()
return response.text
except requests.RequestException:
if attempt + 1 == MAX_RETRIES:
raise
time.sleep(2 ** attempt)
raise RuntimeError("unreachable")
def parse_player(html: str, source_url: str) -> dict[str, Any]:
soup = BeautifulSoup(html, "html.parser")
name = soup.select_one("h1")
if not name:
raise ValueError("expected player heading was not found")
return {
"name": name.get_text(" ", strip=True),
"source_url": source_url,
"retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"parser_version": "1.0.0",
}
@app.get("/players")
def player(url: str = Query(..., description="Approved source URL")):
try:
html = fetch(url)
record = parse_player(html, url)
return {"data": record, "source": "authorized"}
except ValueError as exc:
raise HTTPException(status_code=422, detail=str(exc))
except requests.RequestException as exc:
raise HTTPException(status_code=502, detail=f"upstream request failed: {exc}")
Run and call it
uvicorn collector_api:app --reload --port 8000
curl --get 'http://127.0.0.1:8000/players' \
--data-urlencode 'url=https://authorized.example/player/123'
In a real deployment, do not accept arbitrary URLs from callers. Map an internal player ID to an allow-listed URL, queue work in a background worker, and require authentication on your FastAPI routes.
5. Add retries, null monitoring, and schema controls
A community acquisition script uses a descriptive User-Agent, up to three retries, and rejects a run when more than 20% of responses are null. These are engineering examples, not Transfermarkt requirements.
- Retry timeouts, connection resets, HTTP 429, and 5xx responses with exponential backoff and jitter.
- Do not retry parser errors indefinitely. Save the raw response and alert on a selector or schema change.
- Track requested, successful, empty, blocked, and malformed responses by entity type and season.
- Fail a batch when its null rate exceeds your agreed threshold; never silently publish partial data.
- Store IDs, source URLs, season, retrieval timestamp, parser version, and a content hash.
- Use database upserts keyed by source entity ID plus competition or season where necessary.
6. Normalize the football hierarchy
Keep dimensions separate from events. A practical schema contains:
clubs(id, name, country_id, source_url, retrieved_at)
competitions(id, name, confederation_id, source_url)
seasons(id, competition_id, label)
players(id, name, birth_date, nationality, source_url)
games(id, season_id, home_club_id, away_club_id, kickoff_at)
appearances(game_id, player_id, minutes, starter)
transfers(id, player_id, from_club_id, to_club_id, fee, transfer_date)
market_values(player_id, observed_at, value, currency)
Keep raw JSON or HTML outside these tables. Add provenance columns and treat missing values as unknown rather than zero. Preserve the original currency and value text when parsing monetary fields so a later parser can correct conversion logic.
7. Expose a small FastAPI surface
Serve stable resources instead of mirroring every source page. Common community examples include:
GET /clubs/{id},GET /clubs?search=..., and club transfer history.GET /leagues/{id}, league search, and league clubs.GET /players/{id}, player search, and player transfer history.GET /players/{id}/market-valuesandGET /players/{id}/transfers.
Add pagination, an explicit schema version, and an updated_at value. Do not expose source credentials or allow callers to submit arbitrary upstream URLs.
8. Performance and reliability
| Concern | Practical control |
|---|---|
| Request volume | Queue jobs, deduplicate IDs, cache approved responses, and agree on pacing with the data owner. |
| Latency | Separate ingestion from reads; serve normalized data from DuckDB or a database. |
| Failures | Use bounded retries, dead-letter queues, and alerts for 429, 403, 5xx, and parser failures. |
| Freshness | Refresh high-change entities more often than historical seasons; record retrieval times. |
| Reproducibility | Version parsers, retain raw responses, and log the exact scope of every run. |
| Redistribution | Check the license and written permission for every field and downstream use. |
Throughput is constrained by the authorization agreement and upstream stability, not by how many concurrent workers you can start. Parallel requests can increase blocking risk and make an incident harder to diagnose.
9. Cost and licensing decisions
Compare a sanctioned licensed API, an authorized direct integration, and an open-source scraper on permission, contract stability, historical depth, rate limits, identity coverage, update latency, maintenance, and redistribution rights. A scraper may have no software fee but still costs engineering time, storage, monitoring, and incident response. A licensed feed may cost more per request while removing parser maintenance and legal uncertainty.
10. Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or challenge page | Access control or bot protection | Stop automated retries; confirm permission and use the sanctioned channel. |
| 429 responses | Rate limit exceeded | Pause the queue, reduce pacing, and agree limits with the owner. |
| 200 response with no records | Changed HTML, consent page, or empty entity | Save the raw response, classify it, and update the parser only after review. |
| CE endpoint returns errors | Undocumented endpoint changed or was blocked | Treat it as unavailable; request a supported feed instead of searching for bypasses. |
| Duplicate players | Name used as the key | Key records by source ID and retain aliases separately. |
| Incorrect transfer fees | Locale, currency, or special fee text | Store raw text, parse currency explicitly, and test edge cases such as loan and undisclosed transfers. |
| FastAPI returns stale data | Serving layer is disconnected from ingestion | Publish batch IDs and timestamps; update atomically after validation. |
11. Or skip the browser setup
If your task is to create visual evidence of a Transfermarkt page rather than extract structured records, ScreenshotNeo provides a one-call website screenshot API. See the API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Cookie banners, newsletter popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. An MCP server lets Claude, Cursor, and other MCP clients call take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
12. FAQ
Does Transfermarkt have an API?
The cited 2020 forum answer said no publicly available API existed then. Verify the current position directly and obtain permission before automating.
Can I use a community Python package in production?
Only if its collection method and your use are authorized. Review its URLs, request behavior, license, maintenance status, and redistribution terms.
Should I scrape every competition recursively?
Usually no. Start with the smallest approved scope, then expand after measuring null rates, parser stability, and operational cost.
Can I publish the collected dataset?
Only when your written permission or license explicitly allows redistribution of the fields and derived data.


