ScreenshotNeo

BlogGuides

Scraping All Bakeries in Paris to Build the Best Morning Bike Route

Build a reproducible Paris bakery dataset, check which shops are open, and optimize a morning ride against distance, time and cycling comfort.

By the ScreenshotNeo team30 September 202612 min read

Scraping All Bakeries in Paris to Build the Best Morning Bike Route

A reliable “best morning bike route” needs more than a list of bakeries and a shortest-path query. Define what counts as a bakery, gather and preserve records with their source and date, check opening hours for the ride window, and route the stops on a complete street network. Then compare routes by time, distance, number of stops, opening feasibility and cycling comfort.

You can automate much of this, but do not claim to have found all Paris bakeries unless you can show that your source covers them. The research sources here do not establish a complete, current official bakery register. Treat the result as a dated, deduplicated candidate set, and explain its coverage.

1. Define “all bakeries” before collecting data

“Bakery” is a data rule, not a self-evident category. A narrow rule might include businesses tagged shop=bakery. A broader rule could include documented pastry-bakery businesses, but may also bring in cafés, restaurants and shops that sell baked goods without being bakeries. Decide whether to include chains, bakeries inside markets, businesses that bake on site, and temporary or seasonal closures.

Write that rule down before collecting records. It determines what your total means. No single source in this research pass proves a complete, current count of every bakery in Paris, and different sources may use different categories or update schedules.

Decision Example rule Why it matters
Business type Include records explicitly classified as bakeries; document any added pastry-bakery category. Prevents a broad “food” search from turning into an unexplained list.
Geography Choose Paris proper or a defined ride area that crosses the city boundary. A route may be useful beyond the city, but its candidate set should say so.
Operating status Keep candidates with uncertain status but flag them for review. A missing or stale status is not proof a business is closed.
Ride window For example, 07:00–10:00 on a specified date. Opening hours and temporary closures are time-sensitive.
Stops Set a maximum stop count and dwell time per stop. “Best” changes when a route is a bakery crawl versus a quick commute.

2. Collect candidates and retain the raw evidence

Start with a city-maintained source if it contains a suitable business layer. The City of Paris describes its open-data catalog as covering commerce and tourism as well as mobility and public space, with datasets published under open-data licenses. Browse the Paris Data catalog and record the dataset identifier, license, export date and any stated coverage. The catalog is a place to look; this research does not confirm that it contains a complete bakery register. The City’s open-data overview explains the catalog’s scope and ODbL licensing.

Keep original records and provenance as you normalize and deduplicate bakery candidates.
Keep original records and provenance as you normalize and deduplicate bakery candidates.

If your chosen source does not provide enough candidates, add another documented source rather than silently filling gaps. For each record, keep:

  • Source name and stable source ID, when available.
  • Original name, address, coordinates and category exactly as collected.
  • Opening-hours string, website and any source status field.
  • Retrieval timestamp, source export date and license.
  • Normalized name and address, plus any geocoder provider and lookup date.
  • A confidence or review note for uncertain coordinates, status or category.

Store raw and normalized values separately. Normalization is useful for matching; the original record is necessary to audit a match or correct a bad one. Never geocode an existing coordinate just because a newly geocoded address looks more precise. Geocode only missing locations, retain the provider and date, and flag ambiguous results for manual review.

Deduplicate conservatively

First merge records with the same stable source identifier. For records without one, normalize accents, whitespace, punctuation and common address abbreviations, then compare name, address and distance. A nearby pair is only a duplicate candidate: two bakeries can share a building or sit on opposite sides of a street. Log every merge and preserve the original IDs. Send uncertain pairs to a manual review list instead of deleting one automatically.

3. Check which candidates can be visited that morning

Opening-hours strings are not a guarantee of current business status. Hours can be missing, irregular, seasonal or stale; holidays and temporary closures can invalidate an otherwise plausible schedule. Parse hours into candidate intervals, then recheck the most important stops close to publication or ride time. If a record’s hours cannot be parsed confidently, label it “unknown” instead of treating it as open.

For a ride window from 07:00 to 10:00, a bakery is feasible only if your estimated arrival overlaps an interval when it is open, allowing for the time spent riding from the previous stop and the planned dwell time. Make the date explicit, including weekday and local timezone. Do not assume that a weekday schedule applies on a public holiday.

A practical review checklist:

  • Confirm the shop still appears to be operating.
  • Check the hours for the actual ride date, especially for early openings.
  • Mark missing, ambiguous and conflicting hours as uncertain.
  • Allow time to order, queue and lock or park the bicycle.
  • Recheck the selected stops before publishing or starting the ride.

4. Build a cycling graph that includes ordinary streets

Paris publishes a Linéaires d’aménagement cyclable layer derived from OpenStreetMap. The City says its Mission Vélo performs quality control on the relevant data, and the extraction includes ways where cycling is allowed and access is restricted for other users. It is a dated extraction, so retain its export date.

Combine a complete street graph with cycling infrastructure and count data as separate context layers.
Combine a complete street graph with cycling infrastructure and count data as separate context layers.

That layer is valuable as an infrastructure attribute, but it is not a complete street-routing graph. Ordinary streets shared by default with motor traffic are not represented. Use a street network and bicycle routing profile that can route across the ordinary street network, then attach the Paris cycling layer as an attribute: protected or restricted segment, shared street, and any other relevant classification available in your data.

For each candidate route, calculate at least total distance and estimated cycling time. Where your routing engine supports it, also calculate exposure to high-stress segments or the proportion of travel on protected/restricted cycling infrastructure. Penalize uncomfortable segments in the route objective, but publish the penalty rule. A vague label such as “safe route” is not supported by infrastructure data alone.

Use counts as context, not a safety score

The City’s historical bicycle-count dataset has data available from 1 January 2016 and includes hourly counter and site information. Its counters are on cycle tracks and some bus lanes open to bicycles; scooters and other vehicles are not counted. The current counter feed covers a rolling 13 months and is updated daily at J-1, but the City warns that the number of counters changes and that counters can be disabled for works or fail temporarily.

Counts are directional and site-specific. Use them to describe observed bicycle activity near monitored sites or compare route alternatives, not to claim universal safety, traffic volume everywhere, or a causal effect of infrastructure. Absence of a nearby count means “unobserved,” not “low cycling.”

5. Optimize the route for the actual ride

Define the origin, return point, date, ride window, maximum number of bakery stops and dwell time. Then obtain bicycle travel times or distances between every pair of candidate stops, plus the origin and return point, from a routing engine configured for cycling. A straight-line distance matrix is not a substitute: rivers, one-way streets, bridges and access restrictions change the ride.

For a small candidate set, enumerate stop orders or use a routing optimizer. For a larger set, filter candidates to the ride area and plausible open window first, then use a vehicle-routing or traveling-salesperson solver with time windows. Score at least these route dimensions:

Measure What it tells you Limitation
Ride distance and time How long the network route is under the chosen profile. Routing estimates do not account for every delay or riding style.
Opening feasibility Whether arrivals fit published or verified opening intervals. Hours may be stale or exceptional for the date.
Stop count and dwell Whether the itinerary is a quick ride or a multi-stop crawl. Queue and service time vary.
Cycling comfort attributes How much travel uses the infrastructure categories you value. The official layer omits ordinary shared streets and is dated.
Count context Observed bicycle flow at nearby monitored sites. Coverage is partial and counters can be unavailable.

A weighted score can rank options, but show the component measures alongside it. For example, a shorter itinerary may use more shared streets; a more comfortable route may add time or skip a bakery that opens later. The reader should be able to understand the trade-off rather than trust one opaque “best” number.

6. A runnable route-ordering example

The following Python script chooses a feasible stop order using a travel-time matrix exported by your bicycle router. The matrix is deliberately an input: this article does not assume a particular routing provider or invent its endpoint. Save candidate stops and a square matrix in JSON, then run the script. Matrix rows and columns must use the same order as places; values are travel minutes, with origin first and return point last.

import json
from datetime import datetime, timedelta
from itertools import permutations

# route-input.json format:
# {
#   "start": "2026-10-01T07:00:00",
#   "end": "2026-10-01T10:00:00",
#   "places": [
#     {"name":"Origin","open":"00:00","close":"23:59","dwell":0},
#     {"name":"Bakery A","open":"07:00","close":"10:00","dwell":10},
#     {"name":"Bakery B","open":"08:00","close":"12:00","dwell":10},
#     {"name":"Return","open":"00:00","close":"23:59","dwell":0}
#   ],
#   "minutes": [[0,12,18,10],[12,0,9,14],[18,9,0,11],[10,14,11,0]]
# }

with open("route-input.json", encoding="utf-8") as f:
    data = json.load(f)
places = data["places"]
travel = data["minutes"]
start = datetime.fromisoformat(data["start"])
end = datetime.fromisoformat(data["end"])

if len(travel) != len(places) or any(len(row) != len(places) for row in travel):
    raise ValueError("minutes must be a square matrix matching places")
if len(places) < 3:
    raise ValueError("include an origin, at least one bakery, and a return point")

def at_clock(day, hhmm):
    hour, minute = map(int, hhmm.split(":"))
    return day.replace(hour=hour, minute=minute, second=0, microsecond=0)

# Indices: 0 is origin, last is return; intermediate records are bakeries.
stop_ids = range(1, len(places) - 1)
best = None
for order in permutations(stop_ids):
    route = (0, *order, len(places) - 1)
    now = start
    feasible = True
    ride_minutes = 0
    for a, b in zip(route, route[1:]):
        minutes = travel[a][b]
        if minutes is None or minutes < 0:
            feasible = False
            break
        now += timedelta(minutes=minutes)
        ride_minutes += minutes
        place = places[b]
        opens = at_clock(now, place["open"])
        closes = at_clock(now, place["close"])
        if now < opens:
            now = opens
        if now >= closes:
            feasible = False
            break
        now += timedelta(minutes=place.get("dwell", 0))
        if now > end:
            feasible = False
            break
    if feasible and (best is None or now < best[0]):
        best = (now, route, ride_minutes)

if best is None:
    print("No feasible order. Recheck hours, ride window, dwell, and matrix.")
else:
    finish, route, ride = best
    print("Route:", " -> ".join(places[i]["name"] for i in route))
    print("Ride minutes:", ride)
    print("Finish:", finish.isoformat())

This small example enumerates every ordering, so its runtime grows factorially with the number of bakeries. Keep the candidate set modest or use a routing solver for a larger set. It assumes each stop has one same-day open interval, uses local naive timestamps, and optimizes earliest finish among feasible routes. Adapt the input for split shifts, overnight hours, date-specific exceptions, multiple return points or an explicit comfort penalty. Validate the routing matrix and timezone before interpreting the result.

7. Troubleshooting data and route failures

Symptom Likely cause Fix
Candidate count seems surprisingly low The source category is narrow, coverage is incomplete or geography is clipped. Check the source’s scope and category schema; add another source transparently and keep provenance.
Two records represent one business Different spellings, addresses or source IDs. Use stable IDs first; review near-name/address matches manually and log merges.
One address maps to a distant point Ambiguous street name, wrong postcode or geocoder mismatch. Retain the original address, geocoder details and confidence; manually resolve before routing.
Route includes a closed bakery Hours are stale, exceptional hours were omitted or arrival-time math is wrong. Recheck the date and local timezone; model dwell and travel time, and flag uncertain hours.
Router says no path Disconnected graph, invalid coordinates, access restriction or overly restrictive profile. Inspect snapped points and graph coverage; verify access and cycling settings before relaxing constraints.
Route ignores a cycle lane The infrastructure layer was used as the entire graph or failed to join to route edges. Route on a complete street graph, then join infrastructure attributes using a documented spatial rule.
Counts vanish or change between runs Counter coverage changes; sites may be offline or under works. Keep dataset retrieval dates and site IDs; report missing observations as unavailable.
Script reports no feasible route Time window is too tight, hours are inconsistent, matrix has impossible values or dwell is excessive. Check matrix order and units, then inspect each stop’s interval and the origin/return timing.

8. Performance, reliability and cost

Data collection is usually the easy part; the expensive work is checking volatile business details and obtaining route costs between many candidate pairs. A matrix for n locations has n² entries, so avoid requesting every possible pair before filtering to the geographic area and morning window. Cache data with its retrieval timestamp, refresh business status and hours close to publication, and keep route results tied to the network and profile version used.

Reliability comes from recording uncertainty rather than hiding it. Store raw snapshots, source identifiers, retrieval times, export dates and transformation steps. Keep a manual review queue for duplicate candidates and ambiguous hours. If a source or counter is unavailable, mark it unavailable and continue with a disclosed limitation; do not silently substitute older data as if it were current.

Respect the dataset’s license and the terms of any additional source or routing service you choose. The City’s open-data overview describes its datasets as published under ODbL; verify the license on the specific dataset and meet its attribution and reuse conditions. Do not republish personal information or imply City endorsement of your route.

9. Publish a reproducible route, not an unsupported superlative

Include the extraction date, source coverage, inclusion rule, license, deduplication approach, geocoding method, routing profile and ride assumptions. State whether opening hours were checked and when. Show distance, estimated time, stop count, stop dwell and the cycling-infrastructure context. If you use a weighted score, publish its inputs and weights.

Phrase the result as “best under these assumptions” and state what it optimizes. A route that minimizes riding time may not maximize bakery variety or cycling comfort. A route with the most cycle-track exposure may be longer or may rely on incomplete infrastructure data. A clear methodology lets another developer reproduce and challenge the choice.

Or skip the browser setup

If you need screenshots of bakery pages to review menus, opening-hour notices or website changes, ScreenshotNeo can capture a URL with one API call. It removes cookie banners, popups and chat widgets before the shot; bot checks, blank pages and failed loads are never billed; its MCP server lets AI agents take screenshots; and 1,000 screenshots a month are free with no card, with paid plans starting at $5 for 3,000. See the API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Use the exact target bakery URL in place of the example. The response is an image or PDF capture; it does not verify that a bakery is open or replace a structured business dataset.

Create a free account for 1,000 screenshots a month, with no card required.

FAQ

Can I honestly say the scrape contains every bakery in Paris?

Only if your source coverage and inclusion rule justify that claim. The sources covered here do not establish a complete, current official bakery count. Prefer a dated claim that describes the source and coverage.

Can the Paris cycle infrastructure dataset produce turn-by-turn routes by itself?

No. The City’s layer omits ordinary streets shared by default with motor traffic. Use a complete street graph for routing and join the dated infrastructure layer as context.

Do bicycle counters show which route is safest?

No. They are observations at specific monitored locations and do not cover all roads or all vehicle types. Use them as context alongside infrastructure and the routing profile.

How often should bakery hours be refreshed?

Refresh close to the ride or publication date, with extra attention to early openings, holidays and uncertain records. Hours and business status can change.

Sources