ScreenshotNeo

BlogGuides

Python Cache: How to Speed Up Your Code With Effective Caching Techniques

Learn when Python caching helps, how to use functools.lru_cache, Django and Redis, and how to handle keys, expiry, invalidation and stale data.

By the ScreenshotNeo team29 September 202611 min read

Python Cache: How to Speed Up Your Code With Effective Caching Techniques

Caching speeds up Python code by saving the result of work that is expensive to repeat and reusing it when the same inputs appear again. Start with functools.lru_cache for deterministic functions inside one process. Use Django’s cache framework for web content and Redis or another shared backend when multiple workers need the same entries. A cache is only correct when its key captures every input that affects the result and its expiration or invalidation policy matches how fresh the result must be.

Caching is not automatically faster. It adds lookup, storage, memory and invalidation costs; it can also return stale or user-specific data to the wrong caller if keys are incomplete. Measure the underlying work first, then track cache hits, misses, latency, memory and stale reads after adding it.

1. When should you cache Python work?

Cache a result when the same computation or read is likely to recur, producing it costs more than looking it up, and reusing it is semantically safe. Examples include parsing stable configuration, computing a deterministic transformation, looking up reference data, or rendering a page fragment that changes infrequently.

Do not memoize a function whose output depends on hidden state unless that state is included in the key or invalidation is explicit. A function that reads the current time, a file that changes, an environment variable, a database row or a user’s permissions can return a different result even with identical explicit arguments.

  • CPU-bound repeat work: local memoization may avoid repeating computation.
  • I/O-bound repeat reads: a cache can avoid database or network round trips, but consider freshness and failure behavior.
  • One process: an in-memory cache may be enough.
  • Several processes or hosts: process-local entries are not shared; use a shared backend if workers must see the same value.

Python’s documentation describes lru_cache as useful when an expensive or I/O-bound function is periodically called with the same arguments. It is thread-safe, but concurrent misses for the same key can still invoke the wrapped function more than once. Python functools documentation.

2. Cache a Python function with functools.lru_cache

functools.lru_cache is the simplest starting point for pure or effectively pure functions. It stores recent calls in a process-local mapping and evicts least-recently-used entries after reaching its capacity.

A process-local memoization cache reuses results for repeated inputs.
A process-local memoization cache reuses results for repeated inputs.
from functools import lru_cache

@lru_cache(maxsize=512)
def normalized_country(code: str) -> str:
    # Imagine this lookup involves expensive parsing or stable reference data.
    return code.strip().upper()

print(normalized_country(" us "))  # computes and stores the result
print(normalized_country(" us "))  # reuses the result
print(normalized_country.cache_info())
# CacheInfo(hits=1, misses=1, maxsize=512, currsize=1)

Save this as cache_demo.py and run python cache_demo.py. In real code, the function body must do meaningful work for caching to be useful; this small example makes the cache mechanics easy to see.

Choose a bounded size and inspect behavior

Use a finite maxsize for long-running services unless the input space is known to be small and bounded. The default is 128; a larger value may improve hit rate but uses more memory. Set maxsize=None only when unbounded growth is safe. Use cache_info() to inspect hits, misses, capacity and current entry count, and cache_clear() to remove entries.

from functools import lru_cache

@lru_cache(maxsize=1024)
def product_price(product_id: int) -> int:
    return load_price_from_source(product_id)

# After a source update, explicitly invalidate this process's memoized values.
product_price.cache_clear()

The example assumes load_price_from_source exists in your application. If a particular product changes, clearing the entire cache is simple but may discard useful entries; a cache with explicit keys and deletion may be more appropriate.

Arguments must be hashable

All positional and keyword arguments used by the call must be hashable. Lists and dictionaries are not. Convert inputs to immutable forms when that conversion preserves meaning, or design an explicit cache key.

from functools import lru_cache

@lru_cache(maxsize=256)
def render_columns(columns: tuple[str, ...]) -> str:
    return ", ".join(columns)

print(render_columns(("name", "email")))

Do not blindly turn nested mutable data into tuples if ordering, duplicate values or mutation semantics matter. Normalize deliberately and ensure equivalent inputs map to the same key.

3. Pick a cache scope that matches your application

Approach Scope Good fit Main trade-off
lru_cache One Python process Repeated function calls with hashable inputs Workers do not share entries; no TTL
Django cache framework Backend-dependent Views, fragments, and application data Requires deliberate keys and expiry
Redis or Memcached Shared across processes/hosts Shared working sets and cross-worker reuse Network, serialization and operations

For an alternate eviction policy or cache collection API, a library may be useful, but check its current documentation and version before selecting it. The research for this article does not establish current cachetools behavior or version details.

4. Cache Django pages and data

Django supports per-site, per-view, template-fragment and low-level caching. Its included backends include local memory, database, filesystem, Memcached and Redis, as well as custom backends. Begin with a built-in backend unless your deployment has a clear need for another option. Django cache framework documentation.

Configure a backend

Here is a local-memory configuration for development or a single-process use case. Django’s local-memory backend is thread-safe, private to each process and uses LRU culling; multiple worker processes will not share it.

# settings.py
CACHES = {
    "default": {
        "BACKEND": "django.core.cache.backends.locmem.LocMemCache",
        "LOCATION": "my-application-cache",
        "TIMEOUT": 300,
        "OPTIONS": {
            "MAX_ENTRIES": 1000,
            "CULL_FREQUENCY": 3,
        },
    }
}

TIMEOUT is in seconds. Django documents a default timeout of 300 seconds; None means no expiry and 0 means immediate expiry. Select an expiry based on freshness needs rather than copying the default. MAX_ENTRIES limits entries for relevant backends, while CULL_FREQUENCY controls culling behavior.

Cache a view with a bounded timeout

from django.views.decorators.cache import cache_page

@cache_page(60 * 5)
def public_catalog(request):
    # Return a response that is the same for all callers of this URL.
    ...

For a view that depends on a user, tenant, language, authentication or request headers, ensure the cache varies on those dimensions or use a key that includes them. URL-only caching can expose one person’s content to another. Django documents Vary and cache key behavior in its cache guidance.

Use the low-level API for explicit keys

from django.core.cache import cache

def get_product(product_id):
    key = f"product:v1:{product_id}"
    value = cache.get(key)
    if value is not None:
        return value

    value = load_product(product_id)
    cache.set(key, value, timeout=300)
    return value

A cached value may legitimately be None, so applications that need to cache it should use a unique sentinel or a separate existence check. Versioned keys such as product:v1: can make broad invalidation easier during schema or logic changes.

5. Use Redis when workers need a shared cache

A shared Redis cache can let multiple workers reuse a common working set. Redis’s Python guide demonstrates preloading reference data, serving reads from Redis, synchronizing mutations, deleting keys on deletion and applying a safety-net TTL. In that design, the preload is intentional and a missing key is treated as an error; that is not the right failure policy for every application. Redis Python client guide.

A shared cache lets workers reuse data while the database remains the source of truth.
A shared cache lets workers reuse data while the database remains the source of truth.

Keep the durable source of truth in a database or another authoritative store. A typical write-through or invalidation flow updates the source, then updates or deletes the relevant cache key. Decide what reads do if Redis is unavailable: for many systems, falling back to the source is safer than failing the request, but only if the source can handle that load.

Redis brings network latency, serialization and operational requirements. It may still be preferable when worker-local copies produce inconsistent freshness or duplicate too much expensive work. Compare total cost and measured request latency, not just the apparent speed of an individual lookup.

6. Design cache keys, TTLs and invalidation

Include every result-changing dimension

A key is a compact statement of what makes a value reusable. Include every input that affects the output. For web responses, that can mean user or authorization scope, tenant, locale, query parameters, relevant headers and representation version. Omitting a dimension can serve incorrect or private data.

  • Use stable, unambiguous key formats with namespaces, such as catalog:v2:tenant-42:en.
  • Normalize equivalent inputs consistently so harmless differences do not create duplicate entries.
  • Avoid putting secrets or personal data directly in keys if keys may appear in logs or metrics.
  • Consider cardinality: arbitrary user input or timestamps can create a huge number of nearly useless entries.

Choose freshness rules explicitly

A TTL limits how long an entry can remain present, but it does not guarantee that the next read is current within that interval. If a change must be visible immediately, invalidate or replace the affected entry when the source changes. If brief staleness is acceptable, a finite TTL can serve as a safety net when invalidation is missed.

Invalidation can be targeted by key, grouped with versioned namespaces, or performed broadly. Broad clearing is easy but can cause a burst of misses. Targeted invalidation preserves more hits but requires reliable knowledge of dependencies.

7. Handle concurrency, failures and security

Two requests can miss the same key simultaneously and perform duplicate expensive work. Python explicitly allows this with lru_cache. For expensive high-demand keys, consider request coalescing, a lock or a single-flight pattern. Locks need timeouts and cleanup so a failed worker does not block future requests indefinitely.

Plan for cache failure as a normal operating condition. A cache miss, eviction or outage should usually lead back to the source of truth when that fallback is safe. Protect the source from a sudden traffic surge if the cache disappears; fallback can itself create a thundering herd. Use sensible timeouts, retries only when appropriate, and avoid retry loops that amplify an outage.

Django’s filesystem cache serializes values with pickle. If an attacker can modify cache files, they may falsify trusted content or execute code when values are loaded. Restrict access to cache storage and do not deserialize untrusted data. Treat cache contents as derived and potentially disposable, not as durable records.

8. Measure performance and operating cost

There is no universal percentage speedup for Python caching. The result depends on the work avoided, hit rate, key construction, backend latency, serialization, contention and the size of cached objects. A remote cache can make a cheap local computation slower.

Before and after a change, observe:

  • Hit and miss counts or rates, by cache and key family.
  • Lookup latency and the duration of the underlying function or query.
  • Evictions, current memory use, entry sizes and key cardinality.
  • Stale-read incidents, invalidation delay and cache backend errors.
  • Source-system load during cache outages or deployments.

Estimate the trade-off: storage and operations cost against saved CPU, database load and request time. High hit rate alone does not prove benefit if the miss cost is small or cache reads are slow. Avoid storing large response objects without understanding memory or serialization costs.

9. Troubleshooting common Python cache problems

Symptom Likely cause Fix
TypeError: unhashable type A list, dictionary or other unhashable argument reaches lru_cache. Use immutable arguments or build a deliberate normalized key.
Changes do not appear An entry outlives the source data, or an invalidation path is missing. Set an appropriate TTL and invalidate on writes that require immediate freshness.
Cache hit rate is low Keys contain volatile values, are inconsistently normalized, or requests rarely repeat. Inspect key distributions and reuse patterns before increasing capacity.
Memory keeps growing The cache is unbounded, entries are large, or input cardinality is unexpectedly high. Bound capacity, use expiry or reduce the value/key footprint; monitor eviction and memory.
Several workers disagree Each process has its own local-memory or lru_cache entries. Use a shared backend or accept per-worker caching with a freshness strategy.
Latency got worse Lookup, key generation, network or serialization costs exceed the work saved. Measure hit and miss paths separately; remove caching where it adds overhead.
Database load spikes after deployment Cold starts, cache outage or mass expiration caused many misses. Warm important keys, stagger expiry where suitable, and ensure fallback capacity.
One user sees another’s response The key omitted user, tenant, auth or language context. Correct the key and vary dimensions; purge unsafe entries and review access scope.
Duplicate work occurs on simultaneous misses Concurrent callers compute before either stores the value. Coalesce requests or lock expensive keys; apply timeouts and measure contention.

10. A practical rollout checklist

  1. Identify a measured expensive operation and verify repeated inputs occur.
  2. Write down every input and hidden dependency that changes the result.
  3. Choose process-local, framework or shared cache scope based on deployment topology.
  4. Set a capacity and, where applicable, a TTL based on memory and freshness needs.
  5. Define invalidation for source updates and behavior for backend failures.
  6. Deploy with hit, miss, latency, memory and stale-data monitoring.
  7. Compare end-to-end performance and source load; remove the cache if it does not help.

Or skip the browser setup

If the repeated work you need to cache is website screenshots, ScreenshotNeo provides a one-call website screenshot API. It returns PNG, JPEG, WebP or PDF, and includes caching with a TTL you choose. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot request failed: ${res.status}`);
await Bun.write('shot.webp', res);

ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before the shot; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month, no card required.

Frequently asked questions

Does lru_cache work across Gunicorn or Django workers?

No. Each process has its own memory and therefore its own entries. Use a shared backend if workers must see the same cached values.

Is it safe to cache a function that queries a database?

It can be, if the key identifies the relevant query inputs and updates invalidate or expire affected results. Consider whether the database read is already fast enough that caching adds needless complexity.

Should I cache errors or missing values?

Only when that result is stable enough to reuse and its expiry is appropriate. A transient backend or source failure should not accidentally become a long-lived cached answer.

Does a TTL replace invalidation?

No. TTL bounds how long an entry may persist; invalidation is needed when changes must be visible sooner.

Can caching replace a database?

No. A cache is temporary derived data. Durable records belong in an authoritative storage system.