Cache connection resilience

Redis and Valkey native clients already reconnect disconnected sockets. The aiodrf backends retain the client and pool; they do not recreate a pool after each error. No additional reconnect wrapper, background monitor or monkeypatch is needed. Configure the driver through Django's CACHES setting as described in async caching.

Reconnection and command retries

These are different operations:

Mechanism Behavior Limitation
Lazy reconnection Opens a socket when a later command borrows a disconnected connection Does not prove whether an earlier write succeeded
health_check_interval Checks a connection before use after the driver's idle interval Not a background heartbeat; adds a round trip when due
retry and retry_on_error Replays a failed command with a bounded native-driver backoff A lost reply can cause a write to execute twice
Sentinel Discovers a new primary after failover In-flight requests can still fail; replication is not exactly-once delivery
Cluster Handles routing, redirects and topology changes in the driver Topology retries and partial pipelines have separate semantics

aiodrf adds no command retries by default. The tested standalone pool path does not replay a command unless configured. redis-py and valkey-py retry Cluster commands by default; the native backends pass a retry without retries for Cluster unless retry, cluster_error_retry_attempts or connection_error_retry_attempts is configured, so a replayed aincr() or aadd() cannot apply twice. Redirects and topology refreshes are the driver's routing, not command retries, and are unchanged. Sentinel defaults can still differ between driver versions, and a driver's direct client constructor can have different defaults from an explicit pool. When the retry budget matters, set it explicitly and test the deployed version. The driver references are redis-py retries and valkey-py retries.

The backend remains usable after a connection error. A caller that receives an error may make another independent request on the same instance. Explicit aclose() is different: a contrib backend closed at lifespan shutdown cannot reopen. Create another instance for another lifespan.

Opt-in retries for repeatable operations

Use an async Retry object, not the similarly named synchronous class. For aiodrf.contrib.redis.AsyncRedisCache, the following defines a separate alias for callers whose operations tolerate replay. Start from the native alias in the cache guide:

from redis.asyncio.retry import Retry
from redis.backoff import FullJitterBackoff
from redis.exceptions import ConnectionError, TimeoutError

CACHES["native_reads"] = {
    **CACHES["native"],
    "OPTIONS": {
        **CACHES["native"]["OPTIONS"],
        "socket_connect_timeout": 1,
        "socket_timeout": 1,
        "health_check_interval": 15,
        "retry": Retry(FullJitterBackoff(base=0.05, cap=0.5), retries=2),
        "retry_on_error": [ConnectionError, TimeoutError],
    },
}

For aiodrf.contrib.valkey.AsyncValkeyCache, use the identical lowercase options with Retry from valkey.asyncio.retry, FullJitterBackoff from valkey.backoff and the exceptions from valkey.exceptions. Never mix Redis and Valkey objects. The alias name is descriptive, not an access-control policy. Restrict its callers to repeatable operations, or enforce appropriate server ACLs. Each alias owns a separate pool, so account for both pools in the connection budget.

Two retries permit up to three attempts at the driver's retry boundary. The overall request can include connection setup, health checks, redirects and backoff; it is not bounded solely by socket_timeout. Apply an application deadline when needed:

import asyncio


async def load_summary(cache):
    async with asyncio.timeout(3):
        return await cache.aget("summary")

cache here is the lifespan-owned native_reads instance. Cancellation propagates; native retry backoff yields to the event loop. Do not use negative retry counts (unbounded retry), and do not retry authentication errors or every exception indiscriminately. See the native asyncio client documentation for connection ownership and cleanup.

django-valkey vendor backend

For django_valkey.async_cache.cache.AsyncValkeyCache, native connection options belong in CONNECTION_POOL_KWARGS, not at the top of OPTIONS. Keep the lifespan-owned factory and CLOSE_CONNECTION=True from the cache guide:

from valkey.asyncio.retry import Retry
from valkey.backoff import FullJitterBackoff
from valkey.exceptions import ConnectionError, TimeoutError

options = CACHES["native"]["OPTIONS"]
options["CONNECTION_POOL_KWARGS"] = {
    **options.get("CONNECTION_POOL_KWARGS", {}),
    "health_check_interval": 15,
    "retry": Retry(FullJitterBackoff(base=0.05, cap=0.5), retries=2),
    "retry_on_error": [ConnectionError, TimeoutError],
}

Apply this only to a replay-tolerant alias. This uses the vendor's existing configuration extension, not an aiodrf-specific retry API. See django-valkey async configuration.

Writes and failure reporting

A timeout or connection error after INCR, EVAL, SET NX or a pipeline does not establish that the command failed on the server. Retrying can increment twice, change an add() result, extend a TTL or overwrite a concurrent writer. Even ordinary SET is not necessarily replay-safe for an application's consistency requirements. Native I/O cannot provide exactly-once execution.

For standalone/Sentinel aliases containing such writes, explicitly disable command replay while retaining lazy socket reconnection:

from redis.asyncio.retry import Retry
from redis.backoff import NoBackoff

CACHES["native"]["OPTIONS"].update(
    retry=Retry(NoBackoff(), retries=0),
    retry_on_error=[],
)

Use the equivalent valkey imports for contrib Valkey; for the vendor backend, put these options in CONNECTION_POOL_KWARGS. These snippets do not configure Cluster's separate topology retry budget. Verify Cluster's settings and failure behavior against the installed driver; do not infer a no-replay guarantee from retry=0 alone. Sentinel discovery options live separately in sentinel_kwargs.

Failures remain exceptions, not fabricated cache misses. aiodrf does not add a global circuit breaker or hide write failures. If cache reads are optional, handle the relevant driver exceptions in the application with a bounded fallback and observability. Session storage, security counters and distributed coordination usually require a different failure policy from response caching.

Verification

The Redis and isolated Valkey service profiles cover:

  • Socket disconnection followed by successful work through the same pool.
  • A lost reply after the server accepted a write, without a duplicate increment when replay is disabled.
  • A repeatable read recovering with retries explicitly enabled.
  • Exhaustion of a two-retry budget and propagation of the final exception.
  • Cancellation during backoff with a single-connection pool, followed by a successful command proving that the connection was returned.

These behaviours hold for both django-valkey's backend and aiodrf's own backends, and Sentinel failover and Cluster commands are tested separately. The failures are simulated by discarding a reply the client has already read; the tests do not reproduce network partitions, so they say nothing about recovery time or availability in production.