Measuring performance¶
Measure the whole request before choosing an optimization. Middleware, database round trips, relation loading, serializer hooks, encoding and calls to other services each dominate different workloads, and an async view does not make the database driver asynchronous.
Separate read and write serializers¶
Input and output can use different serializers without changing any DRF field.
After validating and saving, represent write.instance with a dedicated read
serializer through await aio.data(). The
serializer backend example does
this while keeping aperform_create(), the request context, success headers and
the usual errors. Reload saved relations when their order or signal handlers
affect the output: validated many-to-many input is not necessarily what was
saved.
Calling another serializer's synchronous .data from to_representation()
keeps DRF's behaviour and does not use aiodrf's compiled serializers. The
selective optimization guide explains how to
choose compiled output, field caching and copy plans independently.
Tools¶
| Question | Tool | Notes |
|---|---|---|
| Where does a request spend its time? | Pyinstrument | Shows await stacks and the event-loop thread; not a complete profile of worker threads |
| Which Python functions use CPU, and how often are they called? | Yappi with its CPU clock | Includes worker threads; use call counts and self time, and do not add up overlapping cumulative times |
| How much native CPU work is done? | perf stat, Valgrind Cachegrind |
Attach to the application process and exclude start-up; simulated instruction counts are not latency |
| Where is memory allocated? | Memray | Shows allocation volume and stacks; allocated bytes alone do not indicate a leak |
| How expensive is building serializer fields? | Timing in your application | Measure the first (cold) construction separately from later (warm) copies |
| How fast are input validation and output? | Your contract tests with a profiler | Include serializers that fall back to DRF |
| How does the deployment behave? | Your server, proxy and database | See deployment behaviour |
Install profilers only in the environment you are inspecting, and use a dedicated database; they are not dependencies of aiodrf.
Pick the tool for the question. Time spent inside sync_to_async in a sampling
profile does not show that switching threads is slow: the worker thread may be
validating input or running SQL. Inspect worker CPU and query timings
separately. aiodrf.test.count_hops() counts aiodrf's switches to a worker
thread, not Django's or asgiref's, so compare it with whole-process call counts
when you assess thread overhead. Keep hop counting and SQL logging off while you
measure throughput.
When you profile an ASGI application directly, note what is missing compared with production: the HTTP transport, the production event loop and concurrent clients. Profile complete ASGI requests, including disconnects, rather than a view method alone, and confirm improvements with separate, uninstrumented HTTP runs.
Request handling and rendering¶
aiodrf recognizes coroutine functions from their code flags without keeping references to callable instances. Marked functions, partials and callable objects are inspected with the general coroutine checks. A synchronous decorator around an async hook is not assumed to be safe on the event loop.
Profile request start-up, response rendering and cleanup as well as the
handler. Django's request signals, database connection cleanup and the async
signal receivers of installed packages can switch threads even when the view
makes no query. Removing signal receivers, or sharing one executor across
requests, changes Django's lifecycle and isolation guarantees; aiodrf does
neither on your behalf, apart from the opt-in REQUEST_THREADS below.
Before rendering JSON on the event loop, aiodrf inspects the response data. Data made only of built-in types is encoded on the event loop; other values, custom time zones and types with custom metaclasses are encoded in a worker thread. The inspection never evaluates lazy strings, querysets or application callbacks. It runs for every response, because middleware and view code can modify the data. Data of built-in types is recognized quickly, in C; other data is walked in Python, which costs CPU time on large nested payloads. A built-in JSON renderer instance with changed attributes is always rendered in a worker thread, because a replaced method or encoder can perform I/O. Each item of a streaming response is handled the same way.
Request threads¶
Under Django's ASGI handler, every request starts a thread for its synchronous
code, needed at least for Django's request_started and request_finished
signals, and another thread to wait for it. On inexpensive endpoints, starting
these two threads is a noticeable part of the CPU time of a request. With
AIODRF["REQUEST_THREADS"] and aiodrf.asgi.get_asgi_application(), the
threads are kept for later requests; each still serves one request at a time.
See the setting and its trade-offs
in the tuned profile.
Memory and garbage collection¶
DRF's objects refer to each other: the view to its request and response and back, a list serializer's child to the list, bound fields to their serializer. Such reference cycles are freed by Python's cycle collector rather than by reference counting, so a request's body, its data and the instances it serialized stay in memory until the next collection, and each collection has to scan every request in progress.
- aiodrf's
Response.close()breaks these references once the response has been sent: between the view, the request and the response, and in the serializer whoseReturnListorReturnDictthe response returned (its cachedfields, which are rebuilt if read, and its list's child). The limitations list what remains readable afterwards. aiodrf.contrib.list_serializersprovides list serializers whose child refers to the list weakly, for data that does not reach a response as a DRFReturnList(Response({"items": list(serializer.data)})) and for serializers used outside requests. Set one asMeta.list_serializer_class, or asdefault_list_serializer_classon a base serializer;SchemaListSerializeris the one for msgspec and Pydantic serializers.AIODRF["MONKEYPATCHES"]applies the same to all of DRF's classes in the process (weak_list_children,release_drf_responses), for projects that also run DRF's own views or third-party list serializers. It patches DRF and must be enabled explicitly. Itscache_model_field_infoandkeep_json_encoderspatches avoid work DRF repeats for each serializer and each response (setting).
aiodrf never calls gc.freeze(), disables collection or changes the collector's
thresholds. Python documents gc.freeze() for a specific
pre-fork procedure, not
for a running server worker, and freezing warmed-up request objects can keep
resources alive. Change garbage collection settings only when measurements of
your application's memory justify it.
Comparing configurations¶
- Keep the interpreter, dependencies, number of workers, proxy, payloads and queries identical, and record package versions and settings with the results.
- Alternate the order in which the configurations run, warm both up, and record several independent runs with their individual samples.
- Check statuses, response bodies and query counts before comparing speed: a fast error response is not an improvement.
- Report average and tail latency, throughput, event-loop delay, memory and connection limits as relevant, and do not combine unrelated workloads into a single score.
- Measure serializer compilation separately from warm execution, and compiled
serializers separately from those that fall back to DRF. Do not time code
under
override_settings(): changing settings clears aiodrf's caches. - CPU pinning, noisy shared machines and background load are part of the environment. Profilers show where work happens; repeated uninstrumented runs measure it.