Self-hosted performance validation

This directory specifies future performance checks on a dedicated runner. There is no active performance release gate. Functional, concurrency and resource-lifecycle tests remain required independently of benchmark results. Historical measurements and cross-framework comparisons are not distributed with this repository.

Runner requirements

Use a reserved Linux host with a recorded CPU model, governor, memory budget, kernel, interpreter build and service image digests. Keep load generation separate from server CPU allocation. Run trusted source only: pull requests from forks must never execute on a persistent self-hosted runner with secrets. Create a disposable database and output directory for each run; clean up only resources whose ownership the harness can verify.

Scenario matrix

Scenario Required variants Correctness checks
Raw response Small and large JSON, no database Status, headers and decoded payload
Serialization Scalar, nested and collection fields; cold and warm plans DRF output/errors; custom-field fallback; object isolation
Optional optimizations Compiler, field cache, recursive copy and batching separately, then combined Default-disabled path and equivalent data access
Database List/detail/create, relation loading, pagination Query count, result order, permissions and writes
I/O concurrency Controlled async HTTP; database plus HTTP Concurrency limits, timeouts, cancellation and loop delay
Streaming SSE/NDJSON, slow clients and disconnects Producer closure, backpressure, bounded retained state
Runtime Supported ordinary and free-threaded Python Actual GIL state, warnings and dependency compatibility
Deployment Uvicorn + Nginx + PostgreSQL Worker/pool budgets and equivalent proxy configuration

Measurement protocol

Build the baseline and candidate wheels from identified source commits. Use the same interpreter, resolved dependencies, settings, dataset and resources for both. Keep installation outside timed regions. Alternate baseline/candidate process order across independent starts; warm each process and retain individual samples. A first release uses repeated identical-wheel runs to characterize noise, not a fabricated previous-release baseline.

Record throughput, latency distribution, errors, CPU, RSS, loop delay, database connections and query counts where applicable. Separate profiler runs from uninstrumented timing. Store wheel hashes, dependency versions, commands, environment metadata and raw samples alongside the report in a uniquely named results directory in the future benchmark repository.

Release policy prerequisites

Do not enable a blocking percentage threshold before runner calibration. Determine the repeatability envelope per scenario, then choose and document a material regression threshold outside that envelope using repeated A/A and known-regression trials. Evaluate scenarios independently. Missing samples, failed correctness checks or excessive noise must not be reported as a pass. Enabling publication blocking and an exception-approval environment requires a separate reviewed workflow change and repository configuration.

For current functional checks, see the deployment validation guide and release procedure.