linux database/postgresql database/databaseengineering
Postgres is half as fast in Linux 7.0, and we always knew why
Core Idea
PostgreSQL runs about half as fast on Linux 7.0 because the kernel dropped
PREEMPT_NONEand can now preempt the process that is holding a Postgres spinlock, leaving every other worker busy-waiting on it.
- Symptom: roughly 50% lower transaction throughput and doubled query latency, most visible on AWS Graviton4 (ARM64) and other high-core CPUs.
- Chain reaction: aggressive preemption → the
s_lockholder gets preempted mid minor page fault → dozens of processes spin waiting, burning up to 55% of CPU for nothing.- Fixes:
huge_pages = on(best, since fewer page faults mean fewer preemptions), boot withpreempt=nonewhenCONFIG_PREEMPT_DYNAMICis available, or pin workers / dropmax_parallel_workers_per_gatheruntil Postgres adopts RSEQ timeslices.
(PostgreSQL) experiences a ~50% drop in transaction throughput and a doubling of average query latency on the Linux 7.0 kernel. This severe regression was notably observed on AWS Graviton4 (ARM64) instances but can impact other high-core architectures.
Root Cause
- Scheduler Changes: Linux 7.0 removed the
PREEMPT_NONEscheduling model to simplify the kernel, defaulting instead to a more aggressive preemption model likePREEMPT_LAZY. - Spinlock Contention: PostgreSQL processes use userspace spinlocks (
s_lock) to protect shared memory buffers. Under the new scheduler, processes holding these locks can be interrupted (preempted) by the kernel, especially when minor page faults occur. - Cascading Wait Times: When the lock holder is preempted, dozens of other concurrent processes fall into a busy-wait loop, which can waste up to 55% of total CPU time just waiting for the lock to be released.
Workarounds & Mitigations
- Enable Huge Pages (Best Solution): Configuring PostgreSQL to use Huge Pages (
huge_pages = on, utilizing 2MB or 1GB pages instead of default 4KB pages) drastically reduces the frequency of minor page faults, which largely eliminates the lock-holding preemption issue. - Modify Kernel Boot Parameters: If the Linux distribution’s kernel was compiled with
CONFIG_PREEMPT_DYNAMIC, you can addpreempt=noneto the kernel boot parameters to restore the legacy scheduling behavior. - Future Code Fixes: Linux kernel maintainers recommend that PostgreSQL developers implement the Restartable Sequences (RSEQ) timeslice extension, a Linux 7.0+ feature that allows applications to request a grace period to avoid being preempted during critical sections. However, this is not yet implemented in PostgreSQL. Temporary database-level tweaks include pinning worker processes to dedicated CPU cores or setting
max_parallel_workers_per_gather = 0to lower contention.