linux database/postgresql database/databaseengineering

Core Idea

PostgreSQL runs about half as fast on Linux 7.0 because the kernel dropped PREEMPT_NONE and can now preempt the process that is holding a Postgres spinlock, leaving every other worker busy-waiting on it.

  • Symptom: roughly 50% lower transaction throughput and doubled query latency, most visible on AWS Graviton4 (ARM64) and other high-core CPUs.
  • Chain reaction: aggressive preemption → the s_lock holder gets preempted mid minor page fault → dozens of processes spin waiting, burning up to 55% of CPU for nothing.
  • Fixes: huge_pages = on (best, since fewer page faults mean fewer preemptions), boot with preempt=none when CONFIG_PREEMPT_DYNAMIC is available, or pin workers / drop max_parallel_workers_per_gather until Postgres adopts RSEQ timeslices.

(PostgreSQL) experiences a ~50% drop in transaction throughput and a doubling of average query latency on the Linux 7.0 kernel. This severe regression was notably observed on AWS Graviton4 (ARM64) instances but can impact other high-core architectures.

Root Cause

  • Scheduler Changes: Linux 7.0 removed the PREEMPT_NONE scheduling model to simplify the kernel, defaulting instead to a more aggressive preemption model like PREEMPT_LAZY.
  • Spinlock Contention: PostgreSQL processes use userspace spinlocks (s_lock) to protect shared memory buffers. Under the new scheduler, processes holding these locks can be interrupted (preempted) by the kernel, especially when minor page faults occur.
  • Cascading Wait Times: When the lock holder is preempted, dozens of other concurrent processes fall into a busy-wait loop, which can waste up to 55% of total CPU time just waiting for the lock to be released.

Workarounds & Mitigations

  1. Enable Huge Pages (Best Solution): Configuring PostgreSQL to use Huge Pages (huge_pages = on, utilizing 2MB or 1GB pages instead of default 4KB pages) drastically reduces the frequency of minor page faults, which largely eliminates the lock-holding preemption issue.
  2. Modify Kernel Boot Parameters: If the Linux distribution’s kernel was compiled with CONFIG_PREEMPT_DYNAMIC, you can add preempt=none to the kernel boot parameters to restore the legacy scheduling behavior.
  3. Future Code Fixes: Linux kernel maintainers recommend that PostgreSQL developers implement the Restartable Sequences (RSEQ) timeslice extension, a Linux 7.0+ feature that allows applications to request a grace period to avoid being preempted during critical sections. However, this is not yet implemented in PostgreSQL. Temporary database-level tweaks include pinning worker processes to dedicated CPU cores or setting max_parallel_workers_per_gather = 0 to lower contention.