How Vacuum and Autovacuum Operations Work in pgrust
Vacuum and autovacuum operations in pgrust replicate PostgreSQL's C implementation in Rust, using heap_vacuum_rel() as the unified entry point for both manual VACUUM commands and background autovacuum workers, with cost-delay mechanisms and eager-freeze optimizations governed by GUC parameters.
pgrust implements PostgreSQL's storage maintenance logic in pure Rust, providing compatible vacuum and autovacuum operations that handle dead tuple cleanup, visibility map updates, and transaction-ID freezing. The architecture separates manual vacuum execution invoked by client commands from the autonomous background processes that spawn workers based on table activity thresholds.
Manual VACUUM Execution
Manual VACUUM operations begin at the heap_vacuum_rel() function defined in crates/backend/access/heap/vacuumlazy/src/vacuum_rel.rs. This entry point initializes the LVRelState struct to track scan progress, records timestamps, and opens indexes via vac_open_indexes() before driving the heap scan.
The process reports progress through pgstat_progress_start_command() and evaluates whether to perform an eager scan by calling heap_vacuum_eager_scan_setup(). During the main phase, lazy_scan_heap walks heap pages to remove dead tuples, updates the visibility map, and may truncate empty pages if the vacuum_truncate GUC is enabled.
Autovacuum Architecture
The autovacuum system consists of a launcher and worker processes. The launcher runs in am_autovacuum_launcher_process() located in crates/backend/postmaster/postmaster_autovacuum/src/lib.rs, spawning a configurable number of workers that sleep according to autovacuum_naptime.
Each worker calls heap_vacuum_rel() under the same code path as manual VACUUM, but only executes when per-table thresholds are met. Workers initialize the same LVRelState structure but set the instrument flag when log_min_duration is configured to enable detailed logging. Unlike manual operations, autovacuum workers typically run in aggressive mode, which disables eager scanning to prioritize completion speed. After completion, workers update pgstat_relation with last_autovacuum_time and autovacuum_count statistics.
Cost-Delay and Resource Management
Vacuum operations implement cost-delay logic to prevent CPU monopolization. The system tracks a cost balance against limits defined in crates/backend/utils/misc/guc_tables/src/tables.rs, including:
vacuum_cost_page_hitandvacuum_cost_page_missfor buffer access costsvacuum_cost_limitfor manual operationsautovacuum_vacuum_cost_limitandautovacuum_vacuum_cost_delayfor background workers
Memory consumption is constrained by vacuum_buffer_usage_limit for manual operations and autovacuum_work_mem (falling back to maintenance_work_mem) for workers. Freeze thresholds prevent transaction-ID wraparound through autovacuum_freeze_max_age and autovacuum_multixact_freeze_max_age, forcing vacuum when table age exceeds these values.
Eager Scan Optimization
The eager scan optimization attempts to freeze pages early by querying the visibility map through visibilitymap_count before performing a full table scan. This optimization activates only when the relation is sufficiently large, the operation is not aggressive, and freeze limits have been exceeded, reducing I/O for tables with mostly frozen pages.
Implementation Examples
The following examples demonstrate how to invoke vacuum operations and configure autovacuum parameters using the pgrust API.
// Example 1: Run a manual VACUUM on a table using the pgrust API
use pgrust::backend::access::heap::vacuumlazy::vacuum_rel::heap_vacuum_rel;
use pgrust::types::{Relation, Mcx, BufferAccessStrategy, VacuumParams};
// Assume `mcx` is a valid memory context and `rel` is an opened heap relation.
let params = VacuumParams {
options: VACOPT_VERBOSE, // Show detailed progress
max_eager_freeze_failure_rate: 0.03, // Enable eager scanning (default)
..Default::default()
};
let bstrategy = BufferAccessStrategy::default();
heap_vacuum_rel(mcx, rel, ¶ms, bstrategy)
.expect("VACUUM failed");
// Example 2: Adjust autovacuum settings at runtime
use pgrust::utils::misc::guc_tables::vars::{
autovacuum_naptime, autovacuum_vacuum_threshold, autovacuum_vacuum_cost_delay,
};
autovacuum_naptime::set(30); // Sleep 30 s between autovacuum cycles
autovacuum_vacuum_threshold::set(100); // Require 100 dead tuples before vacuum
autovacuum_vacuum_cost_delay::set(5.0); // 5 ms delay per cost limit breach
// Example 3: Query autovacuum statistics for a table
use pgrust::utils::activity::pgstat_relation::pgstat_fetch_stat_dbentry;
let db_id = 1; // OID of the database
if let Some(entry) = pgstat_fetch_stat_dbentry(db_id) {
println!("Last autovacuum: {}", entry.last_autovacuum_time);
println!("Total autovacuum time: {} ms", entry.total_autovacuum_time);
}
Key Source Files
The vacuum and autovacuum implementation spans several critical modules:
crates/backend/access/heap/vacuumlazy/src/vacuum_rel.rs– Containsheap_vacuum_rel()and eager scan setup logiccrates/backend/access/heap/vacuumlazy/src/vacuum_phase.rs– Handles vacuum phase transitions and state managementcrates/backend/postmaster/postmaster_autovacuum/src/lib.rs– Implements the autovacuum launcher and worker orchestrationcrates/backend/utils/misc/guc_tables/src/tables.rs– Defines GUC parameters for vacuum and autovacuum tuningcrates/backend/utils/activity/pgstat_relation/src/lib.rs– Tracks autovacuum statistics and timing inpgstat_relation
Summary
- Unified entry point: Both manual VACUUM and autovacuum workers execute through
heap_vacuum_rel()invacuum_rel.rs. - Eager scan optimization:
heap_vacuum_eager_scan_setup()enables early freezing for large tables when not in aggressive mode, utilizing the visibility map. - Autovacuum triggers: The launcher process in
postmaster_autovacuum/src/lib.rsspawns workers based on dead tuple thresholds andautovacuum_naptimeintervals. - Resource governance: Cost-delay mechanisms and GUC parameters in
guc_tables.rsprevent vacuum operations from overwhelming system resources. - Statistics tracking: Autovacuum activity is logged via
pgstat_relationwithlast_autovacuum_timeandautovacuum_countmetrics.
Frequently Asked Questions
What is the difference between manual VACUUM and autovacuum in pgrust?
Manual VACUUM is invoked directly by client SQL commands through heap_vacuum_rel(), while autovacuum runs as a background daemon through am_autovacuum_launcher_process() that spawns workers when tables exceed configured dead tuple thresholds. Both use the same core scanning logic, but autovacuum workers typically operate in aggressive mode with eager scanning disabled and respect separate cost-delay limits defined by autovacuum_vacuum_cost_delay and autovacuum_vacuum_cost_limit.
How does the eager scan optimization work in pgrust vacuum operations?
The eager scan optimization, implemented in heap_vacuum_eager_scan_setup(), attempts to freeze pages early by querying the visibility map through visibilitymap_count before performing a full table scan. This optimization activates only when the relation is sufficiently large, the operation is not aggressive, and freeze limits have been exceeded, reducing I/O for tables with mostly frozen pages.
What GUC parameters control autovacuum behavior in pgrust?
Autovacuum behavior is governed by parameters defined in crates/backend/utils/misc/guc_tables/src/tables.rs, including autovacuum_naptime (sleep interval between checks), autovacuum_vacuum_threshold (minimum dead tuples required), autovacuum_vacuum_cost_delay and autovacuum_vacuum_cost_limit (cost-delay throttling), and autovacuum_work_mem (memory allocation for workers).
Where does the autovacuum launcher spawn worker processes?
The autovacuum launcher spawns workers in am_autovacuum_launcher_process() within crates/backend/postmaster/postmaster_autovacuum/src/lib.rs. This function maintains a pool of worker processes that periodically wake according to autovacuum_naptime, scan the system catalog for tables needing maintenance, and invoke heap_vacuum_rel() to perform the actual vacuuming.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →