How the DeusData Background Watcher Detects and Processes File Changes for Automatic Synchronization
The DeusData codebase-memory-mcp server runs a dedicated background watcher thread that polls registered Git repositories every 5 seconds (adaptive based on file count), executes git status --porcelain to detect working tree changes, and automatically triggers a full re-index via the watcher_index_fn callback whenever modifications are found.
The DeusData codebase-memory-mcp repository implements a robust background watcher mechanism that keeps your codebase search index synchronized with on-disk changes without manual intervention. This git-aware polling system continuously monitors project roots in a separate POSIX thread, validating file system state and automatically triggering re-indexing pipelines when modifications are detected.
Watcher Initialization and Thread Launch
When the MCP server starts, it instantiates the global watcher object at line 771 of src/main.c:
g_watcher = cbm_watcher_new(watch_store, watcher_index_fn, NULL);
The cbm_watcher_new function accepts three parameters: the persistent store for project root resolution, the watcher_index_fn callback that executes full re-indexes (defined at lines 59-78), and optional user data.
Immediately after creation, the server spawns a dedicated POSIX thread to run the watcher loop asynchronously:
static void *watcher_thread(void *arg) {
cbm_watcher_t *w = arg;
#define WATCHER_BASE_INTERVAL_MS 5000
cbm_watcher_run(w, WATCHER_BASE_INTERVAL_MS);
return NULL;
}
This thread executes independently of the main server loop, ensuring that file system monitoring never blocks request handling.
The Core Polling Loop in cbm_watcher_run
The function cbm_watcher_run (implemented at src/watcher/watcher.c:18-45) serves as the heart of the background synchronization system. It enters a while loop that continues until the atomic stopped flag is set:
while (!atomic_load(&w->stopped)) {
cbm_watcher_poll_once(w);
// Sleep in chunks to maintain responsiveness
}
Rather than sleeping for the full interval duration, the loop breaks sleep into smaller chunks. This design allows the watcher to respond immediately to shutdown signals while still maintaining the configured polling frequency.
Thread-Safe Project Snapshotting
Inside cbm_watcher_poll_once (lines 681-706), the watcher creates a snapshot of all registered projects to isolate the critical section from I/O operations. It allocates an array of project pointers and populates it using cbm_ht_foreach with a snapshot callback:
project_state_t **snap = malloc(n * sizeof(project_state_t *));
cbm_ht_foreach(w->projects, snapshot_project, &sc);
...
for (int i = 0; i < sc.count; i++) {
poll_project(NULL, snap[i], &ctx);
}
This snapshotting technique ensures that the global hash-table lock (w->projects) is held only briefly during the copy operation, while the expensive poll_project function executes outside the critical section.
Detecting File Changes in poll_project
The poll_project function performs a multi-stage validation process to determine if a repository requires re-indexing:
Root Validation. First, the function verifies the project root still exists at src/watcher/watcher.c:572-605. If the directory disappeared, it invokes prune_missing_project to remove the stale entry from the watch list.
Git Repository Verification. The watcher explicitly ignores non-Git projects, returning early if !s->is_git (lines 618-620). This design optimizes for Git-based workflows where change detection relies on SHA comparison.
Adaptive Interval Gating. Each project maintains its own next_poll_ns timestamp. If the current time is earlier than this value, the poll is skipped to prevent excessive I/O on large repositories (lines 622-624).
Dirty Check Execution. When the interval permits, the watcher executes git status --porcelain (plus submodule checks) via the git_is_dirty function (lines ~152-188) to detect uncommitted changes in the working tree.
Triggering Automatic Re-Index Operations
When git_is_dirty returns true, the watcher immediately invokes the index_fn callback (lines 628-646):
ctx->w->index_fn(s->project_name, s->root_path, ctx->w->user_data);
After the callback completes successfully, the watcher updates the project's baseline state by recording the current HEAD hash and file count, then recalculates the next poll interval based on repository size.
Adaptive Polling Intervals
To prevent performance degradation on massive repositories, the watcher implements an adaptive interval calculation in cbm_watcher_poll_interval_ms (lines 98-104):
#define POLL_BASE_MS 5000
#define POLL_FILE_STEP 500 // +1s per 500 files
#define POLL_MAX_MS 60000
int cbm_watcher_poll_interval_ms(int file_count) {
int extra = (file_count / POLL_FILE_STEP) * 1000;
int interval = POLL_BASE_MS + extra;
return (interval > POLL_MAX_MS) ? POLL_MAX_MS : interval;
}
This algorithm adds one second of delay for every 500 tracked files, capping the maximum interval at 60 seconds. A repository with 10,000 files polls every 25 seconds, while small repositories maintain the 5-second base rate.
Graceful Shutdown Handling
The watcher respects the global atomic flag g_shutdown. When the server receives a termination signal, request_shutdown (lines 72-80) sets this flag and calls cbm_watcher_stop, which atomically marks the watcher as stopped. The next iteration of the polling loop detects this change and exits cleanly without orphaning resources.
Summary
- The watcher operates in a dedicated POSIX thread spawned at server startup in
src/main.c. - It polls every 5 seconds by default, dynamically scaling up to 60 seconds based on the number of tracked files per repository.
- Change detection relies exclusively on
git status --porcelain, filtering out non-Git projects automatically. - Thread-safe snapshotting isolates hash-table locks from file system I/O to prevent contention.
- Upon detecting changes, the watcher executes
watcher_index_fnto trigger a supervised re-indexing pipeline automatically.
Frequently Asked Questions
How often does the DeusData watcher check for file changes?
The background watcher polls every 5 seconds by default, but it calculates an adaptive interval using cbm_watcher_poll_interval_ms that adds one second of delay for every 500 tracked files, up to a maximum of 60 seconds. This prevents excessive CPU and I/O usage on very large repositories.
What happens when the watcher detects a modification in a Git repository?
When git_is_dirty reports uncommitted changes, the watcher immediately invokes the watcher_index_fn callback that was registered during initialization at src/main.c:771. This callback acquires a pipeline lock and executes cbm_pipeline_run to perform a full re-index of the modified project, then updates the stored HEAD hash and file count baseline.
How does the watcher handle non-Git projects or deleted directories?
The watcher explicitly ignores non-Git projects by checking is_git at src/watcher/watcher.c:618-620 and returning early. For directories that have been deleted since the last poll, the root_status check triggers prune_missing_project, which removes the stale project entry from the internal hash table to prevent unnecessary polling attempts.
How does the background watcher shut down gracefully when the server stops?
The server signals shutdown by setting the atomic g_shutdown flag via request_shutdown at src/main.c:72-80, which then calls cbm_watcher_stop. This function atomically sets the stopped flag in the watcher structure, causing the while (!atomic_load(&w->stopped)) loop in cbm_watcher_run to exit cleanly on its next iteration.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →