# How the DeusData Background Watcher Detects and Processes File Changes for Automatic Synchronization

> Discover how the DeusData background watcher detects and processes file changes for automatic synchronization. Learn about its efficient Git polling and re-indexing mechanisms.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: how-to-guide
- Published: 2026-07-07

---

**The DeusData `codebase-memory-mcp` server runs a dedicated background watcher thread that polls registered Git repositories every 5 seconds (adaptive based on file count), executes `git status --porcelain` to detect working tree changes, and automatically triggers a full re-index via the `watcher_index_fn` callback whenever modifications are found.**

The DeusData `codebase-memory-mcp` repository implements a robust **background watcher** mechanism that keeps your codebase search index synchronized with on-disk changes without manual intervention. This git-aware polling system continuously monitors project roots in a separate POSIX thread, validating file system state and automatically triggering re-indexing pipelines when modifications are detected.

## Watcher Initialization and Thread Launch

When the MCP server starts, it instantiates the global watcher object at line 771 of [`src/main.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/main.c):

```c
g_watcher = cbm_watcher_new(watch_store, watcher_index_fn, NULL);

```

The `cbm_watcher_new` function accepts three parameters: the persistent store for project root resolution, the `watcher_index_fn` callback that executes full re-indexes (defined at lines 59-78), and optional user data.

Immediately after creation, the server spawns a dedicated POSIX thread to run the watcher loop asynchronously:

```c
static void *watcher_thread(void *arg) {
    cbm_watcher_t *w = arg;
    #define WATCHER_BASE_INTERVAL_MS 5000
    cbm_watcher_run(w, WATCHER_BASE_INTERVAL_MS);
    return NULL;
}

```

This thread executes independently of the main server loop, ensuring that file system monitoring never blocks request handling.

## The Core Polling Loop in `cbm_watcher_run`

The function `cbm_watcher_run` (implemented at `src/watcher/watcher.c:18-45`) serves as the heart of the background synchronization system. It enters a `while` loop that continues until the atomic `stopped` flag is set:

```c
while (!atomic_load(&w->stopped)) {
    cbm_watcher_poll_once(w);
    // Sleep in chunks to maintain responsiveness
}

```

Rather than sleeping for the full interval duration, the loop breaks sleep into smaller chunks. This design allows the watcher to respond immediately to shutdown signals while still maintaining the configured polling frequency.

## Thread-Safe Project Snapshotting

Inside `cbm_watcher_poll_once` (lines 681-706), the watcher creates a **snapshot** of all registered projects to isolate the critical section from I/O operations. It allocates an array of project pointers and populates it using `cbm_ht_foreach` with a snapshot callback:

```c
project_state_t **snap = malloc(n * sizeof(project_state_t *));
cbm_ht_foreach(w->projects, snapshot_project, &sc);
...
for (int i = 0; i < sc.count; i++) {
    poll_project(NULL, snap[i], &ctx);
}

```

This snapshotting technique ensures that the global hash-table lock (`w->projects`) is held only briefly during the copy operation, while the expensive `poll_project` function executes outside the critical section.

## Detecting File Changes in `poll_project`

The `poll_project` function performs a multi-stage validation process to determine if a repository requires re-indexing:

**Root Validation.** First, the function verifies the project root still exists at `src/watcher/watcher.c:572-605`. If the directory disappeared, it invokes `prune_missing_project` to remove the stale entry from the watch list.

**Git Repository Verification.** The watcher explicitly ignores non-Git projects, returning early if `!s->is_git` (lines 618-620). This design optimizes for Git-based workflows where change detection relies on SHA comparison.

**Adaptive Interval Gating.** Each project maintains its own `next_poll_ns` timestamp. If the current time is earlier than this value, the poll is skipped to prevent excessive I/O on large repositories (lines 622-624).

**Dirty Check Execution.** When the interval permits, the watcher executes `git status --porcelain` (plus submodule checks) via the `git_is_dirty` function (lines ~152-188) to detect uncommitted changes in the working tree.

## Triggering Automatic Re-Index Operations

When `git_is_dirty` returns true, the watcher immediately invokes the `index_fn` callback (lines 628-646):

```c
ctx->w->index_fn(s->project_name, s->root_path, ctx->w->user_data);

```

After the callback completes successfully, the watcher updates the project's baseline state by recording the current `HEAD` hash and file count, then recalculates the next poll interval based on repository size.

## Adaptive Polling Intervals

To prevent performance degradation on massive repositories, the watcher implements an adaptive interval calculation in `cbm_watcher_poll_interval_ms` (lines 98-104):

```c
#define POLL_BASE_MS   5000
#define POLL_FILE_STEP 500   // +1s per 500 files
#define POLL_MAX_MS    60000

int cbm_watcher_poll_interval_ms(int file_count) {
    int extra = (file_count / POLL_FILE_STEP) * 1000;
    int interval = POLL_BASE_MS + extra;
    return (interval > POLL_MAX_MS) ? POLL_MAX_MS : interval;
}

```

This algorithm adds one second of delay for every 500 tracked files, capping the maximum interval at 60 seconds. A repository with 10,000 files polls every 25 seconds, while small repositories maintain the 5-second base rate.

## Graceful Shutdown Handling

The watcher respects the global atomic flag `g_shutdown`. When the server receives a termination signal, `request_shutdown` (lines 72-80) sets this flag and calls `cbm_watcher_stop`, which atomically marks the watcher as stopped. The next iteration of the polling loop detects this change and exits cleanly without orphaning resources.

## Summary

- The watcher operates in a dedicated POSIX thread spawned at server startup in [`src/main.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/main.c).
- It polls every 5 seconds by default, dynamically scaling up to 60 seconds based on the number of tracked files per repository.
- Change detection relies exclusively on `git status --porcelain`, filtering out non-Git projects automatically.
- Thread-safe snapshotting isolates hash-table locks from file system I/O to prevent contention.
- Upon detecting changes, the watcher executes `watcher_index_fn` to trigger a supervised re-indexing pipeline automatically.

## Frequently Asked Questions

### How often does the DeusData watcher check for file changes?

The background watcher polls every **5 seconds** by default, but it calculates an adaptive interval using `cbm_watcher_poll_interval_ms` that adds one second of delay for every 500 tracked files, up to a maximum of 60 seconds. This prevents excessive CPU and I/O usage on very large repositories.

### What happens when the watcher detects a modification in a Git repository?

When `git_is_dirty` reports uncommitted changes, the watcher immediately invokes the `watcher_index_fn` callback that was registered during initialization at `src/main.c:771`. This callback acquires a pipeline lock and executes `cbm_pipeline_run` to perform a full re-index of the modified project, then updates the stored HEAD hash and file count baseline.

### How does the watcher handle non-Git projects or deleted directories?

The watcher explicitly ignores non-Git projects by checking `is_git` at `src/watcher/watcher.c:618-620` and returning early. For directories that have been deleted since the last poll, the `root_status` check triggers `prune_missing_project`, which removes the stale project entry from the internal hash table to prevent unnecessary polling attempts.

### How does the background watcher shut down gracefully when the server stops?

The server signals shutdown by setting the atomic `g_shutdown` flag via `request_shutdown` at `src/main.c:72-80`, which then calls `cbm_watcher_stop`. This function atomically sets the `stopped` flag in the watcher structure, causing the `while (!atomic_load(&w->stopped))` loop in `cbm_watcher_run` to exit cleanly on its next iteration.