Performance Tuning for MemPalace Hooks Under 500ms Latency

Configure SAVE_INTERVAL=50, pin MEMPAL_PYTHON to your venv binary, and switch to the all-MiniLM-L6-v2 embedder to eliminate UI freezing and keep hook execution under 500ms.

MemPalace is an MCP-compatible memory system that persists conversation context through Bash-based automation. When performance tuning for MemPalace hooks under 500ms latency becomes critical for responsive Claude Code sessions, you must optimize both the asynchronous save triggers and the synchronous pre-compact mining operations defined in hooks/mempal_save_hook.sh and hooks/mempal_precompact_hook.sh.

Understanding the MemPalace Hook Architecture

The MemPalace/mempalace repository provides two Bash-based hooks invoked by MCP clients during chat sessions. Both share common parsing logic in mempalace/hooks_cli.py and validation routines, but differ in execution timing and blocking behavior.

The Save Hook (mempal_save_hook.sh)

Located at hooks/mempal_save_hook.sh, this hook executes on the Stop signal after every assistant response. Lines 162-227 handle message counting and state tracking, blocking the stop signal every SAVE_INTERVAL exchanges to trigger an asynchronous mempalace mine … --mode convos in the background. The install and configuration logic spans lines 34-84, while Python-based parsing occupies lines 81-130.

The Pre-Compact Hook (mempal_precompact_hook.sh)

Located at hooks/mempal_precompact_hook.sh, this hook runs synchronously just before conversation compression. Lines 162-200 execute a blocking mempalace mine … --mode convos to ensure no data is lost before context window reduction. The configuration section resides in lines 58-94, with parsing logic in lines 101-130.

Identifying Latency Bottlenecks

High latency in MemPalace hooks stems from four primary sources:

  1. Python interpreter resolution: The hooks check $MEMPAL_PYTHON then fall back to command -v python3. If the resolved interpreter lacks the mempalace entry-point, the subprocess spawns and retries, adding seconds of overhead.

  2. Synchronous mining: The pre-compact hook runs mempalace mine in the foreground (lines 162-200 of mempal_precompact_hook.sh). Large transcript directories cause multi-second pauses where the Claude Code UI appears frozen.

  3. Background competition: While the save hook uses asynchronous execution (&), the background process still competes for CPU and disk I/O, potentially delaying the next assistant response.

  4. Log contention: Both hooks append to ~/.mempalace/hook_state/hook.log. Concurrent sessions create I/O contention on busy systems.

Environment Variables for Sub-500ms Performance

Tune these environment variables to reduce hook latency:

  • SAVE_INTERVAL: Increase from the default 15 to 30 or 50 to reduce the frequency of background mine spawns.

  • MEMPAL_PYTHON: Export the absolute path to a Python binary that already has mempalace installed (e.g., /home/user/.venvs/mempalace/bin/python3) to bypass interpreter discovery.

  • MEMPALACA_HOOKS_AUTO_SAVE: Set to false (or 0, no) to disable both hooks entirely, eliminating all automated latency if you prefer manual mining.

  • MEMPAL_DIR: Clear this variable (set to empty) to skip the optional project-directory mining step and remove an extra mempalace mine invocation.

  • MEMPAL_VERBOSE: Keep as false (default) to return {} silently rather than blocking the UI with a JSON decision payload.

  • STATE_DIR: Redirect to an in-memory tmpfs (e.g., /dev/shm/mempalace_hook) to speed up log writes and eliminate disk I/O bottlenecks.

Pinning the Interpreter and Raising Save Intervals

Add these exports to your shell profile (~/.bashrc or ~/.zshrc) to eliminate interpreter discovery overhead and reduce background process frequency:

export SAVE_INTERVAL=40
export MEMPAL_PYTHON=$HOME/.venvs/mempalace/bin/python3
export MEMPALACA_HOOKS_AUTO_SAVE=true
export MEMPAL_DIR=""  # Disable project-wide mining

export STATE_DIR=/dev/shm/mempalace_hook

Optimizing Synchronous Pre-Compact Mining

The pre-compact hook deliberately blocks to guarantee data persistence. Reduce its impact with these strategies:

Switch to a lightweight embedder: The default embeddinggemma-300m model consumes ~300MB and heavy CPU. Download and use all-MiniLM-L6-v2 (~30MB) instead:

mempalace embedder download all-MiniLM-L6-v2

Restrict mining scope: Ensure your transcript directory ($(dirname "$TRANSCRIPT_PATH")) contains only current session files. Shallow folders with fewer large files complete mining faster.

Cache the embedder: Keep the embedding process alive via uv run mempalace serve so subsequent pre-compact runs reuse the cached model rather than cold-loading.

Eliminating I/O Contention

For high-throughput environments running many concurrent Claude Code sessions, move hook state to memory and manage log growth:

Use tmpfs for state storage:

export STATE_DIR=/dev/shm/mempalace_hook

This places hook.log on a RAM disk, eliminating disk I/O bottlenecks during rapid append operations.

Implement log rotation: While the repository truncates error dumps to 4KB in hooks/mempal_save_hook.sh, you can extend this pattern to hook.log by setting a maximum size threshold:

export MEMPAL_HOOK_LOG_MAX_SIZE=1048576  # 1 MiB

Modify the hook or maintain a custom copy in your path to enforce this limit and prevent log bloat.

Summary

  • Increase SAVE_INTERVAL from 15 to 30-50 to reduce background mining frequency.
  • Pin MEMPAL_PYTHON to your venv Python path to avoid interpreter discovery delays.
  • Clear MEMPAL_DIR to disable project-wide mining and eliminate extra mempalace mine invocations.
  • Switch embedders from embeddinggemma-300m to all-MiniLM-L6-v2 for faster synchronous pre-compact operations.
  • Move STATE_DIR to a tmpfs mount like /dev/shm/mempalace_hook to eliminate log write bottlenecks.
  • Disable auto-save entirely with MEMPALACA_HOOKS_AUTO_SAVE=false if preferring manual mining in CI or batch environments.

Frequently Asked Questions

Why does the PreCompact hook cause UI freezing?

The pre-compact hook runs mempalace mine … --mode convos synchronously in the foreground (lines 162-200 of hooks/mempal_precompact_hook.sh). This blocking behavior ensures conversation data persists before context compression, but large transcript directories or slow embedder loading cause visible pauses in the Claude Code interface.

How do I completely disable automatic mining?

Set export MEMPALACA_HOOKS_AUTO_SAVE=false (or 0, no) before starting your Claude Code session. This disables both the save hook and pre-compact hook entirely, allowing you to manually trigger mempalace mine when convenient.

Can I use a custom Python interpreter without editing the hook files?

Yes. Export MEMPAL_PYTHON=/absolute/path/to/your/python3 before launching Claude Code. The hook checks this variable before running command -v python3, ensuring the correct interpreter with the mempalace entry-point is used without modifying hooks/mempal_save_hook.sh.

Use all-MiniLM-L6-v2 instead of the default embeddinggemma-300m. At approximately 30MB versus 300MB, the smaller model loads faster and consumes less CPU during the synchronous pre-compact mining phase, keeping hook execution under 500ms.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →