Input Validation and Sanitization in MemPalace config.py: A Complete Guide

MemPalace's config.py module centralizes all input validation and sanitization into strict, reusable functions that strip malformed Unicode, enforce length limits, and validate ISO-8601 dates before any data reaches the storage layer.

MemPalace relies on a single source-of-truth configuration module to validate every piece of user-supplied data before it is stored or processed. The mempalace/config.py file implements a small, well-contained API that normalizes names, dates, and content strings, protects the system from malformed Unicode, and sanitizes knowledge-graph values. Understanding these validation boundaries is essential for anyone extending the MemPalace ingestion pipeline or troubleshooting data errors.

Core Sanitization Utilities in mempalace/config.py

The validation layer in MemPalace is built around explicit sanitization functions rather than hidden side effects. Each utility raises a clear ValueError when a rule is violated, allowing upstream callers to handle bad input before it ever touches the database.

sanitize_name: Normalizing Identifiers

The sanitize_name(value: str, field_name: str = "name") → str function, implemented in mempalace/config.py, cleans wing, room, entity, and other identifier strings. It strips surrounding whitespace, removes control characters and invalid surrogate pairs, collapses multiple spaces into a single space, and enforces a maximum length of 256 characters. If the resulting string is empty, the function raises a ValueError with a field-specific message, preventing downstream "not-found" errors caused by invisible or malformed tokens.

sanitize_kg_value: Cleaning Knowledge-Graph Values

Graph predicates and values in the SQLite backend must remain comparable and searchable. The sanitize_kg_value(value: str, field_name: str = "value") → str function calls sanitize_name internally to reuse the same Unicode and whitespace rules, then additionally strips trailing periods to keep graph keys tidy. This guarantees that knowledge-graph lookups remain deterministic and free of hidden characters.

sanitize_iso_temporal and sanitize_iso_date: Strict Date Validation

Temporal data entering MemPalace must conform to ISO-8601 standards. The sanitize_iso_temporal(value, field_name: str = "date") → str function verifies that the supplied string matches an ISO-8601 regex pattern—accepting dates, times, datetimes, or intervals—and then calls the private _validate_iso_temporal_calendar helper to ensure calendar-level validity. For stricter pure-date requirements, sanitize_iso_date(value, field_name: str = "date") → str reuses the same regex but forces the pattern to contain only the date component, raising a ValueError if a time component is present. This distinction prevents datetime strings from leaking into fields that expect only a calendar date.

sanitize_content: Protecting Verbatim Text Stores

MemPalace stores verbatim transcripts and notes inside drawers, so it must clean text without altering its meaning. The sanitize_content(value: str, max_length: int = 100_000) → str function normalizes line endings to \n, strips leading and trailing whitespace from each line, removes isolated surrogate code points, and enforces a hard length limit. Transcripts that exceed the default 100 KB threshold trigger a clear ValueError, protecting the ingestion pipeline from out-of-memory crashes during large mining operations.

Configuration Helpers and Safe Type Coercion

Beyond string cleaning, mempalace/config.py provides helper utilities such as _try_coerce_int and _validated_chunk_config that safely convert user-provided numeric settings. These helpers convert strings to integers with optional minimum-value checks, then fall back to documented defaults when conversion fails. Because the Config class exposes chunk size, backend selection, and embedding model settings through these guarded getters, a typo in a user-provided config.json never silently disables core behaviors.

Practical Examples

The following code demonstrates how to use the public sanitization API to clean user-supplied data before passing it to the MemPalace storage layer.

from mempalace.config import (
    sanitize_name,
    sanitize_iso_date,
    sanitize_content,
)

# Example 1 – cleaning a wing name supplied by a user

raw_wing = "  John Doe’s 🧠  "
clean_wing = sanitize_name(raw_wing, field_name="wing")

# clean_wing == "John Doe’s 🧠"

# Example 2 – enforcing a proper ISO‑date

raw_date = "2023-02-30"          # Invalid – February never has 30 days

try:
    good_date = sanitize_iso_date(raw_date)
except ValueError as e:
    print(f"Bad date: {e}")

# Example 3 – trimming a huge transcript before mining

raw_transcript = open("large.txt").read()
try:
    safe_transcript = sanitize_content(raw_transcript)
except ValueError as e:
    # Likely the transcript exceeds 100 KB

    raise RuntimeError("Transcript too large") from e

In mempalace/config.py, the sanitize_name implementation appears near line 49, sanitize_iso_temporal near line 138, sanitize_iso_date near line 175, and sanitize_content near line 186. The Config class ties these helpers together to load and validate all runtime settings.

Why This Matters for MemPalace Data Integrity

Strict validation inside mempalace/config.py underpins three critical system guarantees.

  • Verbatim-first guarantee. MemPalace never stores paraphrased or lossy data. By sanitizing only characters that could corrupt storage—such as lone surrogates or mixed line endings—the system preserves the exact user words while protecting the database from malformed Unicode.
  • Robust knowledge-graph indexing. Clean identifiers and values let the SQLite knowledge-graph enforce reliable primary-key constraints and fast lookups without hidden whitespace breaking equality checks.
  • Predictable ingestion performance. Hard length limits and strict ISO-date handling stop accidental huge payloads from slowing down mining or triggering memory exhaustion.

Because every sanitization function raises ValueError with a descriptive message, callers such as the CLI or MCP server can surface user-friendly errors immediately and halt processing before bad data is committed.

Summary

  • mempalace/config.py is the single source of truth for input validation and sanitization in MemPalace.
  • sanitize_name strips whitespace, control characters, and surrogates from identifiers while capping length at 256 characters.
  • sanitize_kg_value layers knowledge-graph-specific cleaning on top of sanitize_name to keep predicates deterministic.
  • sanitize_iso_temporal and sanitize_iso_date enforce strict ISO-8601 compliance, with the latter rejecting any time component.
  • sanitize_content normalizes line endings and enforces a 100 KB default limit to protect verbatim stores.
  • Helper utilities like _try_coerce_int and _validated_chunk_config guard numeric configuration values against typos and invalid types.

Frequently Asked Questions

What happens when sanitize_name receives an empty string?

The function raises a ValueError with a field-specific message. This early rejection prevents empty identifiers from propagating into the knowledge-graph or storage layer, where they would cause ambiguous lookup failures.

Can sanitize_iso_date accept a datetime string such as 2023-01-01T00:00:00?

No. sanitize_iso_date strictly enforces the pure date component (YYYY-MM-DD). If a time component is present, the function raises a ValueError, ensuring that date-only fields remain free of unexpected timestamps.

Why does MemPalace enforce a 100 KB limit in sanitize_content instead of allowing unlimited transcripts?

The default max_length of 100,000 characters in sanitize_content acts as a circuit breaker during ingestion. Payloads that exceed this limit raise a ValueError, which prevents out-of-memory crashes and keeps mining operations predictable without sacrificing the verbatim integrity of appropriately sized inputs.

Where are configuration helpers like _try_coerce_int used?

These helpers power the Config class inside mempalace/config.py, converting raw JSON settings—such as chunk size or backend flags—into validated integers with safe fallbacks. This ensures that user typos in config.json never silently disable chunking or other core pipeline behaviors.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →