How memU Handles Cross-References Between Memory Items (Like Symlinks)

memU implements cross-references between memory items as lightweight symlinks using deterministic six-character short IDs embedded in category summaries and resolved via JSON queries on the extra column.

Cross-references between memory items enable AI systems to build knowledge graphs without duplicating content across the database. The open-source NevaMind-AI/memU repository implements this capability through a file system-inspired symlink architecture that links textual summaries to concrete memory entries using compact, deterministic identifiers.

Rather than embedding full UUIDs or content directly into text fields, memU uses a lightweight indirection mechanism. Each memory item receives a short identifier derived from its full UUID, functioning analogously to a file system symlink: the reference text lives in the category summary, while the target data remains in its original MemoryItem row.

Short-ID Generation

When a memory item is created, memU generates a deterministic six-character identifier by stripping hyphens from the full UUID and truncating to the first six characters. This occurs in src/memu/app/memorize.py:

short_id = item_id.replace("-", "")[:6]        # src/memu/app/memorize.py#L81-L83

This compact ID minimizes storage overhead in textual summaries while maintaining uniqueness within typical usage patterns.

Embedding References in Memory Items

After updating a category summary, memU scans the text for [ref:xxx] patterns. Upon detecting a valid short ID, the system persists it into the underlying MemoryItem row's extra column under the ref_id key:

store.memory_item_repo.update_item(
    item_id=matched_item_id,
    extra={"ref_id": short_id},
)                                            # src/memu/app/memorize.py#L30-L36

This storage mechanism decouples the reference text from the target item's location, enabling the system to resolve links dynamically during retrieval.

Reference Syntax and Text Processing

References appear within summaries using the [ref:abc123] syntax. The memu.utils.references module provides specialized utilities for handling these markers throughout the codebase:

  • extract_references(text) – Extracts all [ref:…] markers and returns a list of short IDs (lines 20-49).
  • strip_references(text) – Removes all reference markers to produce clean display text (lines 52-74).
  • format_references_as_citations(text) – Converts markers into numbered citations for presentation (lines 77-92).

These functions are implemented in src/memu/utils/references.py and provide the parsing layer necessary for both storage and display operations.

Cross-Reference Resolution During Retrieval

The retrieval pipeline in src/memu/app/retrieve.py resolves these symbolic links during query execution. It first extracts all reference IDs from generated category summaries, then fetches the corresponding concrete items:

ref_ids.extend(extract_references(summary))
items_pool = store.memory_item_repo.list_items_by_ref_ids(ref_ids, where_filters)

See lines 629-637 in src/memu/app/retrieve.py. The retrieved items are then fed back to the LLM as context, allowing the model to "follow the link" and incorporate the referenced memory's full content.

Database Implementation by Backend

Each storage backend implements list_items_by_ref_ids to resolve short IDs against the JSON extra column where ref_id is stored.

SQLite uses the json_extract function to query the nested field:

json_extract(extra, '$.ref_id') IN (:ref_ids)

See the implementation in src/memu/database/sqlite/repositories/memory_item_repo.py lines 119-144.

PostgreSQL achieves equivalent functionality using the ->> operator on the JSONB column, while the in-memory repository filters the cached object pool directly. The MemoryItem data model defined in src/memu/database/models.py provides the schema foundation for this JSON storage.

Practical Implementation Examples

Creating a Memory Item with a Reference


# Create a memory item (full UUID returned)

item = store.memory_item_repo.create_item(
    resource_id="doc-42",
    memory_type="note",
    summary="User loves coffee",
    embedding=[0.1, 0.2, ...],
    user_data={},
)

# In a category summary we can now write a reference:

summary = "User preferences include coffee [ref:" + memu.memorize._build_item_ref_id(item.id) + "]."

Extracting and Cleaning References

from memu.utils.references import extract_references, strip_references

refs = extract_references(summary)              # → ['abc123']

clean = strip_references(summary)               # "User preferences include coffee."

# Persist short-ID into the item's extra column (handled automatically by

# `Memorize._persist_item_references`).

Retrieving Referenced Items in a Query


# Inside the retrieve pipeline

ref_ids = extract_references(category_summary)  # get all short IDs used

items = store.memory_item_repo.list_items_by_ref_ids(ref_ids)

# `items` now contains the full MemoryItem objects whose extra.ref_id matches.

Summary

  • memU generates six-character short IDs from UUIDs to serve as lightweight handles for cross-references.
  • The [ref:xxx] syntax embedded in category summaries creates textual links without storing full content.
  • The extra.ref_id JSON field in MemoryItem rows stores the target identifier, queryable via backend-specific JSON operators.
  • SQLite uses json_extract(extra, '$.ref_id') while PostgreSQL uses the ->> operator for resolution.
  • The retrieval pipeline dynamically resolves these references, mimicking file system symlinks by separating link text from target data.

Frequently Asked Questions

How does memU generate the short IDs for cross-references?

memU generates deterministic six-character short IDs by removing hyphens from the full UUID and taking the first six characters. This logic resides in src/memu/app/memorize.py and ensures consistent, reproducible identifiers for each memory item.

Can I manually create cross-references between existing memory items?

Yes. You can manually insert [ref:shortid] markers into any category summary. When the system processes these summaries via the memorization pipeline, it automatically extracts the references and updates the corresponding MemoryItem rows with the appropriate ref_id value in the extra column.

What database backends support the cross-reference lookup?

All official backends implement list_items_by_ref_ids. SQLite queries using json_extract(extra, '$.ref_id'), PostgreSQL uses the JSONB ->> operator, and the in-memory repository filters cached objects. See the SQLite implementation in src/memu/database/sqlite/repositories/memory_item_repo.py lines 119-144.

Are cross-references resolved during the retrieval phase?

Yes. During retrieval (src/memu/app/retrieve.py lines 629-637), the system extracts all [ref:…] markers from category summaries, queries the repository for items matching those short IDs, and includes the full target items in the context window provided to the LLM.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →