How Cross-Session Long-Term Memory Is Auto-Extracted and Stored in WeKnora

WeKnora automatically extracts and persists cross-session long-term memory through a background queue worker that distills conversation turns using a configurable LLM and stores the results as semantically-indexed memory items for future retrieval.

WeKnora is an open-source knowledge-aware conversational AI framework that maintains persistent user memories across independent chat sessions. The system implements cross-session long-term memory auto-extraction through a sophisticated pipeline that operates asynchronously without requiring explicit user commands. This article examines the exact implementation details found in the Tencent/WeKnora repository, tracing the journey from conversation turn to persisted memory item.

The Three-Component Architecture

The auto-extraction system comprises three tightly integrated layers that handle configuration, scheduling, and execution.

Memory Configuration Layer

Workspace-level settings determine whether auto-extraction is active and which model performs the distillation. In internal/types/memory.go, the MemoryConfig struct defines the extraction parameters:

  • ExtractModelID: The LLM identifier used for generating memory summaries
  • ExtractDelaySeconds: Debounce timer preventing immediate extraction after every turn
  • ExtractMinIntervalSeconds: Minimum cooldown between extraction runs for the same subject
  • ExtractInstructions: Custom system prompts guiding the extraction style

The (*MemoryConfig) AutoExtractEnabled() method returns true only when the configuration exists and ExtractModelID is non-empty, serving as the primary gatekeeper for the entire pipeline.

Background Distillation Task

The system decouples extraction from the request-response cycle using a task queue defined in internal/types/task.go. The QueueMemory worker processes jobs of type TypeMemoryExtract = "memory:extract". The MemoryExtractPayload carries the essential context:

type MemoryExtractPayload struct {
    SubjectID uint64
    TenantID  uint64
    TurnIDs   []uint64
}

This asynchronous design ensures that memory extraction never blocks user interactions, allowing the system to batch multiple turns or retry failed extractions without impacting chat latency.

Extraction Service

The core logic resides in internal/application/service/memory/extract.go. This service orchestrates the actual LLM interaction and database persistence. It validates the configuration via cfg.AutoExtractEnabled(), constructs prompts containing recent conversation history, and manages concurrent access through lease-based locking.

The Auto-Extraction Workflow

When a conversation turn completes, the system initiates a six-stage pipeline to transform transient chat content into persistent structured memory.

Turn Completion and Eligibility

Upon finishing a QA turn, the application checks the tenant's MemoryConfig. Only workspaces with valid configurations trigger the extraction pathway. The system records the turn content and marks the session as eligible for background processing.

Debounce and Scheduling

To prevent LLM API flooding, WeKnora implements intelligent throttling. The scheduler respects ExtractDelaySeconds (how long to wait after activity ceases) and ExtractMinIntervalSeconds (global rate limiting per subject). When these timers expire, the server enqueues a memory:extract task via the QueueMemory interface.

Lease Acquisition and Concurrency Control

The extraction worker must obtain a MemoryExtractionSession lease before processing, as defined in internal/types/memory_extraction.go. This mechanism guarantees that only one worker handles a given subject at a time. If the lease expires or is stolen by another instance, the worker aborts with ErrMemoryExtractionLeaseLost, preventing duplicate memory creation and ensuring exactly-once semantics.

Prompt Construction and Model Invocation

The service builds structured prompts that include recent conversation turns identified in MemoryExtractPayload.TurnIDs, custom instructions from cfg.ExtractInstructions, and the <user_memory> envelope tag signaling the model to produce extractable notes. The configured model (cfg.ExtractModelID) processes this input and returns candidate memory strings—typically concise factual statements or user preference summaries.

Persistence and State Update

Each valid candidate becomes a MemoryItem with Origin set to MemoryOriginExtracted, linking back to the source turn via SourceTurnID. The system generates vector embeddings through MemoryItemEmbedding for semantic retrieval in future sessions. Finally, the worker updates the subject's MemoryExtractionState, advancing the LastExtractedAt timestamp and ExtractCursor to mark successful completion.

Configuration and Data Models

Understanding the underlying data structures clarifies how WeKnora maintains extraction state across distributed workers.

The MemoryConfig Gate

The AutoExtractEnabled() method in internal/types/memory.go provides a concise validation pattern:

func (c *MemoryConfig) AutoExtractEnabled() bool {
    return c != nil && c.ExtractModelID != ""
}

This boolean check appears in both the scheduling logic and the extraction worker, ensuring consistent behavior across the pipeline.

Extraction State Tracking

The MemoryExtractionState struct (paired with MemoryExtractionSession in internal/types/memory_extraction.go) tracks LastExtractedAt (timestamp of the most recent successful extraction), ExtractCursor (progress marker preventing reprocessing of old turns), and lease metadata for distributed coordination. These fields enable the system to resume interrupted extractions and respect the ExtractMinIntervalSeconds constraint across server restarts.

Implementation Example

The following pattern demonstrates the complete extraction flow as implemented in the service layer:

// Scheduling phase after turn completion
if cfg.AutoExtractEnabled() {
    queue.Enqueue(task.Task{
        Type: task.TypeMemoryExtract,
        Payload: task.MemoryExtractPayload{
            SubjectID: subject.ID,
            TenantID:  tenant.ID,
            TurnIDs:   []uint64{turn.ID},
        },
    })
}

// Worker implementation in internal/application/service/memory/extract.go
func (s *Service) extractMemory(ctx context.Context, p task.MemoryExtractPayload) error {
    cfg := p.Config
    if !cfg.AutoExtractEnabled() {
        return nil
    }
    
    // Build prompt with conversation context
    prompt := buildExtractionPrompt(p.TurnIDs, cfg.ExtractInstructions)
    
    // Invoke extraction model
    results, err := s.llmClient.Generate(ctx, cfg.ExtractModelID, prompt)
    if err != nil {
        return err
    }
    
    // Persist extracted memories
    for _, content := range results {
        item := types.MemoryItem{
            SubjectID:    p.SubjectID,
            Origin:       types.MemoryOriginExtracted,
            Content:      content,
            SourceTurnID: p.TurnIDs[0],
        }
        if err := s.repo.CreateMemoryItem(ctx, &item); err != nil {
            return err
        }
    }
    
    // Advance extraction cursor
    return s.repo.UpdateExtractionState(ctx, p.SubjectID, time.Now())
}

The buildExtractionPrompt helper wraps user content with the <user_memory> envelope, ensuring the LLM returns structured data suitable for long-term storage.

Summary

WeKnora implements cross-session long-term memory auto-extraction through a resilient, queue-based architecture:

  • Configuration-driven activation via MemoryConfig.AutoExtractEnabled() in internal/types/memory.go determines which workspaces utilize automatic extraction
  • Asynchronous processing through QueueMemory tasks (TypeMemoryExtract) decouples memory creation from chat responses
  • Distributed-safe execution using MemoryExtractionSession leases prevents concurrent processing conflicts
  • Semantic persistence stores distilled memories as MemoryItem records with MemoryOriginExtracted origin and vector embeddings for cross-session retrieval

Frequently Asked Questions

What triggers memory extraction in WeKnora?

Memory extraction triggers automatically after conversation turn completion when the tenant's MemoryConfig has a valid ExtractModelID. The system enqueues a background task rather than processing immediately, respecting debounce timers (ExtractDelaySeconds and ExtractMinIntervalSeconds) to prevent excessive LLM calls.

How does WeKnora prevent duplicate memory extraction?

The system uses lease-based concurrency control through MemoryExtractionSession defined in internal/types/memory_extraction.go. Before processing, workers must acquire a lease for the specific subject; if another worker holds the lease, the task aborts with ErrMemoryExtractionLeaseLost. Additionally, the ExtractCursor in MemoryExtractionState ensures previously processed turns are not re-extracted.

Which LLM model performs the extraction?

The extraction model is configurable per workspace via the ExtractModelID field in MemoryConfig. The system invokes this model in internal/application/service/memory/extract.go, passing conversation context wrapped in the <user_memory> envelope along with any custom ExtractInstructions defined by the tenant administrator.

Where are extracted memories stored for cross-session retrieval?

Extracted memories persist as MemoryItem records in the database, defined in internal/types/memory.go. Each item stores its content, origin (MemoryOriginExtracted), source turn reference, and vector embeddings (MemoryItemEmbedding). These records associate with a specific SubjectID, making them available across all future chat sessions for that user or tenant.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →