# How Cross-Session Long-Term Memory Is Auto-Extracted and Stored in WeKnora

> Discover how WeKnora automatically extracts and stores cross-session long-term memory using LLMs and a background queue worker for efficient retrieval.

- Repository: [Tencent/WeKnora](https://github.com/tencent/WeKnora)
- Tags: deep-dive
- Published: 2026-09-12

---

**WeKnora automatically extracts and persists cross-session long-term memory through a background queue worker that distills conversation turns using a configurable LLM and stores the results as semantically-indexed memory items for future retrieval.**

WeKnora is an open-source knowledge-aware conversational AI framework that maintains persistent user memories across independent chat sessions. The system implements **cross-session long-term memory auto-extraction** through a sophisticated pipeline that operates asynchronously without requiring explicit user commands. This article examines the exact implementation details found in the Tencent/WeKnora repository, tracing the journey from conversation turn to persisted memory item.

## The Three-Component Architecture

The auto-extraction system comprises three tightly integrated layers that handle configuration, scheduling, and execution.

### Memory Configuration Layer

Workspace-level settings determine whether auto-extraction is active and which model performs the distillation. In [`internal/types/memory.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory.go), the `MemoryConfig` struct defines the extraction parameters:

- `ExtractModelID`: The LLM identifier used for generating memory summaries
- `ExtractDelaySeconds`: Debounce timer preventing immediate extraction after every turn
- `ExtractMinIntervalSeconds`: Minimum cooldown between extraction runs for the same subject
- `ExtractInstructions`: Custom system prompts guiding the extraction style

The `(*MemoryConfig) AutoExtractEnabled()` method returns **true** only when the configuration exists and `ExtractModelID` is non-empty, serving as the primary gatekeeper for the entire pipeline.

### Background Distillation Task

The system decouples extraction from the request-response cycle using a task queue defined in [`internal/types/task.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/task.go). The **QueueMemory** worker processes jobs of type `TypeMemoryExtract = "memory:extract"`. The `MemoryExtractPayload` carries the essential context:

```go
type MemoryExtractPayload struct {
    SubjectID uint64
    TenantID  uint64
    TurnIDs   []uint64
}

```

This asynchronous design ensures that memory extraction never blocks user interactions, allowing the system to batch multiple turns or retry failed extractions without impacting chat latency.

### Extraction Service

The core logic resides in [`internal/application/service/memory/extract.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/memory/extract.go). This service orchestrates the actual LLM interaction and database persistence. It validates the configuration via `cfg.AutoExtractEnabled()`, constructs prompts containing recent conversation history, and manages concurrent access through lease-based locking.

## The Auto-Extraction Workflow

When a conversation turn completes, the system initiates a six-stage pipeline to transform transient chat content into persistent structured memory.

### Turn Completion and Eligibility

Upon finishing a QA turn, the application checks the tenant's `MemoryConfig`. Only workspaces with valid configurations trigger the extraction pathway. The system records the turn content and marks the session as eligible for background processing.

### Debounce and Scheduling

To prevent LLM API flooding, WeKnora implements intelligent throttling. The scheduler respects `ExtractDelaySeconds` (how long to wait after activity ceases) and `ExtractMinIntervalSeconds` (global rate limiting per subject). When these timers expire, the server enqueues a `memory:extract` task via the QueueMemory interface.

### Lease Acquisition and Concurrency Control

The extraction worker must obtain a `MemoryExtractionSession` lease before processing, as defined in [`internal/types/memory_extraction.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory_extraction.go). This mechanism guarantees that only one worker handles a given subject at a time. If the lease expires or is stolen by another instance, the worker aborts with `ErrMemoryExtractionLeaseLost`, preventing duplicate memory creation and ensuring exactly-once semantics.

### Prompt Construction and Model Invocation

The service builds structured prompts that include recent conversation turns identified in `MemoryExtractPayload.TurnIDs`, custom instructions from `cfg.ExtractInstructions`, and the `<user_memory>` envelope tag signaling the model to produce extractable notes. The configured model (`cfg.ExtractModelID`) processes this input and returns candidate memory strings—typically concise factual statements or user preference summaries.

### Persistence and State Update

Each valid candidate becomes a `MemoryItem` with `Origin` set to `MemoryOriginExtracted`, linking back to the source turn via `SourceTurnID`. The system generates vector embeddings through `MemoryItemEmbedding` for semantic retrieval in future sessions. Finally, the worker updates the subject's `MemoryExtractionState`, advancing the `LastExtractedAt` timestamp and `ExtractCursor` to mark successful completion.

## Configuration and Data Models

Understanding the underlying data structures clarifies how WeKnora maintains extraction state across distributed workers.

### The MemoryConfig Gate

The `AutoExtractEnabled()` method in [`internal/types/memory.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory.go) provides a concise validation pattern:

```go
func (c *MemoryConfig) AutoExtractEnabled() bool {
    return c != nil && c.ExtractModelID != ""
}

```

This boolean check appears in both the scheduling logic and the extraction worker, ensuring consistent behavior across the pipeline.

### Extraction State Tracking

The `MemoryExtractionState` struct (paired with `MemoryExtractionSession` in [`internal/types/memory_extraction.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory_extraction.go)) tracks `LastExtractedAt` (timestamp of the most recent successful extraction), `ExtractCursor` (progress marker preventing reprocessing of old turns), and lease metadata for distributed coordination. These fields enable the system to resume interrupted extractions and respect the `ExtractMinIntervalSeconds` constraint across server restarts.

## Implementation Example

The following pattern demonstrates the complete extraction flow as implemented in the service layer:

```go
// Scheduling phase after turn completion
if cfg.AutoExtractEnabled() {
    queue.Enqueue(task.Task{
        Type: task.TypeMemoryExtract,
        Payload: task.MemoryExtractPayload{
            SubjectID: subject.ID,
            TenantID:  tenant.ID,
            TurnIDs:   []uint64{turn.ID},
        },
    })
}

// Worker implementation in internal/application/service/memory/extract.go
func (s *Service) extractMemory(ctx context.Context, p task.MemoryExtractPayload) error {
    cfg := p.Config
    if !cfg.AutoExtractEnabled() {
        return nil
    }
    
    // Build prompt with conversation context
    prompt := buildExtractionPrompt(p.TurnIDs, cfg.ExtractInstructions)
    
    // Invoke extraction model
    results, err := s.llmClient.Generate(ctx, cfg.ExtractModelID, prompt)
    if err != nil {
        return err
    }
    
    // Persist extracted memories
    for _, content := range results {
        item := types.MemoryItem{
            SubjectID:    p.SubjectID,
            Origin:       types.MemoryOriginExtracted,
            Content:      content,
            SourceTurnID: p.TurnIDs[0],
        }
        if err := s.repo.CreateMemoryItem(ctx, &item); err != nil {
            return err
        }
    }
    
    // Advance extraction cursor
    return s.repo.UpdateExtractionState(ctx, p.SubjectID, time.Now())
}

```

The `buildExtractionPrompt` helper wraps user content with the `<user_memory>` envelope, ensuring the LLM returns structured data suitable for long-term storage.

## Summary

WeKnora implements **cross-session long-term memory auto-extraction** through a resilient, queue-based architecture:

- **Configuration-driven activation** via `MemoryConfig.AutoExtractEnabled()` in [`internal/types/memory.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory.go) determines which workspaces utilize automatic extraction
- **Asynchronous processing** through `QueueMemory` tasks (`TypeMemoryExtract`) decouples memory creation from chat responses
- **Distributed-safe execution** using `MemoryExtractionSession` leases prevents concurrent processing conflicts
- **Semantic persistence** stores distilled memories as `MemoryItem` records with `MemoryOriginExtracted` origin and vector embeddings for cross-session retrieval

## Frequently Asked Questions

### What triggers memory extraction in WeKnora?

Memory extraction triggers automatically after conversation turn completion when the tenant's `MemoryConfig` has a valid `ExtractModelID`. The system enqueues a background task rather than processing immediately, respecting debounce timers (`ExtractDelaySeconds` and `ExtractMinIntervalSeconds`) to prevent excessive LLM calls.

### How does WeKnora prevent duplicate memory extraction?

The system uses lease-based concurrency control through `MemoryExtractionSession` defined in [`internal/types/memory_extraction.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory_extraction.go). Before processing, workers must acquire a lease for the specific subject; if another worker holds the lease, the task aborts with `ErrMemoryExtractionLeaseLost`. Additionally, the `ExtractCursor` in `MemoryExtractionState` ensures previously processed turns are not re-extracted.

### Which LLM model performs the extraction?

The extraction model is configurable per workspace via the `ExtractModelID` field in `MemoryConfig`. The system invokes this model in [`internal/application/service/memory/extract.go`](https://github.com/Tencent/WeKnora/blob/main/internal/application/service/memory/extract.go), passing conversation context wrapped in the `<user_memory>` envelope along with any custom `ExtractInstructions` defined by the tenant administrator.

### Where are extracted memories stored for cross-session retrieval?

Extracted memories persist as `MemoryItem` records in the database, defined in [`internal/types/memory.go`](https://github.com/Tencent/WeKnora/blob/main/internal/types/memory.go). Each item stores its content, origin (`MemoryOriginExtracted`), source turn reference, and vector embeddings (`MemoryItemEmbedding`). These records associate with a specific `SubjectID`, making them available across all future chat sessions for that user or tenant.