Retain vs Retain-Files in Hindsight: Understanding the Two Memory Ingestion Commands

The retain command ingests raw text directly via JSON payloads, while retain-files uploads binary files through multipart form data, converts them to markdown, and then processes the extracted text through the same underlying memory pipeline.

The vectorize-io/hindsight repository provides two high-level CLI verbs for populating memory banks with structured data. While both commands ultimately feed into identical fact-extraction, embedding, and deduplication pipelines, they differ fundamentally in transport protocol, payload structure, and preprocessing requirements. Choosing between them depends on whether your source data exists as plain text or as binary documents requiring conversion.

Core Differences Between retain and retain-files

The retain Command: Direct Text Ingestion

The retain command stores a single piece of text—or a batch of text items—as discrete memory units within a specified bank. According to the source code in hindsight-clients/go/api_memory.go, this command maps to the RetainMemories method, which sends a POST request to /v1/default/banks/{bank_id}/memories.

The payload consists of a JSON RetainRequest object (defined in hindsight-clients/go/model_retain_request.go) containing an array of MemoryItem structures. Each item includes a content field for the text body, along with optional metadata such as document_id and tags. The client sets Content-Type: application/json and transmits the structured data directly to the backend.

The retain-files Command: Binary File Processing

The retain-files command handles one or many binary files—including PDF, DOCX, TXT, images, and audio—by first converting them to markdown format before retention. As implemented in hindsight-clients/go/api_files.go, this command invokes the FileRetain method, which transmits a POST request to /v1/default/banks/{bank_id}/files/retain.

Unlike the JSON-based retain flow, this endpoint expects multipart/form-data encoding. The request combines file streams with a JSON FileRetainRequest describing optional parameters such as parser configuration and document_tags. Because file conversion can be computationally intensive, the retain-files endpoint always processes asynchronously by default, returning an operation ID that users can poll via the Operations API.

Architectural Flow and Pipeline Convergence

Despite differing entry points, both commands converge onto the identical backend processing pipeline.

retain Pipeline:

  1. The CLI constructs a RetainRequest containing raw text items.
  2. MemoryAPIService.RetainMemoriesExecute validates the JSON and forwards it to the backend.
  3. The system extracts facts, generates embeddings, deduplicates entries, and links entities.
  4. If async=true, the call returns immediately with an operation ID; otherwise, it blocks until completion.

retain-files Pipeline:

  1. The CLI reads files from the filesystem and packs them into a multipart body alongside a FileRetainRequest.
  2. FilesAPIService.FileRetainExecute streams the binary data to the server-side "file-convert-retain" pipeline.
  3. The backend first converts files to markdown using pluggable parsers, then injects the extracted text into the standard retain pipeline described above.
  4. The operation always returns an operation ID for asynchronous status polling.

Practical Implementation Examples

CLI Usage Patterns

Store a single text snippet immediately:

hindsight memory retain my-bank "Alice works at Google as a software engineer"

Ingest text with metadata and asynchronous processing:

hindsight memory retain my-bank "Q3 Planning Notes" --context "strategy" --tags "quarterly,planning" --async

Upload individual files or entire directories:


# Single file ingestion

hindsight memory retain-files my-bank quarterly-report.pdf

# Recursive directory upload with background processing

hindsight memory retain-files my-bank ./documents/ --async

Go Client Integration

Plain Text Retention:

import (
    "context"
    "fmt"
    hindsight "github.com/vectorize-io/hindsight/hindsight-clients/go"
)

func retainText() {
    cfg := hindsight.NewConfiguration()
    client := hindsight.NewAPIClient(cfg)

    // Construct RetainRequest per model_retain_request.go
    req := hindsight.RetainRequest{
        Items: []hindsight.MemoryItem{
            {
                Content: hindsight.PtrString("Alice works at Google."),
                Tags:    &[]string{"personnel", "engineering"},
            },
        },
        Async: hindsight.PtrBool(false),
    }

    // Execute via MemoryAPIService.RetainMemories (api_memory.go)
    resp, _, err := client.MemoryAPI.RetainMemories(context.Background(), "my-bank").
        RetainRequest(req).Execute()
    if err != nil {
        panic(err)
    }
    fmt.Println("Memory ID:", resp.GetId())
}

File-Based Retention:

import (
    "context"
    "fmt"
    "os"
    hindsight "github.com/vectorize-io/hindsight/hindsight-clients/go"
)

func retainFile() {
    cfg := hindsight.NewConfiguration()
    client := hindsight.NewAPIClient(cfg)

    // Open local file for streaming
    file, err := os.Open("contract.pdf")
    if err != nil {
        panic(err)
    }
    defer file.Close()

    // JSON configuration for the file processing pipeline
    requestBody := `{"parser":"default","document_tags":["legal","contracts"]}`

    // Execute via FilesAPIService.FileRetain (api_files.go)
    resp, _, err := client.FilesAPI.FileRetain(context.Background(), "my-bank").
        Files([]*os.File{file}).
        Request(requestBody).
        Execute()
    if err != nil {
        panic(err)
    }
    fmt.Println("Operation ID:", resp.GetOperationId())
}

Summary

  • retain accepts JSON payloads containing raw text via POST /v1/default/banks/{bank_id}/memories, implemented in hindsight-clients/go/api_memory.go as RetainMemories.
  • retain-files streams binary files via multipart/form-data to POST /v1/default/banks/{bank_id}/files/retain, implemented in hindsight-clients/go/api_files.go as FileRetain.
  • Both commands ultimately process data through the same fact-extraction, embedding, and deduplication pipeline, but retain-files adds a file-to-markdown conversion step.
  • The retain command supports synchronous or asynchronous execution, while retain-files always operates asynchronously due to conversion overhead.

Frequently Asked Questions

Can I use retain to store file contents if I read the files myself?

Yes. If you extract text from files manually using external libraries, you can pass that content as string values in a RetainRequest to the retain command. However, you lose the automatic format detection, markdown conversion, and parsing optimizations built into the retain-files pipeline, which handles binary conversions server-side according to the logic in hindsight-clients/go/api_files.go.

Does retain-files always process asynchronously?

Yes. Because file conversion—particularly for PDFs, images, and audio—requires significant computational resources, the FileRetain method always initiates an asynchronous job. The API returns an operation ID immediately, which you can poll through the Operations API to track completion status, unlike the retain command which supports blocking synchronous calls.

What file formats does retain-files support?

The command supports diverse binary formats including PDF, DOCX, TXT, images, and audio files. The backend employs pluggable parsers (configurable via the parser field in FileRetainRequest) to convert these to markdown before retention. For specific format limitations, consult the API documentation at hindsight-docs/versioned_docs/version-0.4/developer/api/retain.mdx.

Do both commands create identical memory objects in the bank?

Yes. After processing, memories created via both paths contain the same underlying structure: extracted facts, vector embeddings, entity links, and metadata tags. The only difference lies in the ingestion transport layer—JSON text versus multipart file streams—and the preprocessing conversion step unique to retain-files.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →