# Retain vs Retain-Files in Hindsight: Understanding the Two Memory Ingestion Commands

> Understand the difference between Hindsight's retain and retain_files commands. Learn how retain ingests JSON text while retain_files processes uploaded binary files through a markdown conversion.

- Repository: [vectorize-io/hindsight](https://github.com/vectorize-io/hindsight)
- Tags: deep-dive
- Published: 2026-03-13

---

**The `retain` command ingests raw text directly via JSON payloads, while `retain-files` uploads binary files through multipart form data, converts them to markdown, and then processes the extracted text through the same underlying memory pipeline.**

The `vectorize-io/hindsight` repository provides two high-level CLI verbs for populating memory banks with structured data. While both commands ultimately feed into identical fact-extraction, embedding, and deduplication pipelines, they differ fundamentally in transport protocol, payload structure, and preprocessing requirements. Choosing between them depends on whether your source data exists as plain text or as binary documents requiring conversion.

## Core Differences Between `retain` and `retain-files`

### The `retain` Command: Direct Text Ingestion

The `retain` command stores a single piece of text—or a batch of text items—as discrete memory units within a specified bank. According to the source code in [`hindsight-clients/go/api_memory.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/api_memory.go), this command maps to the `RetainMemories` method, which sends a `POST` request to `/v1/default/banks/{bank_id}/memories`.

The payload consists of a JSON `RetainRequest` object (defined in [`hindsight-clients/go/model_retain_request.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/model_retain_request.go)) containing an array of `MemoryItem` structures. Each item includes a `content` field for the text body, along with optional metadata such as `document_id` and `tags`. The client sets `Content-Type: application/json` and transmits the structured data directly to the backend.

### The `retain-files` Command: Binary File Processing

The `retain-files` command handles one or many binary files—including PDF, DOCX, TXT, images, and audio—by first converting them to markdown format before retention. As implemented in [`hindsight-clients/go/api_files.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/api_files.go), this command invokes the `FileRetain` method, which transmits a `POST` request to `/v1/default/banks/{bank_id}/files/retain`.

Unlike the JSON-based `retain` flow, this endpoint expects `multipart/form-data` encoding. The request combines file streams with a JSON `FileRetainRequest` describing optional parameters such as `parser` configuration and `document_tags`. Because file conversion can be computationally intensive, the `retain-files` endpoint always processes asynchronously by default, returning an operation ID that users can poll via the Operations API.

## Architectural Flow and Pipeline Convergence

Despite differing entry points, both commands converge onto the identical backend processing pipeline.

**`retain` Pipeline:**

1. The CLI constructs a `RetainRequest` containing raw text items.
2. `MemoryAPIService.RetainMemoriesExecute` validates the JSON and forwards it to the backend.
3. The system extracts facts, generates embeddings, deduplicates entries, and links entities.
4. If `async=true`, the call returns immediately with an operation ID; otherwise, it blocks until completion.

**`retain-files` Pipeline:**

1. The CLI reads files from the filesystem and packs them into a multipart body alongside a `FileRetainRequest`.
2. `FilesAPIService.FileRetainExecute` streams the binary data to the server-side "file-convert-retain" pipeline.
3. The backend first converts files to markdown using pluggable parsers, then injects the extracted text into the standard retain pipeline described above.
4. The operation always returns an operation ID for asynchronous status polling.

## Practical Implementation Examples

### CLI Usage Patterns

Store a single text snippet immediately:

```bash
hindsight memory retain my-bank "Alice works at Google as a software engineer"

```

Ingest text with metadata and asynchronous processing:

```bash
hindsight memory retain my-bank "Q3 Planning Notes" --context "strategy" --tags "quarterly,planning" --async

```

Upload individual files or entire directories:

```bash

# Single file ingestion

hindsight memory retain-files my-bank quarterly-report.pdf

# Recursive directory upload with background processing

hindsight memory retain-files my-bank ./documents/ --async

```

### Go Client Integration

**Plain Text Retention:**

```go
import (
    "context"
    "fmt"
    hindsight "github.com/vectorize-io/hindsight/hindsight-clients/go"
)

func retainText() {
    cfg := hindsight.NewConfiguration()
    client := hindsight.NewAPIClient(cfg)

    // Construct RetainRequest per model_retain_request.go
    req := hindsight.RetainRequest{
        Items: []hindsight.MemoryItem{
            {
                Content: hindsight.PtrString("Alice works at Google."),
                Tags:    &[]string{"personnel", "engineering"},
            },
        },
        Async: hindsight.PtrBool(false),
    }

    // Execute via MemoryAPIService.RetainMemories (api_memory.go)
    resp, _, err := client.MemoryAPI.RetainMemories(context.Background(), "my-bank").
        RetainRequest(req).Execute()
    if err != nil {
        panic(err)
    }
    fmt.Println("Memory ID:", resp.GetId())
}

```

**File-Based Retention:**

```go
import (
    "context"
    "fmt"
    "os"
    hindsight "github.com/vectorize-io/hindsight/hindsight-clients/go"
)

func retainFile() {
    cfg := hindsight.NewConfiguration()
    client := hindsight.NewAPIClient(cfg)

    // Open local file for streaming
    file, err := os.Open("contract.pdf")
    if err != nil {
        panic(err)
    }
    defer file.Close()

    // JSON configuration for the file processing pipeline
    requestBody := `{"parser":"default","document_tags":["legal","contracts"]}`

    // Execute via FilesAPIService.FileRetain (api_files.go)
    resp, _, err := client.FilesAPI.FileRetain(context.Background(), "my-bank").
        Files([]*os.File{file}).
        Request(requestBody).
        Execute()
    if err != nil {
        panic(err)
    }
    fmt.Println("Operation ID:", resp.GetOperationId())
}

```

## Summary

- **`retain`** accepts JSON payloads containing raw text via `POST /v1/default/banks/{bank_id}/memories`, implemented in [`hindsight-clients/go/api_memory.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/api_memory.go) as `RetainMemories`.
- **`retain-files`** streams binary files via `multipart/form-data` to `POST /v1/default/banks/{bank_id}/files/retain`, implemented in [`hindsight-clients/go/api_files.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/api_files.go) as `FileRetain`.
- Both commands ultimately process data through the same fact-extraction, embedding, and deduplication pipeline, but `retain-files` adds a file-to-markdown conversion step.
- The `retain` command supports synchronous or asynchronous execution, while `retain-files` always operates asynchronously due to conversion overhead.

## Frequently Asked Questions

### Can I use `retain` to store file contents if I read the files myself?

Yes. If you extract text from files manually using external libraries, you can pass that content as string values in a `RetainRequest` to the `retain` command. However, you lose the automatic format detection, markdown conversion, and parsing optimizations built into the `retain-files` pipeline, which handles binary conversions server-side according to the logic in [`hindsight-clients/go/api_files.go`](https://github.com/vectorize-io/hindsight/blob/main/hindsight-clients/go/api_files.go).

### Does `retain-files` always process asynchronously?

Yes. Because file conversion—particularly for PDFs, images, and audio—requires significant computational resources, the `FileRetain` method always initiates an asynchronous job. The API returns an operation ID immediately, which you can poll through the Operations API to track completion status, unlike the `retain` command which supports blocking synchronous calls.

### What file formats does `retain-files` support?

The command supports diverse binary formats including PDF, DOCX, TXT, images, and audio files. The backend employs pluggable parsers (configurable via the `parser` field in `FileRetainRequest`) to convert these to markdown before retention. For specific format limitations, consult the API documentation at `hindsight-docs/versioned_docs/version-0.4/developer/api/retain.mdx`.

### Do both commands create identical memory objects in the bank?

Yes. After processing, memories created via both paths contain the same underlying structure: extracted facts, vector embeddings, entity links, and metadata tags. The only difference lies in the ingestion transport layer—JSON text versus multipart file streams—and the preprocessing conversion step unique to `retain-files`.