# How Document Chunking and Embedding Processing Works During Upload in AnythingLLM

> Discover how AnythingLLM handles document chunking and embedding processing during upload. Learn about file conversion, text splitting, vector generation, and caching in this technical deep dive.

- Repository: [Mintplex Labs/anything-llm](https://github.com/Mintplex-Labs/anything-llm)
- Tags: deep-dive
- Published: 2026-03-07

---

**AnythingLLM processes uploads through a multi-stage pipeline that converts files to JSON via the Collector API, splits text using a model-aware TextSplitter, generates embeddings through the configured engine, and caches vectors in `vector-cache/` to eliminate redundant processing.**

When you upload a document to AnythingLLM, the system orchestrates a sophisticated workflow that bridges file ingestion, intelligent text segmentation, and vector generation. This article examines the complete flow—from the initial API request in [`server/endpoints/api/document/index.js`](https://github.com/Mintplex-Labs/anything-llm/blob/main/server/endpoints/api/document/index.js) through the recursive chunking logic in `TextSplitter` to the final vector persistence—using the actual source code from the Mintplex-Labs/anything-llm repository.

## The Upload Pipeline Architecture

AnythingLLM separates document processing concerns across three distinct layers: the API server, the Collector microservice, and the vector database providers. This architecture enables handling of diverse file formats while maintaining consistent **document chunking and embedding processing** strategies across backends like Weaviate, Qdrant, and Chroma.

### API Endpoint: Receiving the Upload

The journey begins at the `/v1/document/upload` endpoint, which authenticates the request and forwards the file to the Collector service:

```javascript
// server/endpoints/api/document/index.js
app.post(
  "/v1/document/upload",
  [validApiKey, handleAPIFileUpload],
  async (request, response) => {
    const { originalname } = request.file;
    const { addToWorkspaces = "", metadata: _metadata = {} } = reqBody(request);
    const metadata = typeof _metadata === "string"
      ? safeJsonParse(_metadata, {})
      : _metadata;

    // Forward the file to the collector service
    const Collector = new CollectorApi();
    const { success, reason, documents } = await Collector.processDocument(
      originalname,
      metadata
    );

```

The `handleAPIFileUpload` middleware extracts the file and metadata, then instantiates `CollectorApi` to delegate processing to the Collector microservice running on port 8888.

### Collector Service: File Type Detection and Conversion

The Collector determines the appropriate converter based on file extension and extracts raw text into a structured document object:

```javascript
// collector/processSingleFile/index.js
async function processSingleFile(targetFilename, options = {}, metadata = {}) {
  const fullFilePath = path.resolve(WATCH_DIRECTORY, normalizePath(targetFilename));
  // …validation omitted for brevity…

  const FileTypeProcessor = require(SUPPORTED_FILETYPE_CONVERTERS[processFileAs]);
  return await FileTypeProcessor({
    fullFilePath,
    filename: targetFilename,
    options,
    metadata,
  });
}

```

For text files, the [`asTxt.js`](https://github.com/Mintplex-Labs/anything-llm/blob/main/asTxt.js) converter reads content, estimates token counts, and constructs a metadata-rich document object:

```javascript
// collector/processSingleFile/convert/asTxt.js
async function asTxt({ fullFilePath, filename, options, metadata }) {
  const content = fs.readFileSync(fullFilePath, "utf8");
  const data = {
    id: v4(),
    url: "file://" + fullFilePath,
    title: metadata.title || filename,
    docAuthor: metadata.docAuthor || "Unknown",
    description: metadata.description || "Unknown",
    docSource: metadata.docSource || "a text file uploaded by the user.",
    chunkSource: metadata.chunkSource || "",
    published: createdDate(fullFilePath),
    wordCount: content.split(" ").length,
    pageContent: content,
    token_count_estimate: tokenizeString(content),
  };
  // ...
}

```

### Persisting Raw Documents

The `writeToServerDocuments` utility persists the document as JSON to the server's filesystem:

```javascript
// collector/utils/files/index.js
function writeToServerDocuments({ data = {}, filename, destinationOverride = null, options = {} }) {
  if (!filename) throw new Error("Filename is required!");
  let destination;
  if (destinationOverride) destination = path.resolve(destinationOverride);
  else if (options.parseOnly) destination = path.resolve(directUploadsFolder);
  else destination = path.resolve(documentsFolder, "custom-documents");

  if (!fs.existsSync(destination)) fs.mkdirSync(destination, { recursive: true });
  const destinationFilePath = normalizePath(path.resolve(destination, filename) + ".json");
  fs.writeFileSync(destinationFilePath, JSON.stringify(data, null, 4), { encoding: "utf-8" });
  return { ...data, location: destinationFilePath.split("/").slice(-2).join("/") };
}

```

The document now resides in `storage/documents/custom-documents/` as a JSON file containing `pageContent` and metadata, ready for chunking.

## Document Chunking and Embedding Processing

Once persisted, the vector database provider loads the document and initiates the **chunking and embedding** sequence.

### TextSplitter: Chunking with Model-Aware Limits

The `TextSplitter` class creates a recursive character-based splitter that respects the configured embedding model's token limitations:

```javascript
// server/utils/vectorDbProviders/weaviate/index.js
const { TextSplitter } = require("../../TextSplitter");
const { getEmbeddingEngineSelection } = require("../../helpers");

// Inside addDocumentToNamespace()
const { pageContent, docId, ...metadata } = documentData;

// Determine chunk size based on embedder limits
const EmbedderEngine = getEmbeddingEngineSelection();
const textSplitter = new TextSplitter({
  chunkSize: TextSplitter.determineMaxChunkSize(
    await SystemSettings.getValueOrFallback({ label: "text_splitter_chunk_size" }),
    EmbedderEngine?.embeddingMaxChunkLength
  ),
  chunkOverlap: await SystemSettings.getValueOrFallback(
    { label: "text_splitter_chunk_overlap" }, 20
  ),
  chunkHeaderMeta: TextSplitter.buildHeaderMeta(metadata),
  chunkPrefix: EmbedderEngine?.embeddingPrefix,
});
const textChunks = await textSplitter.splitText(pageContent);

```

The `determineMaxChunkSize` method ensures chunks never exceed the embedding model's context window (e.g., 8191 tokens for OpenAI embeddings), while `chunkHeaderMeta` prepends document metadata (title, author, source) to each chunk for improved retrieval context.

### Embedding Engine Integration

After splitting, each chunk is sent to the selected embedding engine via the `embedChunks` method:

```javascript
// server/utils/vectorDbProviders/weaviate/index.js
const EmbedderEngine = getEmbeddingEngineSelection(); // e.g., OpenAiEmbedder
const vectorValues = await EmbedderEngine.embedChunks(textChunks);

```

The embedding engine—which may be OpenAI, Voyage AI, or a local provider—returns a dense vector for each text chunk. This abstraction allows AnythingLLM to switch embedding providers without modifying the chunking logic.

### Vector Persistence and Caching

The provider bulk-inserts vectors into the database while simultaneously caching results to disk:

```javascript
// After embedding
const documentVectors = [];
const vectors = [];
for (const [i, vector] of vectorValues.entries()) {
  const id = uuidv4();
  documentVectors.push({ docId, vectorId: id });
  vectors.push({
    id,
    class: camelCase(namespace),
    vector,
    properties: { ...flattenedMetadata, text: textChunks[i] },
  });
}
await this.addVectors(client, vectors);      // bulk insert into Weaviate
await storeVectorResult(chunks, fullFilePath); // cache chunk+vector data on disk

```

The `storeVectorResult` function generates a deterministic cache key using UUIDv5:

```javascript
// collector/utils/files/index.js
async function storeVectorResult(vectorData = [], filename = null) {
  if (!filename) return;
  const digest = uuidv5(filename, uuidv5.URL);
  const writeTo = path.resolve(vectorCachePath, `${digest}.json`);
  fs.writeFileSync(writeTo, JSON.stringify(vectorData), "utf-8");
}

```

This caching mechanism prevents redundant API calls to embedding providers when re-processing identical files, as the system checks `cachedVectorInformation` before initiating new embedding requests.

## Summary

- **Upload Flow**: The `/v1/document/upload` endpoint forwards files to the Collector microservice, which converts them to JSON documents stored in `storage/documents/custom-documents/`.

- **Chunking Strategy**: The `TextSplitter` class creates chunks respecting the embedding model's `embeddingMaxChunkLength`, with configurable overlap and metadata headers.

- **Embedding Generation**: The `getEmbeddingEngineSelection` helper routes chunks to the configured provider (OpenAI, Voyage AI, etc.) via the `embedChunks` method.

- **Persistence**: Vectors are bulk-inserted into the vector database (Weaviate, Qdrant, etc.) while raw chunks and vectors are cached in `vector-cache/<uuid>.json` using a filename-derived UUIDv5 digest.

- **Optimization**: The caching layer eliminates redundant embedding API calls when identical files are re-uploaded.

## Frequently Asked Questions

### What determines the chunk size when processing documents in AnythingLLM?

The chunk size is determined by `TextSplitter.determineMaxChunkSize`, which compares the user's configured `text_splitter_chunk_size` setting against the `embeddingMaxChunkLength` property of the selected embedding engine. This ensures chunks never exceed the model's token limit—for example, 8191 tokens for OpenAI's embedding models—preventing API errors while maximizing context utilization.

### How does AnythingLLM avoid re-embedding the same document?

The system uses `storeVectorResult` to cache embeddings in `vector-cache/` using a UUIDv5 hash of the filename as the cache key. When processing begins, the `cachedVectorInformation` function checks for existing cache entries. If found, the system reuses the stored vectors and chunks, bypassing the embedding engine entirely and eliminating redundant API costs.

### Where does AnythingLLM store documents before they are embedded?

Raw documents are stored as JSON files in `storage/documents/custom-documents/` by the `writeToServerDocuments` function in [`collector/utils/files/index.js`](https://github.com/Mintplex-Labs/anything-llm/blob/main/collector/utils/files/index.js). These files contain the full `pageContent`, metadata (title, author, source), and token count estimates. The vector database providers read from this location to perform chunking and embedding.

### Which embedding providers does AnythingLLM support for document processing?

According to the source code in [`server/utils/helpers/index.js`](https://github.com/Mintplex-Labs/anything-llm/blob/main/server/utils/helpers/index.js), AnythingLLM supports multiple embedding engines including OpenAI, Voyage AI, Azure OpenAI, and local embedding models. The `getEmbeddingEngineSelection` function instantiates the appropriate embedder class (e.g., `OpenAiEmbedder`), which implements the `embedChunks` method used during the vectorization phase.