How Semantic Search Indexing Works in Forge: A Technical Deep Dive

Forge’s semantic search indexing uses a hash-based differential sync to transmit only changed file content to the workspace server, then executes vector similarity queries against embedded representations of your codebase.

Forge, the AI-assisted coding platform from the antinomyhq/forgecode repository, implements semantic search indexing through a client-server architecture that minimizes data transfer while maintaining accurate vector representations of your codebase. The system operates through a precise two-phase workflow: differential workspace synchronization followed by embedding-based retrieval. Understanding these mechanics reveals how Forge balances privacy, performance, and search accuracy.

The Two-Phase Semantic Indexing Architecture

Forge’s semantic search relies on distinct synchronization and query phases, each optimized to reduce network overhead while maintaining index freshness.

Phase 1: Differential Workspace Synchronization

The indexing process begins when ForgeWorkspaceService::sync_workspace—defined in [crates/forge_services/src/context_engine.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/context_engine.rs)—orchestrates a differential sync with the remote workspace server. Rather than uploading your entire codebase, the system computes content hashes locally and transmits only modifications.

The heavy lifting occurs in WorkspaceSyncEngine within [crates/forge_services/src/sync.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/sync.rs), which executes a four-stage pipeline:

  1. Discovery: The engine calls discover_sync_file_paths to identify all files subject to indexing.
  2. Local Hashing: Via read_hashes, the system computes content hashes without loading full file contents into memory.
  3. Remote Comparison: The client fetches existing remote hashes using fetch_remote_hashes, then computes a diff through WorkspaceStatus::new.
  4. Selective Transfer: Only new or modified files proceed to upload_files, while deleted paths trigger delete_files.

This differential approach ensures that only changed content traverses the network, significantly reducing bandwidth usage for large codebases.

Phase 2: Semantic Query Execution

Once synchronized, semantic searches execute through ForgeWorkspaceService::query_workspace (also in [context_engine.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/context_engine.rs)). The client constructs a SearchParams object containing your natural language query and optional reranking parameters, then forwards the request to the infrastructure layer via infra.search(&search_query, &token).

The server transforms your text query into a vector embedding and performs similarity matching against the stored file-content embeddings. Results return as Node structures—defined in [crates/forge_domain/src/node.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_domain/src/node.rs)—containing precise file:line references and surrounding code snippets.

What Data Is Transmitted to the Workspace Server?

Understanding the exact payload structure helps clarify Forge’s privacy model and network efficiency.

During Synchronization

The sync operation transmits three distinct data types:

  • File hashes: A Vec<FileHash> containing content hashes for all discovered files, enabling the server to identify differences without accessing file contents.
  • File contents: For each modified or new file, a FileRead payload containing the absolute path string and UTF-8 encoded file content.
  • Deletion markers: A CodeBase request containing path strings for files that no longer exist locally.

The hash-first methodology ensures that unchanged files—regardless of size—never transfer their contents across the network.

During Search Queries

Query operations transmit minimal data:

Practical Implementation Examples

Command-Line Synchronization

Initiate indexing through the Forge CLI:


# Creates or updates the remote semantic index

forge workspace sync

The CLI implementation in [crates/forge_main/src/ui.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_main/src/ui.rs) wires this command to the service layer, streaming progress events as synchronization occurs.

Programmatic Sync Usage

For custom tooling, interact with the service directly:

use std::path::PathBuf;

let service = workspace_service.clone();
let mut stream = service.sync_workspace(PathBuf::from(".")).await?;

while let Some(event) = stream.next().await {
    println!("Sync progress: {:?}", event?);
}

Executing Semantic Searches

After synchronization, query the indexed workspace:


# Search for implementation patterns

forge workspace query "exponential backoff retry mechanism"

Or via the Rust API:

use forge_domain::SearchParams;

let params = SearchParams::new(
    "exponential backoff retry mechanism".into(),
    Some("implementation".into())
);

let results = workspace_service
    .query_workspace(PathBuf::from("."), params)
    .await?;

for node in results {
    println!("{}:{} – {}", node.file, node.line, node.snippet);
}

Summary

Frequently Asked Questions

Does Forge send my entire codebase to the workspace server?

No. Forge’s WorkspaceSyncEngine computes content hashes locally and compares them against remote hashes before transferring any data. Only files with new or modified content are uploaded as FileRead payloads, while unchanged files transmit only their hash identifiers.

What personal information is included in semantic search requests?

Search requests include your user ID, workspace ID, and the natural language query string within a CodeBase structure. The actual source code contents remain on the server from the previous sync; queries only transmit the search parameters needed to generate vector embeddings and retrieve relevant Node references.

How does Forge handle file deletions during synchronization?

When WorkspaceStatus::new detects files present in the remote index but absent locally, it generates deletion instructions. These transmit as path strings within a CodeBase request to the delete_files endpoint, ensuring the remote vector index stays consistent with your local filesystem.

Can I use Forge’s semantic search without internet access?

No. The semantic search implementation requires connectivity to the workspace server for both synchronization (uploading hashes and changed file contents) and query execution (server-side embedding generation and vector similarity search against the hosted index).

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →