How Semantic Search Indexing Works in Forge: A Technical Deep Dive
Forge’s semantic search indexing uses a hash-based differential sync to transmit only changed file content to the workspace server, then executes vector similarity queries against embedded representations of your codebase.
Forge, the AI-assisted coding platform from the antinomyhq/forgecode repository, implements semantic search indexing through a client-server architecture that minimizes data transfer while maintaining accurate vector representations of your codebase. The system operates through a precise two-phase workflow: differential workspace synchronization followed by embedding-based retrieval. Understanding these mechanics reveals how Forge balances privacy, performance, and search accuracy.
The Two-Phase Semantic Indexing Architecture
Forge’s semantic search relies on distinct synchronization and query phases, each optimized to reduce network overhead while maintaining index freshness.
Phase 1: Differential Workspace Synchronization
The indexing process begins when ForgeWorkspaceService::sync_workspace—defined in [crates/forge_services/src/context_engine.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/context_engine.rs)—orchestrates a differential sync with the remote workspace server. Rather than uploading your entire codebase, the system computes content hashes locally and transmits only modifications.
The heavy lifting occurs in WorkspaceSyncEngine within [crates/forge_services/src/sync.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/sync.rs), which executes a four-stage pipeline:
- Discovery: The engine calls
discover_sync_file_pathsto identify all files subject to indexing. - Local Hashing: Via
read_hashes, the system computes content hashes without loading full file contents into memory. - Remote Comparison: The client fetches existing remote hashes using
fetch_remote_hashes, then computes a diff throughWorkspaceStatus::new. - Selective Transfer: Only new or modified files proceed to
upload_files, while deleted paths triggerdelete_files.
This differential approach ensures that only changed content traverses the network, significantly reducing bandwidth usage for large codebases.
Phase 2: Semantic Query Execution
Once synchronized, semantic searches execute through ForgeWorkspaceService::query_workspace (also in [context_engine.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/context_engine.rs)). The client constructs a SearchParams object containing your natural language query and optional reranking parameters, then forwards the request to the infrastructure layer via infra.search(&search_query, &token).
The server transforms your text query into a vector embedding and performs similarity matching against the stored file-content embeddings. Results return as Node structures—defined in [crates/forge_domain/src/node.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_domain/src/node.rs)—containing precise file:line references and surrounding code snippets.
What Data Is Transmitted to the Workspace Server?
Understanding the exact payload structure helps clarify Forge’s privacy model and network efficiency.
During Synchronization
The sync operation transmits three distinct data types:
- File hashes: A
Vec<FileHash>containing content hashes for all discovered files, enabling the server to identify differences without accessing file contents. - File contents: For each modified or new file, a
FileReadpayload containing the absolute path string and UTF-8 encoded file content. - Deletion markers: A
CodeBaserequest containing path strings for files that no longer exist locally.
The hash-first methodology ensures that unchanged files—regardless of size—never transfer their contents across the network.
During Search Queries
Query operations transmit minimal data:
- Identity context: Your user ID and workspace ID within a
CodeBasestructure. - Search parameters: A
SearchParamsobject (from [crates/forge_domain/src/search.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_domain/src/search.rs)) containing the embedding query string and optional reranking "use case" identifiers. - Vector embeddings: Generated server-side from your query text, never exposing raw code during the search phase.
Practical Implementation Examples
Command-Line Synchronization
Initiate indexing through the Forge CLI:
# Creates or updates the remote semantic index
forge workspace sync
The CLI implementation in [crates/forge_main/src/ui.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_main/src/ui.rs) wires this command to the service layer, streaming progress events as synchronization occurs.
Programmatic Sync Usage
For custom tooling, interact with the service directly:
use std::path::PathBuf;
let service = workspace_service.clone();
let mut stream = service.sync_workspace(PathBuf::from(".")).await?;
while let Some(event) = stream.next().await {
println!("Sync progress: {:?}", event?);
}
Executing Semantic Searches
After synchronization, query the indexed workspace:
# Search for implementation patterns
forge workspace query "exponential backoff retry mechanism"
Or via the Rust API:
use forge_domain::SearchParams;
let params = SearchParams::new(
"exponential backoff retry mechanism".into(),
Some("implementation".into())
);
let results = workspace_service
.query_workspace(PathBuf::from("."), params)
.await?;
for node in results {
println!("{}:{} – {}", node.file, node.line, node.snippet);
}
Summary
- Differential syncing via content hashing minimizes data transfer by transmitting only new, modified, or deleted files.
- Workspace synchronization is handled by
ForgeWorkspaceService::sync_workspaceandWorkspaceSyncEnginein theforge_servicescrate. - Hash comparison occurs locally; only
Vec<FileHash>metadata and changedFileReadpayloads cross the network. - Semantic queries transmit only
SearchParams(text queries and reranking hints), receiving lightweightNoderesults with file locations and snippets. - Core files include [
context_engine.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/context_engine.rs) for orchestration and [sync.rs](https://github.com/antinomyhq/forgecode/blob/main/crates/forge_services/src/sync.rs) for the hashing and diff logic.
Frequently Asked Questions
Does Forge send my entire codebase to the workspace server?
No. Forge’s WorkspaceSyncEngine computes content hashes locally and compares them against remote hashes before transferring any data. Only files with new or modified content are uploaded as FileRead payloads, while unchanged files transmit only their hash identifiers.
What personal information is included in semantic search requests?
Search requests include your user ID, workspace ID, and the natural language query string within a CodeBase structure. The actual source code contents remain on the server from the previous sync; queries only transmit the search parameters needed to generate vector embeddings and retrieve relevant Node references.
How does Forge handle file deletions during synchronization?
When WorkspaceStatus::new detects files present in the remote index but absent locally, it generates deletion instructions. These transmit as path strings within a CodeBase request to the delete_files endpoint, ensuring the remote vector index stays consistent with your local filesystem.
Can I use Forge’s semantic search without internet access?
No. The semantic search implementation requires connectivity to the workspace server for both synchronization (uploading hashes and changed file contents) and query execution (server-side embedding generation and vector similarity search against the hosted index).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →