# How index_status Reports parse_partial Files and What to Do About Them

> Discover how index_status flags parse_partial files, identifies indexing gaps, and guides you to fix syntax or encoding issues for complete graph coverage.

- Repository: [Martin Vogel/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp)
- Tags: how-to-guide
- Published: 2026-07-10

---

**`index_status` returns a coverage report that flags `parse_partial` files—documents that were indexed but contain line ranges where tree-sitter failed to parse, stored in the file node's `detail` field—and you should inspect those ranges, fix underlying syntax errors or encoding issues, and re-index to achieve complete graph coverage.**

The `index_status` tool in the **DeusData/codebase-memory-mcp** repository provides a critical visibility layer into your project's indexing health. While it reports high-level metrics like node and edge counts, its **coverage report** specifically distinguishes between files that were skipped entirely and those that were only partially parsed. Understanding the `parse_partial` status is essential for maintaining a complete knowledge graph, as unresolved gaps can cause `search_graph` and `trace_path` queries to miss symbol definitions located inside unparsed ranges.

## What parse_partial Means in index_status

When the indexing pipeline processes a source file, tree-sitter attempts to build a complete Abstract Syntax Tree (AST). If the parser encounters errors, missing tokens, or unrecoverable syntax issues, it cannot generate nodes for those specific regions. Instead of failing the entire file, the pipeline creates a **"parse_partial"** node.

According to the source code in [`src/store/store.h`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/store/store.h) (line 404), a file node with `kind="parse_partial"` indicates that the file **was indexed** but the parse tree contained errors or missing regions. The exact locations are recorded in the `detail` field as comma-separated line ranges, for example `"12-34, 78-89"`.

The generation logic resides in [`src/pipeline/pipeline.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c) within the function `add_parse_partial_summary`. This routine assembles the JSON payload that `index_status` ultimately returns, populating the `parse_partial` and `parse_partial_count` fields. The tool definition itself, including the description of these coverage flags, is located in [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c) between lines 5030-5033. Additionally, the UI exposes this data via the `/api/index-status` endpoint implemented in [`src/ui/http_server.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/ui/http_server.c) in the `handle_index_status` function.

## Why parse_partial Files Matter for Graph Queries

The knowledge graph only contains symbols that were successfully parsed. Any code constructs—functions, classes, or variables—residing inside a `parse_partial` range are **not guaranteed to be present** in the graph. This directly impacts the reliability of downstream tools:

- **`search_graph`** may fail to return definitions located in unparsed blocks.
- **`trace_path`** can produce incomplete call chains if a function definition falls within a gap.
- **Code navigation** in the UI will not provide "Go to Definition" for symbols in those ranges.

Treating `parse_partial` entries as actionable signals rather than warnings ensures your semantic queries return complete and accurate results.

## How to Identify parse_partial Ranges

When you invoke `index_status`, the returned JSON includes a coverage section. Each `parse_partial` entry contains:

- **`kind`**: Set to `"parse_partial"`.
- **`detail`**: A string listing the problematic line ranges (e.g., `"45-50, 120-135"`).

For a programmatic view of these gaps, use the `query_graph` tool with the parameter `graph="missed"` (as documented in [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c)). This returns `File` nodes and their associated `kind`/`detail` fields, which is useful for automation or CI pipelines.

## Resolving parse_partial Files

Addressing these files requires a systematic approach from inspection to remediation.

### Inspecting the Reported Line Ranges

Start by opening the flagged file and navigating to the exact lines specified in the `detail` field. If you need to locate specific symbols that might be missing, run a text search limited to those ranges:

```bash
sed -n '45,50p' problematic_file.py | grep -n "function_name"

```

### Common Root Causes and Fixes

The majority of `parse_partial` errors stem from fixable source code issues:

- **Syntax errors or incomplete code**: Stray brackets, missing semicolons, or unclosed blocks prevent the parser from building a valid AST.
- **Missing includes or imports**: If the file references symbols from headers or modules that are not resolvable, tree-sitter may fail to parse dependent constructs.
- **Non-UTF-8 characters**: Invalid encoding can break the lexer before it reaches critical tokens.
- **Conditional compilation**: Preprocessor directives that leave the file in an syntactically unexpected state (common in C/C++) can confuse the parser.

Correct these issues in the source file directly. If the file is generated code or intentionally uses non-standard syntax that cannot be parsed, proceed to the exclusion step instead.

### Re-indexing After Fixes

Once you have corrected the underlying issues—or added necessary include paths—trigger a fresh indexing run:

```json
{
  "tool": "index_repository",
  "params": {}
}

```

Upon completion, invoke `index_status` again. The file should no longer appear in the `parse_partial` list if the parser successfully processed the previously problematic ranges.

### Excluding Intentionally Broken Files

If a file is meant to be unparseable (e.g., machine-generated code, proprietary binary formats with code extensions), add it to the `.cbmignore` file or the appropriate ignore list. This moves the file from the `parse_partial` category to the deliberately ignored "not indexed" category, cleaning up your coverage report and preventing false positives in health checks.

## Programmatically Querying Missed Symbols

For automation workflows, you can extract a structured list of all gaps without parsing the `index_status` JSON manually. The `query_graph` tool supports a specific graph mode for this purpose:

```json
{
  "tool": "query_graph",
  "params": {
    "graph": "missed"
  }
}

```

This returns nodes where `kind="parse_partial"` alongside their `detail` fields, allowing scripts to generate tickets or reports for developers to fix syntax errors before they affect production queries.

## Summary

- **`index_status`** reports `parse_partial` files in its coverage section, populated by `add_parse_partial_summary` in [`src/pipeline/pipeline.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/pipeline/pipeline.c).
- These files were indexed but contain line ranges where tree-sitter failed, recorded in the `detail` field as `"start-end"` pairs.
- Unparsed ranges cause **missing symbols** in graph queries like `search_graph` and `trace_path`.
- **Remediation** involves inspecting the line ranges, fixing syntax/encoding issues, and re-indexing.
- **Intentionally unparseable** files should be added to `.cbmignore` to exclude them from coverage metrics.
- Use **`query_graph`** with `graph="missed"` to programmatically audit these gaps.

## Frequently Asked Questions

### How is parse_partial different from skipped in index_status?

According to the `index_status` implementation in [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c), **`skipped`** files were not indexed at all—typically due to size limits, read permissions, or outright parse failures—while **`parse_partial`** files were successfully read and partially indexed, but contain specific line ranges where the tree-sitter parser encountered errors. The `parse_partial` status includes a `detail` field with the specific line ranges, whereas `skipped` files simply lack any representation in the graph.

### Can I still query code inside parse_partial ranges?

No. Because the AST contains gaps in those specific line ranges, the graph does not contain nodes for symbols defined within them. Queries such as `search_graph` or `trace_path` will not return results for definitions located inside the ranges listed in the `detail` field. You must fix the source code and re-index to make those symbols queryable.

### What should I do if a parse_partial file contains valid code that tree-sitter cannot parse?

If the code is valid in a specific language dialect or uses macros that tree-sitter does not understand, you have two options: First, attempt to simplify the syntax or add necessary includes/imports that help the parser resolve the constructs. Second, if the file is valid but unparseable by design (e.g., heavy template metaprogramming), add the file pattern to `.cbmignore` to exclude it from indexing entirely. This removes it from the `parse_partial` list and prevents it from polluting your coverage metrics.

### How do I automate detection of parse_partial files in CI?

Use the `query_graph` tool with the `graph="missed"` parameter, as defined in [`src/mcp/mcp.c`](https://github.com/DeusData/codebase-memory-mcp/blob/main/src/mcp/mcp.c). This returns a structured list of all `parse_partial` file nodes and their `detail` fields. You can script this call in your CI pipeline to fail builds or generate reports when the `parse_partial_count` exceeds zero, ensuring syntax errors are caught before they impact downstream semantic analysis.