# Does code-review-graph Support Analysis of Jupyter Notebooks?

> Yes code-review-graph supports Jupyter Notebook analysis by extracting code cells and parsing them with tree-sitter preserving metadata.

- Repository: [Tirth Kanani/code-review-graph](https://github.com/tirth8205/code-review-graph)
- Tags: how-to-guide
- Published: 2026-08-13

---

**Yes, code-review-graph fully supports Jupyter Notebook analysis by treating `.ipynb` files as multi-language source containers, extracting code cells, detecting languages via magic commands, and parsing each cell with tree-sitter while preserving cell metadata.**

The **code-review-graph** project provides static analysis capabilities for polyglot codebases. Understanding how this tool handles Jupyter notebooks is essential for data science teams who need to map dependencies across Python, SQL, and R cells within a single notebook file.

## How Notebook Detection Works

The parser identifies Jupyter notebooks through an extension-to-language mapping defined in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py). At **lines 554-557**, the `EXTENSION_TO_LANGUAGE` dictionary includes the entry `".ipynb": "notebook"`, routing all files with the notebook extension to specialized parsing logic.

When you invoke `CodeParser.parse_file()` on a notebook path, the method checks this mapping and delegates to the internal `_parse_notebook` method rather than the standard single-language parser.

## The Notebook Parsing Pipeline

According to the source code in `tirth8205/code-review-graph`, notebook analysis follows a structured extraction pipeline:

- **Entry Point**: At **lines 2638-2642** in [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py), `parse_file` detects the `"notebook"` language tag and calls `_parse_notebook`.
- **JSON Extraction**: The `_parse_notebook` method (**lines 3211-3225**) reads the notebook JSON structure, discovers the kernel language, and constructs a list of `CellInfo` objects representing each cell.
- **Cell Grouping**: At **lines 3295-3350**, cells are grouped by detected language, concatenated into source buffers, and parsed with the appropriate tree-sitter parser. Line numbers are rebased per language group, and cell indexes are stored in each node’s `extra` field.

## Cell Language Detection and Magic Command Handling

The parser implements sophisticated magic command detection at **lines 33-66** of [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) to handle polyglot notebooks:

- **Language Switching**: Prefixes like `%python`, `%sql`, and `%r` switch the cell language to the specified interpreter.
- **Cell Skipping**: Magics such as `%scala`, `%md`, and `%sh` cause the cell to be skipped entirely, filtering out markdown and shell cells that do not contribute executable code to the graph.

For Python and R cells, the system sanitizes interactive content. At **lines 67-73**, any line beginning with `%` or `!` is removed before parsing to prevent tree-sitter syntax errors on magic commands and shell escapes.

## SQL Cell Analysis and Table Dependencies

Beyond standard programming languages, code-review-graph extracts database dependencies from SQL cells. When a cell is identified as SQL (via `%sql` magic), the parser applies the `_SQL_TABLE_RE` regex at **lines 3266-3278** to identify table references. It then generates `IMPORTS_FROM` edges linking the notebook file node to the discovered database tables, enabling cross-language dependency tracking from Python analysis code to SQL data sources.

## Practical Example: Parsing a Notebook with CodeParser

You can analyze Jupyter notebooks programmatically using the `CodeParser` class. The following example demonstrates extraction of functions, cell index tracking, and SQL dependency identification:

```python
from pathlib import Path
from code_review_graph.parser import CodeParser

# Parse a Jupyter notebook

parser = CodeParser()
nodes, edges = parser.parse_file(Path("example_notebook.ipynb"))

# Inspect the top-level File node

file_node = next(n for n in nodes if n.kind == "File")
print("File language:", file_node.language)        # → python (or r)

print("Notebook format tag:", file_node.extra.get("notebook_format"))

# Show functions defined in cell 2

funcs = [n for n in nodes if n.kind == "Function" and n.extra.get("cell_index") == 2]
for f in funcs:
    print(f.name, f.extra)   # e.g. add, multiply

# Show imported SQL tables from a %sql cell

sql_imports = [e for e in edges if e.kind == "IMPORTS_FROM"]
for imp in sql_imports:
    print("SQL table:", imp.target)

```

Running this against a multi-cell notebook produces output that tracks the origin of each code element:

```

File language: python
add {'cell_index': 2}
multiply {'cell_index': 2}
SQL table: catalog.schema.raw_data

```

## Summary

- **code-review-graph** registers `.ipynb` files via `EXTENSION_TO_LANGUAGE` in [`code_review_graph/parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/code_review_graph/parser.py) (**lines 554-557**).
- The `_parse_notebook` method (**lines 3211-3225**) extracts cell content and metadata from notebook JSON.
- Magic commands drive language detection (**lines 33-66**), while shell-style lines are filtered (**lines 67-73**) before parsing.
- SQL cells generate dependency edges via `_SQL_TABLE_RE` parsing (**lines 3266-3278**).
- Cell provenance is preserved through the `cell_index` field in node metadata, enabling precise mapping between graph nodes and original notebook cells.

## Frequently Asked Questions

### How does code-review-graph handle magic commands in Jupyter notebooks?

The parser inspects cell content for magic prefixes at **lines 33-66** of [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py). Commands like `%python` or `%r` switch the cell's language context, while `%md` triggers cell skipping. For Python and R cells, lines starting with `%` or `!` are stripped (**lines 67-73**) before tree-sitter parsing to avoid syntax errors.

### Can code-review-graph extract dependencies from SQL cells in notebooks?

Yes. When a cell is detected as SQL (via `%sql` magic), the parser applies the `_SQL_TABLE_RE` regex at **lines 3266-3278** to identify referenced tables. It creates `IMPORTS_FROM` edges linking the notebook file node to the discovered table entities, allowing the graph to track data dependencies from analysis code to database sources.

### Where is the notebook parsing logic tested in the repository?

The test suite in [`tests/test_notebook.py`](https://github.com/tirth8205/code-review-graph/blob/main/tests/test_notebook.py) validates the complete notebook pipeline, including language detection, cell-index tracking, cross-cell call graphs, import extraction, magic-line filtering, and handling of empty or malformed notebooks.

### What happens to markdown cells when parsing a Jupyter notebook?

Markdown cells are explicitly skipped. The logic at **lines 33-66** of [`parser.py`](https://github.com/tirth8205/code-review-graph/blob/main/parser.py) identifies `%md` magic commands and markdown cell types, excluding them from the parse tree since they do not contain executable code relevant to the dependency graph.