Does code-review-graph Support Analysis of Jupyter Notebooks?

Yes, code-review-graph fully supports Jupyter Notebook analysis by treating .ipynb files as multi-language source containers, extracting code cells, detecting languages via magic commands, and parsing each cell with tree-sitter while preserving cell metadata.

The code-review-graph project provides static analysis capabilities for polyglot codebases. Understanding how this tool handles Jupyter notebooks is essential for data science teams who need to map dependencies across Python, SQL, and R cells within a single notebook file.

How Notebook Detection Works

The parser identifies Jupyter notebooks through an extension-to-language mapping defined in code_review_graph/parser.py. At lines 554-557, the EXTENSION_TO_LANGUAGE dictionary includes the entry ".ipynb": "notebook", routing all files with the notebook extension to specialized parsing logic.

When you invoke CodeParser.parse_file() on a notebook path, the method checks this mapping and delegates to the internal _parse_notebook method rather than the standard single-language parser.

The Notebook Parsing Pipeline

According to the source code in tirth8205/code-review-graph, notebook analysis follows a structured extraction pipeline:

  • Entry Point: At lines 2638-2642 in parser.py, parse_file detects the "notebook" language tag and calls _parse_notebook.
  • JSON Extraction: The _parse_notebook method (lines 3211-3225) reads the notebook JSON structure, discovers the kernel language, and constructs a list of CellInfo objects representing each cell.
  • Cell Grouping: At lines 3295-3350, cells are grouped by detected language, concatenated into source buffers, and parsed with the appropriate tree-sitter parser. Line numbers are rebased per language group, and cell indexes are stored in each node’s extra field.

Cell Language Detection and Magic Command Handling

The parser implements sophisticated magic command detection at lines 33-66 of parser.py to handle polyglot notebooks:

  • Language Switching: Prefixes like %python, %sql, and %r switch the cell language to the specified interpreter.
  • Cell Skipping: Magics such as %scala, %md, and %sh cause the cell to be skipped entirely, filtering out markdown and shell cells that do not contribute executable code to the graph.

For Python and R cells, the system sanitizes interactive content. At lines 67-73, any line beginning with % or ! is removed before parsing to prevent tree-sitter syntax errors on magic commands and shell escapes.

SQL Cell Analysis and Table Dependencies

Beyond standard programming languages, code-review-graph extracts database dependencies from SQL cells. When a cell is identified as SQL (via %sql magic), the parser applies the _SQL_TABLE_RE regex at lines 3266-3278 to identify table references. It then generates IMPORTS_FROM edges linking the notebook file node to the discovered database tables, enabling cross-language dependency tracking from Python analysis code to SQL data sources.

Practical Example: Parsing a Notebook with CodeParser

You can analyze Jupyter notebooks programmatically using the CodeParser class. The following example demonstrates extraction of functions, cell index tracking, and SQL dependency identification:

from pathlib import Path
from code_review_graph.parser import CodeParser

# Parse a Jupyter notebook

parser = CodeParser()
nodes, edges = parser.parse_file(Path("example_notebook.ipynb"))

# Inspect the top-level File node

file_node = next(n for n in nodes if n.kind == "File")
print("File language:", file_node.language)        # → python (or r)

print("Notebook format tag:", file_node.extra.get("notebook_format"))

# Show functions defined in cell 2

funcs = [n for n in nodes if n.kind == "Function" and n.extra.get("cell_index") == 2]
for f in funcs:
    print(f.name, f.extra)   # e.g. add, multiply

# Show imported SQL tables from a %sql cell

sql_imports = [e for e in edges if e.kind == "IMPORTS_FROM"]
for imp in sql_imports:
    print("SQL table:", imp.target)

Running this against a multi-cell notebook produces output that tracks the origin of each code element:


File language: python
add {'cell_index': 2}
multiply {'cell_index': 2}
SQL table: catalog.schema.raw_data

Summary

  • code-review-graph registers .ipynb files via EXTENSION_TO_LANGUAGE in code_review_graph/parser.py (lines 554-557).
  • The _parse_notebook method (lines 3211-3225) extracts cell content and metadata from notebook JSON.
  • Magic commands drive language detection (lines 33-66), while shell-style lines are filtered (lines 67-73) before parsing.
  • SQL cells generate dependency edges via _SQL_TABLE_RE parsing (lines 3266-3278).
  • Cell provenance is preserved through the cell_index field in node metadata, enabling precise mapping between graph nodes and original notebook cells.

Frequently Asked Questions

How does code-review-graph handle magic commands in Jupyter notebooks?

The parser inspects cell content for magic prefixes at lines 33-66 of parser.py. Commands like %python or %r switch the cell's language context, while %md triggers cell skipping. For Python and R cells, lines starting with % or ! are stripped (lines 67-73) before tree-sitter parsing to avoid syntax errors.

Can code-review-graph extract dependencies from SQL cells in notebooks?

Yes. When a cell is detected as SQL (via %sql magic), the parser applies the _SQL_TABLE_RE regex at lines 3266-3278 to identify referenced tables. It creates IMPORTS_FROM edges linking the notebook file node to the discovered table entities, allowing the graph to track data dependencies from analysis code to database sources.

Where is the notebook parsing logic tested in the repository?

The test suite in tests/test_notebook.py validates the complete notebook pipeline, including language detection, cell-index tracking, cross-cell call graphs, import extraction, magic-line filtering, and handling of empty or malformed notebooks.

What happens to markdown cells when parsing a Jupyter notebook?

Markdown cells are explicitly skipped. The logic at lines 33-66 of parser.py identifies %md magic commands and markdown cell types, excluding them from the parse tree since they do not contain executable code relevant to the dependency graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →