Hub Nodes vs Bridge Nodes in code-review-graph: Identifying Critical Code Architecture Patterns
Hub nodes are highly-connected components ranked by total degree, while bridge nodes are architectural connectors ranked by betweenness centrality—both reveal different risks in your codebase graph.
In the tirth8205/code-review-graph open-source project, hub and bridge detection provides two complementary lenses for analyzing code structure. These graph-based metrics help developers identify hotspots that could cascade failures and chokepoints that hide coupling between modules.
What Hub Nodes Represent in Code Graphs
Hub nodes capture the most connected functions, classes, or symbols in your codebase based on total degree—the sum of incoming plus outgoing edges.
How Hub Detection Works
The find_hub_nodes function in code_review_graph/analysis.py (lines 14-55) implements this by:
- Aggregating degree counts from every edge in the graph
- Filtering nodes with
total_degree > 0 - Sorting by connectivity score
A hub represents a "hot spot" in your architecture. Changes to hub nodes affect many callers and callees, making them high-impact targets for testing and careful review.
Using the Hub Detection Tool
from code_review_graph.main import get_hub_nodes_tool
hub_info = get_hub_nodes_tool(top_n=5)
print("Top hubs:")
for h in hub_info["hub_nodes"]:
print(f"- {h['qualified_name']} (degree={h['total_degree']})")
The get_hub_nodes_tool wrapper returns a dictionary with hub_nodes, count, and provenance metadata. Use this to ask: "Does this high-degree node have adequate test coverage?"
What Bridge Nodes Represent in Code Graphs
Bridge nodes identify architectural chokepoints—components that sit on many shortest paths between otherwise separate parts of the codebase.
How Bridge Detection Works
The find_bridge_nodes function in code_review_graph/analysis.py (lines 58-112) computes betweenness centrality using NetworkX:
- Exact calculation for graphs with ≤ 5,000 nodes
- Sampled approximation for larger graphs
High betweenness centrality indicates a node acts as a connector between communities. If a bridge fails, entire subgraphs can become disconnected.
Using the Bridge Detection Tool
from code_review_graph.main import get_bridge_nodes_tool
bridge_info = get_bridge_nodes_tool(top_n=5)
print("\nTop bridges:")
for b in bridge_info["bridge_nodes"]:
print(f"- {b['qualified_name']} (betweenness={b['betweenness']})")
The get_bridge_nodes_tool helps answer: "Why does this node connect communities A and B?" Bridges often reveal hidden coupling between modules that should be decoupled.
Key Differences Between Hub and Bridge Nodes
| Dimension | Hub Nodes | Bridge Nodes |
|---|---|---|
| Centrality metric | Degree centrality (total connections) | Betweenness centrality (path intermediacy) |
| Structural role | Local popularity | Global connectivity |
| Risk type | Cascading impact from changes | Disconnection of subsystems |
| Detection function | find_hub_nodes (lines 14-55) |
find_bridge_nodes (lines 58-112) |
| Tool wrapper | get_hub_nodes_tool |
get_bridge_nodes_tool |
Practical Applications in Code Review
Prioritize Testing Efforts
Target hub nodes first when allocating test resources—their high connectivity means bugs propagate widely.
Identify Architectural Debt
Bridge nodes often indicate accidental coupling between modules. Consider refactoring to eliminate unnecessary bridges and improve modularity.
Evaluate Refactoring Impact
Before removing a node, check both its hub score (how many dependencies break) and bridge score (whether communities split).
Summary
- Hub nodes measure local connectivity via
total_degree—catch hotspots with widespread impact - Bridge nodes measure global intermediacy via
betweenness—catch architectural chokepoints - Both metrics are implemented in
code_review_graph/analysis.pywith CLI tools incode_review_graph/main.py - Use hubs to prioritize testing; use bridges to detect hidden coupling
Frequently Asked Questions
How do I interpret a node that scores high on both hub and bridge metrics?
This indicates a critical architectural component that is both heavily connected and structurally irreplaceable. Such nodes demand the highest scrutiny—their failure would both break many direct dependencies and disconnect entire subsystems. Consider adding circuit breakers, extensive tests, or refactoring to distribute the load.
Does the betweenness calculation use exact or approximate methods?
According to the code-review-graph source code, find_bridge_nodes uses exact betweenness centrality for graphs with 5,000 nodes or fewer. For larger codebases, it automatically switches to sampled approximation to maintain performance. This threshold balances accuracy with computational feasibility.
What input format does the hub and bridge detection expect?
The detection functions operate on parsed code graphs built from your repository. The tools get_hub_nodes_tool and get_bridge_nodes_tool handle graph construction internally—you only specify top_n to control result set size. The underlying analysis expects NetworkX-compatible graph structures with qualified node names representing functions, classes, or symbols.
Where are these tools documented beyond the source code?
Command-line usage is documented in docs/COMMANDS.md at the repository root. The high-level concepts appear in README.md under the "Hub & bridge detection" section. For implementation details, reference the inline comments in code_review_graph/analysis.py lines 14-112.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →