Calculating Centrality Algorithms in Semantica: Degree, Betweenness, and Closeness
Semantica calculates centrality algorithms through the CentralityCalculator class in semantica/kg/centrality_calculator.py, which automatically uses NetworkX for performance when available or falls back to pure-Python BFS implementations.
The Semantica knowledge-graph toolkit provides a dedicated engine for ranking nodes using classic graph centrality measures. Whether analyzing social networks, knowledge graphs, or dependency trees, calculating centrality algorithms like degree, betweenness, and closeness helps identify the most influential nodes. The implementation automatically optimizes for performance while maintaining accuracy across both accelerated and fallback code paths.
Architecture of the Centrality Engine
The centrality system separates algorithmic logic from visualization and registration concerns. The CentralityCalculator class serves as the core computational engine, while AlgorithmRegistry in semantica/kg/registry.py handles metadata and capability discovery for extensibility.
The calculator integrates with NetworkX when the optional dependency is installed, delegating to optimized C-based implementations. When NetworkX is unavailable, it falls back to pure-Python implementations using adjacency lists and breadth-first search (BFS) algorithms. All calculations wrap ProgressTracker calls from semantica/utils/progress_tracker.py for live pipeline status updates, and use the centralized logger from semantica/utils/logging.py for debugging.
Degree Centrality Implementation
Degree centrality measures node importance by counting direct connections and normalizing by n-1 (where n is the total node count).
Implementation Details:
- NetworkX path: Delegates to
nx.degree_centralityfor maximum performance - Fallback path: Builds an adjacency list via
_build_adjacency()and iterates nodes to compute raw degrees, then applies normalization - Source location: Lines 29-50 in
semantica/kg/centrality_calculator.py
The method calculate_degree_centrality() returns a dictionary containing centrality scores and rankings, consumable by visualization utilities in semantica/visualization/analytics_visualizer.py.
Betweenness Centrality Implementation
Betweenness centrality identifies bridge nodes by counting how often a node appears on shortest paths between all pairs of nodes.
Algorithm Logic:
- For every node pair, finds all shortest paths
- Increments counters for each intermediate node lying on those paths
- Normalizes final scores by the total number of possible node pairs
Implementation Details:
- NetworkX path: Uses
nx.betweenness_centralitywith optimized graph algorithms - Fallback path: Runs BFS from each source node using
_bfs_shortest_paths()to collect all shortest paths and aggregate counts - Source location: Lines 68-86 in
semantica/kg/centrality_calculator.py(corrected from the documented range)
The calculate_betweenness_centrality() method handles both weighted and unweighted graph structures through the same interface.
Closeness Centrality Implementation
Closeness centrality ranks nodes by their average distance to all other reachable nodes, computing the reciprocal of the average shortest-path length.
Algorithm Logic:
- Computes average shortest-path distance from a node to all reachable nodes
- Returns the reciprocal of that average as the closeness score
- Handles disconnected components by considering only reachable nodes
Implementation Details:
- NetworkX path: Delegates to
nx.closeness_centrality - Fallback path: Executes
_bfs_distances()from each node to collect distances, then calculatesreachable / total_distance - Source location: Lines 52-65 in
semantica/kg/centrality_calculator.py
Practical Code Examples
Calculating Degree, Betweenness, and Closeness
from semantica.kg.centrality_calculator import CentralityCalculator
# Initialize the calculator
calculator = CentralityCalculator()
# Assume graph follows Semantica's dict format or NetworkX graph
graph = {
"entities": [{"id": "A"}, {"id": "B"}, {"id": "C"}],
"relationships": [{"source": "A", "target": "B"}, {"source": "B", "target": "C"}]
}
# 1. Degree centrality
deg_result = calculator.calculate_degree_centrality(graph)
print("Top node by degree:", deg_result["rankings"][0])
# 2. Betweenness centrality
bet_result = calculator.calculate_betweenness_centrality(graph)
print("Highest betweenness:", bet_result["rankings"][0])
# 3. Closeness centrality
close_result = calculator.calculate_closeness_centrality(graph)
print("Most central (closeness):", close_result["rankings"][0])
Accessing PageRank via AlgorithmRegistry
from semantica.kg.registry import algorithm_registry
from semantica.kg.centrality_calculator import CentralityCalculator
calculator = CentralityCalculator()
# Retrieve PageRank metadata from registry
pagerank_cls = algorithm_registry.get("centrality", "pagerank")
# Calculate PageRank with custom parameters
pagerank = calculator.calculate_pagerank(
graph,
node_labels=None, # Consider all node types
relationship_types=None, # Consider all edge types
max_iterations=30,
damping_factor=0.85,
)
print("PageRank of node 'A':", pagerank["centrality"].get("A"))
Visualizing Centrality Rankings
from semantica.visualization.analytics_visualizer import visualize_centrality_rankings
# Visualize degree centrality results
visualize_centrality_rankings(
deg_result,
title="Degree Centrality Rankings",
output="interactive", # Renders HTML widget
)
Summary
- The
CentralityCalculatorclass insemantica/kg/centrality_calculator.pyprovides unified access to degree, betweenness, and closeness centrality algorithms. - The implementation automatically leverages NetworkX for performance-critical calculations, falling back to pure-Python BFS algorithms when unavailable.
- Degree centrality counts normalized connections, betweenness measures path interception, and closeness calculates reciprocal average distances.
- The
AlgorithmRegistryinsemantica/kg/registry.pyenables extensibility and lazy loading of algorithm implementations like PageRank. - Built-in progress tracking and logging support long-running graph analyses without custom instrumentation.
- Results integrate directly with
semantica/visualization/analytics_visualizer.pyfor immediate interactive or static visualization.
Frequently Asked Questions
Does Semantica require NetworkX to calculate centrality algorithms?
No, NetworkX is an optional dependency. The CentralityCalculator automatically detects NetworkX availability and delegates to nx.degree_centrality, nx.betweenness_centrality, and nx.closeness_centrality when present. Without NetworkX, it falls back to pure-Python implementations using _build_adjacency() and _bfs_shortest_paths() methods, ensuring functionality across all environments.
How does the AlgorithmRegistry work with centrality calculations?
The AlgorithmRegistry in semantica/kg/registry.py serves as a metadata catalog for algorithm categories including centrality. It registers built-in algorithms like PageRank (lines 58-70) and allows retrieval via algorithm_registry.get("centrality", "pagerank"). The CentralityCalculator uses this registry for lazy loading and extensibility, enabling developers to swap or extend algorithm implementations without modifying the core calculator class.
Can I filter centrality calculations by specific node or relationship types?
Yes, the calculate_pagerank() method accepts node_labels and relationship_types parameters to filter the graph before computation. While degree, betweenness, and closeness methods typically operate on the full graph structure, you can pre-filter your graph dictionary before passing it to these methods to achieve similar subset analysis.
What visualization options exist for centrality results?
The calculator outputs dictionaries compatible with visualize_centrality() and visualize_centrality_rankings() from semantica/visualization/analytics_visualizer.py. These functions accept the result dictionaries directly and support both interactive output (HTML widgets) and static formats, automatically extracting the "rankings" and "centrality" keys from the calculator output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →