Calculating Centrality Algorithms in Semantica: Degree, Betweenness, and Closeness

Semantica calculates centrality algorithms through the CentralityCalculator class in semantica/kg/centrality_calculator.py, which automatically uses NetworkX for performance when available or falls back to pure-Python BFS implementations.

The Semantica knowledge-graph toolkit provides a dedicated engine for ranking nodes using classic graph centrality measures. Whether analyzing social networks, knowledge graphs, or dependency trees, calculating centrality algorithms like degree, betweenness, and closeness helps identify the most influential nodes. The implementation automatically optimizes for performance while maintaining accuracy across both accelerated and fallback code paths.

Architecture of the Centrality Engine

The centrality system separates algorithmic logic from visualization and registration concerns. The CentralityCalculator class serves as the core computational engine, while AlgorithmRegistry in semantica/kg/registry.py handles metadata and capability discovery for extensibility.

The calculator integrates with NetworkX when the optional dependency is installed, delegating to optimized C-based implementations. When NetworkX is unavailable, it falls back to pure-Python implementations using adjacency lists and breadth-first search (BFS) algorithms. All calculations wrap ProgressTracker calls from semantica/utils/progress_tracker.py for live pipeline status updates, and use the centralized logger from semantica/utils/logging.py for debugging.

Degree Centrality Implementation

Degree centrality measures node importance by counting direct connections and normalizing by n-1 (where n is the total node count).

Implementation Details:

  • NetworkX path: Delegates to nx.degree_centrality for maximum performance
  • Fallback path: Builds an adjacency list via _build_adjacency() and iterates nodes to compute raw degrees, then applies normalization
  • Source location: Lines 29-50 in semantica/kg/centrality_calculator.py

The method calculate_degree_centrality() returns a dictionary containing centrality scores and rankings, consumable by visualization utilities in semantica/visualization/analytics_visualizer.py.

Betweenness Centrality Implementation

Betweenness centrality identifies bridge nodes by counting how often a node appears on shortest paths between all pairs of nodes.

Algorithm Logic:

  • For every node pair, finds all shortest paths
  • Increments counters for each intermediate node lying on those paths
  • Normalizes final scores by the total number of possible node pairs

Implementation Details:

  • NetworkX path: Uses nx.betweenness_centrality with optimized graph algorithms
  • Fallback path: Runs BFS from each source node using _bfs_shortest_paths() to collect all shortest paths and aggregate counts
  • Source location: Lines 68-86 in semantica/kg/centrality_calculator.py (corrected from the documented range)

The calculate_betweenness_centrality() method handles both weighted and unweighted graph structures through the same interface.

Closeness Centrality Implementation

Closeness centrality ranks nodes by their average distance to all other reachable nodes, computing the reciprocal of the average shortest-path length.

Algorithm Logic:

  • Computes average shortest-path distance from a node to all reachable nodes
  • Returns the reciprocal of that average as the closeness score
  • Handles disconnected components by considering only reachable nodes

Implementation Details:

  • NetworkX path: Delegates to nx.closeness_centrality
  • Fallback path: Executes _bfs_distances() from each node to collect distances, then calculates reachable / total_distance
  • Source location: Lines 52-65 in semantica/kg/centrality_calculator.py

Practical Code Examples

Calculating Degree, Betweenness, and Closeness

from semantica.kg.centrality_calculator import CentralityCalculator

# Initialize the calculator

calculator = CentralityCalculator()

# Assume graph follows Semantica's dict format or NetworkX graph

graph = {
    "entities": [{"id": "A"}, {"id": "B"}, {"id": "C"}],
    "relationships": [{"source": "A", "target": "B"}, {"source": "B", "target": "C"}]
}

# 1. Degree centrality

deg_result = calculator.calculate_degree_centrality(graph)
print("Top node by degree:", deg_result["rankings"][0])

# 2. Betweenness centrality

bet_result = calculator.calculate_betweenness_centrality(graph)
print("Highest betweenness:", bet_result["rankings"][0])

# 3. Closeness centrality

close_result = calculator.calculate_closeness_centrality(graph)
print("Most central (closeness):", close_result["rankings"][0])

Accessing PageRank via AlgorithmRegistry

from semantica.kg.registry import algorithm_registry
from semantica.kg.centrality_calculator import CentralityCalculator

calculator = CentralityCalculator()

# Retrieve PageRank metadata from registry

pagerank_cls = algorithm_registry.get("centrality", "pagerank")

# Calculate PageRank with custom parameters

pagerank = calculator.calculate_pagerank(
    graph,
    node_labels=None,          # Consider all node types

    relationship_types=None,   # Consider all edge types

    max_iterations=30,
    damping_factor=0.85,
)
print("PageRank of node 'A':", pagerank["centrality"].get("A"))

Visualizing Centrality Rankings

from semantica.visualization.analytics_visualizer import visualize_centrality_rankings

# Visualize degree centrality results

visualize_centrality_rankings(
    deg_result,
    title="Degree Centrality Rankings",
    output="interactive",    # Renders HTML widget

)

Summary

  • The CentralityCalculator class in semantica/kg/centrality_calculator.py provides unified access to degree, betweenness, and closeness centrality algorithms.
  • The implementation automatically leverages NetworkX for performance-critical calculations, falling back to pure-Python BFS algorithms when unavailable.
  • Degree centrality counts normalized connections, betweenness measures path interception, and closeness calculates reciprocal average distances.
  • The AlgorithmRegistry in semantica/kg/registry.py enables extensibility and lazy loading of algorithm implementations like PageRank.
  • Built-in progress tracking and logging support long-running graph analyses without custom instrumentation.
  • Results integrate directly with semantica/visualization/analytics_visualizer.py for immediate interactive or static visualization.

Frequently Asked Questions

Does Semantica require NetworkX to calculate centrality algorithms?

No, NetworkX is an optional dependency. The CentralityCalculator automatically detects NetworkX availability and delegates to nx.degree_centrality, nx.betweenness_centrality, and nx.closeness_centrality when present. Without NetworkX, it falls back to pure-Python implementations using _build_adjacency() and _bfs_shortest_paths() methods, ensuring functionality across all environments.

How does the AlgorithmRegistry work with centrality calculations?

The AlgorithmRegistry in semantica/kg/registry.py serves as a metadata catalog for algorithm categories including centrality. It registers built-in algorithms like PageRank (lines 58-70) and allows retrieval via algorithm_registry.get("centrality", "pagerank"). The CentralityCalculator uses this registry for lazy loading and extensibility, enabling developers to swap or extend algorithm implementations without modifying the core calculator class.

Can I filter centrality calculations by specific node or relationship types?

Yes, the calculate_pagerank() method accepts node_labels and relationship_types parameters to filter the graph before computation. While degree, betweenness, and closeness methods typically operate on the full graph structure, you can pre-filter your graph dictionary before passing it to these methods to achieve similar subset analysis.

What visualization options exist for centrality results?

The calculator outputs dictionaries compatible with visualize_centrality() and visualize_centrality_rankings() from semantica/visualization/analytics_visualizer.py. These functions accept the result dictionaries directly and support both interactive output (HTML widgets) and static formats, automatically extracting the "rankings" and "centrality" keys from the calculator output.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →