How to Use Semantica's Graph Analytics Algorithms: A Complete Guide to Centrality and Community Detection
Semantica's graph analytics algorithms are exposed through two main classes—CentralityCalculator and CommunityDetector—which compute node centrality scores and detect communities via methods like calculate_pagerank() and detect_communities(), with optional NetworkX acceleration and pure-Python fallbacks.
Semantica provides a comprehensive suite of graph analytics utilities for analyzing knowledge graphs, ranging from PageRank centrality to Louvain community detection. These capabilities are implemented in the semantica-agi/semantica repository and exposed both as direct Python APIs and as LLM-accessible tools through the Multi-Channel Processor (MCP). Understanding how to leverage these algorithms enables you to extract structural insights, identify influential nodes, and discover community patterns within your graph data.
Core Components for Graph Analytics
The analytics engine centers on two primary classes defined in the knowledge graph module. Both are designed with optional-dependency awareness, attempting to import NetworkX for optimized performance while gracefully falling back to pure-Python implementations when NetworkX is unavailable.
CentralityCalculator
Located in [semantica/kg/centrality_calculator.py](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/centrality_calculator.py), this class calculates node importance through five distinct metrics. The implementation spans lines 70–670, with each algorithm providing specific fallback logic for environments without NetworkX.
Key methods include:
calculate_degree_centrality()– Normalizes scores by (n-1) and falls back to basic adjacency list traversal (lines 70–119).calculate_betweenness_centrality()– Uses NetworkX'sbetweenness_centralitywhen available; otherwise performs BFS for every source-target pair (lines 85–125).calculate_closeness_centrality()– Computes reciprocal average shortest path length with BFS distance fallbacks (lines 151–165).calculate_eigenvector_centrality()– Attempts NetworkX's eigenvector solver before running power-iteration on a dense adjacency matrix (lines 191–226).calculate_pagerank()– Implements the standard random-walk formula using sparse CSR matrices for efficiency, supporting label and relationship filtering (lines 483–670).
CommunityDetector
Found in [semantica/kg/community_detector.py](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/community_detector.py), this class detects structural groupings through multiple algorithms:
- Louvain – Uses NetworkX's
greedy_modularity_communities(lines 64–115), returning early for empty graphs (lines 76–84). - Leiden – Currently delegates to the Louvain routine while preserving output format consistency (lines 126–136).
- Overlapping – Leverages NetworkX's
k_clique_communitieswith a dense-subgraph fallback (lines 204–226). - Label Propagation – Offers standard in-memory (lines 332–452) and chunked variants for very large graphs (lines 456–522).
All methods utilize shared utilities for adjacency list construction (_build_adjacency), NetworkX conversion (_to_networkx), and modularity calculation (_calculate_modularity).
MCP Tool Integration
The Multi-Channel Processor exposes these analytics through the graph tool definition in [semantica_mcp/mcp/tools/graph.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py). The handle_get_graph_analytics function (lines 77–112) instantiates these calculators and dispatches requests based on the metrics parameter, enabling LLM agents to request centrality scores and community detection without direct Python coding.
Computing Node Centrality
Centrality analysis identifies influential nodes within your knowledge graph. The CentralityCalculator class provides both individual metric calculations and batch processing capabilities.
Supported Centrality Algorithms
Each algorithm returns a standardized dictionary containing a centrality mapping (node → score) and a rankings list ordered by descending score:
- Degree Centrality – Measures direct connections, normalized by the maximum possible edges.
- Betweenness Centrality – Quantifies bridge nodes that lie on shortest paths between other nodes.
- Closeness Centrality – Calculates how quickly a node can reach all other nodes in the graph.
- Eigenvector Centrality – Weights connections by the importance of neighboring nodes.
- PageRank – Computes the probability distribution of a random walk with damping.
Advanced PageRank Configuration
The calculate_pagerank() method offers sophisticated filtering through private helpers _filter_nodes_by_labels and _get_filtered_neighbors. You can constrain the random walk to specific node types and relationship categories:
from semantica.kg import CentralityCalculator
calc = CentralityCalculator()
result = calc.calculate_pagerank(
graph=my_graph,
node_labels=["Person", "Organization"],
relationship_types=["RELATED_TO"],
damping_factor=0.9,
max_iterations=30
)
This filters the graph to only traverse "Person" and "Organization" nodes via "RELATED_TO" edges before computing rankings.
Detecting Graph Communities
Community detection reveals natural clusters within your knowledge graph, helping identify functional groups or organizational boundaries.
Louvain and Leiden Methods
The detect_communities() method defaults to the Louvain algorithm, which maximizes modularity through greedy optimization. According to the source code in community_detector.py, this uses NetworkX's greedy_modularity_communities when available (lines 64–115). The Leiden algorithm (lines 126–136) currently delegates to Louvain while maintaining API compatibility for future optimization.
Results include:
communities– List of node sets representing each cluster.node_assignments– Dictionary mapping nodes to community IDs.modularity– Quality score measuring the density of intra-community edges versus random expectation.
Overlapping and Label Propagation
For graphs where nodes belong to multiple communities, the overlapping algorithm uses k-clique detection (lines 204–226). The label propagation method offers both standard and chunked processing:
- Standard mode (lines 332–452) keeps the entire graph in memory for rapid convergence.
- Chunked mode (lines 456–522) processes
chunk_sizenodes at a time, enabling analysis of graphs that exceed available RAM.
Architecture and Implementation Details
Understanding the internal mechanics ensures optimal deployment across diverse environments.
NetworkX Acceleration with Pure-Python Fallbacks
Both calculator classes implement defensive import patterns:
try:
import networkx as nx
HAS_NETWORKX = True
except ImportError:
HAS_NETWORKX = False
When NetworkX is present, algorithms leverage optimized C-backed implementations. In restricted environments, pure-Python fallbacks ensure functionality:
- Betweenness falls back to all-pairs BFS (lines 85–125).
- Eigenvector uses power-iteration on dense matrices (lines 191–226).
- Closeness implements BFS distance accumulation (lines 151–165).
Result Normalization and Output Format
All centrality methods return standardized structures:
{
"centrality": {"node_id": 0.85, ...},
"rankings": [("node_id", 0.85), ...] # Sorted descending
}
Community methods append node_assignments and quality metrics like modularity. This normalization allows the MCP's handle_get_graph_analytics function to aggregate diverse metrics into a unified LLM response.
Practical Usage Examples
The following examples demonstrate common analytics patterns using the Semantica API. Replace my_graph with a graph object obtained via get_graph() from [semantica_mcp/mcp/session.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/session.py) or constructed directly.
Calculate All Centrality Measures
Compute every centrality metric in a single batch operation:
from semantica.kg import CentralityCalculator
calc = CentralityCalculator()
results = calc.calculate_all_centrality(my_graph)
# Access specific metrics
pagerank_scores = results["pagerank"]["centrality"]
top_degree_nodes = results["degree"]["rankings"][:10]
Filtered PageRank Analysis
Analyze influence within specific semantic contexts:
from semantica.kg import CentralityCalculator
calc = CentralityCalculator()
pr_result = calc.calculate_pagerank(
graph=my_graph,
node_labels=["Person", "Organization"],
relationship_types=["RELATED_TO"],
max_iterations=30,
damping_factor=0.9
)
top_5 = pr_result["rankings"][:5]
print(f"Top influential entities: {top_5}")
Community Detection with Modularity
Detect clusters using the Louvain algorithm and evaluate quality:
from semantica.kg import CommunityDetector
detector = CommunityDetector()
result = detector.detect_communities(
graph=my_graph,
algorithm="louvain",
resolution=1.2
)
print(f"Communities detected: {len(result['communities'])}")
print(f"Modularity score: {result['modularity']:.3f}")
Chunked Processing for Large Graphs
Process massive graphs using label propagation with memory constraints:
from semantica.kg import CommunityDetector
detector = CommunityDetector()
lp_result = detector.detect_communities(
graph=my_graph,
algorithm="label_propagation",
max_iterations=200,
chunk_size=5000 # Process 5k nodes per batch
)
LLM-Driven Analytics via MCP
Invoke analytics through the MCP tool interface for agentic workflows:
{
"name": "get_graph_analytics",
"arguments": {
"metrics": ["pagerank", "degree", "communities"],
"top_n": 10
}
}
The MCP handler instantiates CentralityCalculator and CommunityDetector internally, returning aggregated results including PageRank scores, degree rankings, and Louvain community assignments.
Key Source Files
The following files comprise the graph analytics subsystem according to the semantica-agi/semantica source code:
- [
semantica/kg/centrality_calculator.py](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/centrality_calculator.py) – Implements degree, betweenness, closeness, eigenvector, and PageRank calculations with NetworkX fallbacks. - [
semantica/kg/community_detector.py](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/community_detector.py) – Provides Louvain, Leiden, overlapping, and label-propagation community detection. - [
semantica_mcp/mcp/tools/graph.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py) – Exposes analytics via theget_graph_analyticsMCP tool function. - [
semantica/kg/_graph_view.py](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/_graph_view.py) – Converts internal graph representations to NetworkX-compatible views. - [
semantica_mcp/mcp/session.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/session.py) – Providesget_graph()for retrieving the current knowledge graph instance.
Summary
- Semantica's graph analytics algorithms are implemented in two primary classes:
CentralityCalculatorandCommunityDetector, located insemantica/kg/centrality_calculator.pyandsemantica/kg/community_detector.pyrespectively. - Centrality measures include degree, betweenness, closeness, eigenvector, and PageRank, with
calculate_pagerank()supporting advanced filtering by node labels and relationship types. - Community detection supports Louvain, Leiden, overlapping k-clique, and label propagation methods, with chunked processing available for memory-constrained environments.
- NetworkX integration provides accelerated performance, while pure-Python fallbacks ensure functionality in restricted deployment environments.
- MCP integration exposes these capabilities to LLM agents through the
get_graph_analyticstool, enabling conversational analytics workflows.
Frequently Asked Questions
What centrality measures does Semantica support?
Semantica supports five centrality measures through the CentralityCalculator class: degree centrality (normalized by n-1), betweenness centrality (identifying bridge nodes), closeness centrality (measuring reachability), eigenvector centrality (weighting by neighbor importance), and PageRank (random-walk probability). Each method is implemented in semantica/kg/centrality_calculator.py with specific fallback logic for environments without NetworkX.
How does Semantica handle large graphs that don't fit in memory?
For memory-constrained scenarios, the CommunityDetector class provides a chunked label propagation algorithm that processes graphs in configurable batches (default 5,000 nodes) rather than loading the entire structure into memory simultaneously. This is implemented in lines 456–522 of semantica/kg/community_detector.py via the chunk_size parameter in detect_communities().
Can I filter nodes and relationships when computing analytics?
Yes, the calculate_pagerank() method specifically supports filtering through the node_labels and relationship_types parameters. Internally, this uses private helpers _filter_nodes_by_labels and _get_filtered_neighbors (lines 483–670 of centrality_calculator.py) to construct a subgraph before running the random-walk calculation, allowing targeted analysis of specific entity types.
What is the difference between Louvain and Leiden community detection?
Currently, both algorithms in Semantica's CommunityDetector produce similar results, as the Leiden implementation (lines 126–136) delegates to the Louvain routine using NetworkX's greedy_modularity_communities. The Leiden option preserves API consistency for future optimization, while Louvain (lines 64–115) provides the active greedy modularity maximization with resolution parameter support.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →