Source Formats Supported by Semantica's Knowledge Graph Builder: The Complete 2024 Guide
Semantica's knowledge graph builder supports four distinct source formats—CSV, JSON, database, and API—each routed through specialized ingestion routines in the SeedManager class.
The semantica-agi/semantica repository provides a flexible data ingestion pipeline that converts diverse data sources into knowledge graphs. Understanding these source formats and their specific configuration requirements is essential for building robust graph structures from your existing data assets.
The Four Core Source Formats
Semantica processes data through a source object containing a mandatory format field. The SeedManager dispatches each format to dedicated ingestor modules based on this field value.
CSV Files
The CSV format triggers the CSV ingestion routine, configurable with delimiters and file paths. In semantica/seed/seed_manager.py lines 591–594, the builder validates the format and routes to the file ingestor handler.
csv_source = {
"format": "csv",
"uri": "file:///data/my_data.csv",
"options": {"delimiter": ","}
}
kg_csv = SeedManager.build_graph(csv_source)
JSON Documents
JSON sources support both local files and remote URLs, parsed through the JSON ingestion routine at lines 598–602. This format handles nested structures and converts them into graph nodes and edges.
json_source = {
"format": "json",
"uri": "https://example.com/data.json"
}
kg_json = SeedManager.build_graph(json_source)
Database Connections
The database format accommodates SQL dialects including SQLite, PostgreSQL, and DuckDB. Lines 605–609 in seed_manager.py route these to semantica/ingest/db_ingestor.py, requiring connection URIs and query strings.
db_source = {
"format": "database",
"dialect": "sqlite",
"uri": "sqlite:///mydb.sqlite",
"query": "SELECT * FROM entities"
}
kg_db = SeedManager.build_graph(db_source)
REST API Endpoints
API sources enable real-time graph construction from REST endpoints. The ingestion routine at lines 612–616 supports configurable HTTP methods, headers, and authentication tokens via semantica/ingest/web_ingestor.py.
api_source = {
"format": "api",
"endpoint": "https://api.example.com/v1/graph",
"method": "GET",
"headers": {"Authorization": "Bearer <token>"}
}
kg_api = SeedManager.build_graph(api_source)
How Format Routing Works in SeedManager
The SeedManager class serves as the central dispatcher for the knowledge graph builder. When SeedManager.build_graph() receives a source object, it inspects the format field and delegates to specialized handlers:
- CSV and JSON route through
semantica/ingest/file_ingestor.py - Database connections utilize
semantica/ingest/db_ingestor.py - API calls execute through
semantica/ingest/web_ingestor.py
If the provided format does not match csv, json, database, or api, the builder raises a ProcessingError at line 620, halting execution and indicating the unsupported source format.
Configuration Requirements by Format
Each source format requires specific mandatory fields in the source object:
| Format | Required Fields | Optional Fields |
|---|---|---|
| CSV | format, uri |
options (delimiter, encoding) |
| JSON | format, uri |
None |
| Database | format, dialect, uri, query |
Connection pool settings |
| API | format, endpoint, method |
headers, body, timeout |
Summary
- Semantica's knowledge graph builder processes four source formats: CSV, JSON, database, and API.
- The
SeedManagerclass insemantica/seed/seed_manager.pyroutes formats to specialized ingestors based on lines 591–616. - CSV and JSON files ingest through the file ingestor module, while database connections use the database ingestor.
- API sources require endpoint configuration and support custom headers for authentication.
- Invalid formats trigger a
ProcessingErrorat line 620 ofseed_manager.py.
Frequently Asked Questions
What happens if I specify an unsupported format in the source object?
The knowledge graph builder raises a ProcessingError indicating the unsupported source format. This validation occurs at line 620 in semantica/seed/seed_manager.py, preventing execution with incompatible data types.
Can I ingest multiple source formats into a single knowledge graph?
While the raw analysis shows individual SeedManager.build_graph() calls for each format, you would invoke the builder separately for each source type and merge the resulting graph objects, as each format requires distinct ingestion routines and configuration schemas.
Does the database format support cloud data warehouses like Snowflake or BigQuery?
The source code references generic database ingestion at lines 605–609 in seed_manager.py and the db_ingestor.py file. The specific dialect support depends on the underlying database driver implementation in the ingestor module, though the architecture supports extensible SQL dialects.
How do I configure authentication for API sources?
Pass authentication tokens or API keys through the headers dictionary in the source object. The web ingestor accepts standard HTTP header configurations, allowing Bearer tokens, API keys, or custom authentication schemes required by your endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →