How the MSConceptGraph Notebook Implements Knowledge Representation with ConceptNet
The MSConceptGraph notebook demonstrates knowledge representation by mapping news headlines to abstract concepts via ConceptNet's semantic network, enabling symbolic AI categorization through weighted "is-a" relationships.
The Microsoft AI-For-Beginners repository includes a practical lesson on symbolic AI that transforms unstructured text into structured knowledge graphs. The MSConceptGraph.ipynb notebook located in lessons/2-Symbolic/ illustrates how to leverage ConceptNet—an open semantic network—to replace the deprecated Microsoft Concept Graph API. This implementation demonstrates core knowledge representation techniques by extracting noun phrases from news headlines and mapping them to hierarchical concept nodes using probabilistic "is-a" relationships.
The Six-Stage Knowledge Representation Pipeline
The notebook implements a complete pipeline for semantic categorization through six distinct stages. Each stage transforms the data from raw text to structured knowledge representations using graph-based reasoning.
-
Data Ingestion — The pipeline begins by fetching recent news headlines using the NewsAPI.org REST endpoint (lines 82-88 of
lessons/2-Symbolic/MSConceptGraph.ipynb). This requires a personal API key to access live news data. -
Surface-Form Extraction — The system employs the
TextBloblibrary to identifynoun_phraseswithin each headline (lines 81-88). This NLTK-backed extraction captures the main topics and entities mentioned in the news titles. -
Concept-Graph Query — Each noun phrase is mapped to higher-level concepts using ConceptNet's
IsArelations. A customqueryfunction (lines 43-62) wraps the ConceptNet REST API athttps://api.conceptnet.io/query, normalizing terms and extracting semantic edges. -
Weight-Based Filtering — The pipeline retains only salient parent concepts whose normalized weight exceeds 0.1 (line 10 of the aggregation cell). This threshold ensures only statistically significant conceptual relationships influence the final categorization.
-
Aggregation and Categorization — A dictionary
waccumulates headlines under their shared abstract parent concepts (lines 6-14). The resulting clusters group news under categories such as economy, nation, and person (lines 26-48). -
Presentation — The final stage prints grouped headlines for human inspection using simple
printstatements (lines 94-96), demonstrating the explainable nature of symbolic AI reasoning.
Querying ConceptNet for Semantic Relationships
The core knowledge representation logic resides in the query function defined at lines 43-62 of the notebook. This function normalizes lexical items and retrieves weighted semantic relationships from ConceptNet's graph structure.
import urllib, json
def http(x):
response = urllib.request.urlopen(x)
return response.read().decode('utf-8')
def query(noun):
concept = noun.lower().replace(' ', '_')
url = f"https://api.conceptnet.io/query?start=/c/en/{urllib.parse.quote(concept)}&rel=/r/IsA&limit=10"
try:
result = json.loads(http(url))
except Exception:
return {}
edges = result.get('edges', [])
if not edges: return {}
total = sum(e['weight'] for e in edges)
if total == 0: return {}
# Normalised weight per parent concept
return {e['end']['label']: e['weight'] / total for e in edges}
The function performs three critical operations. First, it normalizes the input noun by converting to lowercase and replacing spaces with underscores to match ConceptNet's URI scheme. Second, it queries the /r/IsA relation type to retrieve hierarchical parent concepts. Third, it normalizes edge weights into a probability distribution by dividing individual edge weights by the total weight sum, enabling probabilistic reasoning over the concept space.
Extracting and Processing News Headlines
Before semantic mapping can occur, the notebook ingests live news data and extracts surface forms. The implementation uses the NewsAPI service to pull top headlines from multiple countries.
newsapi_key = '<your API key here>'
def get_news(country='us'):
url = f"https://newsapi.org/v2/top-headlines?country={country}&apiKey={newsapi_key}"
res = json.loads(http(url))
return res['articles']
all_titles = [a['title'] for a in get_news('us') + get_news('gb')]
Once headlines are collected, the TextBlob library processes each title to extract noun_phrases. This step identifies the key entities and topics that serve as input nodes for the concept graph traversal, bridging the gap between raw text and structured semantic nodes.
Weight-Based Concept Filtering and Aggregation
The aggregation logic applies a threshold filter to ensure only meaningful conceptual relationships survive. The notebook implements this through a dictionary comprehension that checks the normalized weight property.
from textblob import TextBlob
w = {}
for title in all_titles:
for noun in TextBlob(title).noun_phrases:
# Get parent concepts from ConceptNet
terms = query(noun)
# Keep only strong parents (weight > 0.1)
for term in [c for c, wt in terms.items() if wt > 0.1]:
w.setdefault(term, []).append(title)
This filtering mechanism (line 10's wt > 0.1 condition) eliminates weak semantic associations that could introduce noise into the knowledge representation. The resulting clusters demonstrate explainable AI—each grouping traces directly to explicit edges in the ConceptNet graph, making the reasoning process transparent and auditable.
The final presentation displays headlines organized by their abstract categories:
print('\nECONOMY:\n' + '\n'.join(w['economy']))
print('\nNATION:\n' + '\n'.join(w['nation']))
print('\nPERSON:\n' + '\n'.join(w['person']))
Why ConceptNet Replaces the Microsoft Concept Graph
The original Microsoft Concept Graph API is no longer publicly available. According to the notebook's opening markdown (lines 9-13) and the lesson README (lessons/2-Symbolic/README.md, lines 13-21), the implementation substitutes ConceptNet as a drop-in replacement.
ConceptNet provides analogous IsA relationships across a multilingual knowledge graph containing millions of nodes and edges. This substitution maintains the pedagogical goal of teaching semantic network-based knowledge representation while ensuring the code remains runnable without access to proprietary Microsoft services. The open-source nature of ConceptNet also provides transparency into the underlying graph structure and edge weighting mechanisms.
Summary
The MSConceptGraph notebook demonstrates practical knowledge representation through a symbolic AI pipeline that transforms unstructured news text into hierarchical concept clusters.
- Semantic Network Architecture — The implementation leverages ConceptNet's "is-a" relations to lift lexical items from raw text into abstract conceptual nodes.
- Probabilistic Weighting — Edge weights from ConceptNet are normalized to create probability distributions, enabling the selection of the most statistically likely parent concepts.
- Explainable Categorization — The 0.1 weight threshold and explicit graph traversal make the clustering decisions transparent and traceable to specific semantic relationships.
- Practical Implementation — Located in
lessons/2-Symbolic/MSConceptGraph.ipynb, the notebook combines NewsAPI data ingestion, TextBlob NLP processing, and RESTful graph queries into a complete knowledge representation system.
Frequently Asked Questions
How does the MSConceptGraph notebook handle API authentication for data sources?
The notebook requires two API keys to function: a NewsAPI.org key for retrieving headlines and implicit access to the public ConceptNet REST endpoint. The NewsAPI key must be inserted into the newsapi_key variable at line 82 before execution. ConceptNet requires no authentication for its query endpoint at https://api.conceptnet.io/query.
What determines which parent concepts the notebook retains for categorization?
The notebook applies a normalized weight threshold of 0.1 to filter ConceptNet results. After calculating the sum of all edge weights returned for a noun phrase, each parent concept's weight is divided by this total to produce a probability score. Only concepts exceeding the 0.1 threshold (10% probability) are retained for final aggregation, ensuring that weak or tangential semantic relationships do not clutter the results.
Why does the notebook use ConceptNet instead of the original Microsoft Concept Graph service?
The Microsoft Concept Graph API has been discontinued and is no longer publicly available. According to the repository documentation in lessons/2-Symbolic/README.md, the notebook substitutes ConceptNet—an open-source, multilingual semantic network that provides equivalent IsA hierarchical relationships. This substitution preserves the educational value of the lesson while ensuring the code remains executable without proprietary API access.
How does the query function normalize terms for the ConceptNet API?
The query function defined at lines 43-62 performs lexical normalization by converting input nouns to lowercase and replacing whitespace with underscores (noun.lower().replace(' ', '_')). This transformation matches ConceptNet's URI encoding scheme for English concepts (/c/en/{concept}), ensuring that multi-word phrases like "artificial intelligence" become artificial_intelligence for proper API querying.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →