# How the MSConceptGraph Notebook Implements Knowledge Representation with ConceptNet

> Discover how the MSConceptGraph notebook uses ConceptNet for knowledge representation. Map news headlines to concepts and enable symbolic AI categorization. Learn more today!

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-23

---

**The MSConceptGraph notebook demonstrates knowledge representation by mapping news headlines to abstract concepts via ConceptNet's semantic network, enabling symbolic AI categorization through weighted "is-a" relationships.**

The Microsoft `AI-For-Beginners` repository includes a practical lesson on symbolic AI that transforms unstructured text into structured knowledge graphs. The `MSConceptGraph.ipynb` notebook located in `lessons/2-Symbolic/` illustrates how to leverage **ConceptNet**—an open semantic network—to replace the deprecated Microsoft Concept Graph API. This implementation demonstrates core knowledge representation techniques by extracting noun phrases from news headlines and mapping them to hierarchical concept nodes using probabilistic "is-a" relationships.

## The Six-Stage Knowledge Representation Pipeline

The notebook implements a complete pipeline for semantic categorization through six distinct stages. Each stage transforms the data from raw text to structured knowledge representations using graph-based reasoning.

- **Data Ingestion** — The pipeline begins by fetching recent news headlines using the **NewsAPI.org** REST endpoint (lines 82-88 of `lessons/2-Symbolic/MSConceptGraph.ipynb`). This requires a personal API key to access live news data.

- **Surface-Form Extraction** — The system employs the `TextBlob` library to identify `noun_phrases` within each headline (lines 81-88). This NLTK-backed extraction captures the main topics and entities mentioned in the news titles.

- **Concept-Graph Query** — Each noun phrase is mapped to higher-level concepts using **ConceptNet's** `IsA` relations. A custom `query` function (lines 43-62) wraps the ConceptNet REST API at `https://api.conceptnet.io/query`, normalizing terms and extracting semantic edges.

- **Weight-Based Filtering** — The pipeline retains only salient parent concepts whose normalized weight exceeds **0.1** (line 10 of the aggregation cell). This threshold ensures only statistically significant conceptual relationships influence the final categorization.

- **Aggregation and Categorization** — A dictionary `w` accumulates headlines under their shared abstract parent concepts (lines 6-14). The resulting clusters group news under categories such as *economy*, *nation*, and *person* (lines 26-48).

- **Presentation** — The final stage prints grouped headlines for human inspection using simple `print` statements (lines 94-96), demonstrating the explainable nature of symbolic AI reasoning.

## Querying ConceptNet for Semantic Relationships

The core knowledge representation logic resides in the `query` function defined at lines 43-62 of the notebook. This function normalizes lexical items and retrieves weighted semantic relationships from ConceptNet's graph structure.

```python
import urllib, json

def http(x):
    response = urllib.request.urlopen(x)
    return response.read().decode('utf-8')

def query(noun):
    concept = noun.lower().replace(' ', '_')
    url = f"https://api.conceptnet.io/query?start=/c/en/{urllib.parse.quote(concept)}&rel=/r/IsA&limit=10"
    try:
        result = json.loads(http(url))
    except Exception:
        return {}
    edges = result.get('edges', [])
    if not edges: return {}
    total = sum(e['weight'] for e in edges)
    if total == 0: return {}
    # Normalised weight per parent concept

    return {e['end']['label']: e['weight'] / total for e in edges}

```

The function performs three critical operations. First, it **normalizes** the input noun by converting to lowercase and replacing spaces with underscores to match ConceptNet's URI scheme. Second, it queries the `/r/IsA` relation type to retrieve hierarchical parent concepts. Third, it **normalizes edge weights** into a probability distribution by dividing individual edge weights by the total weight sum, enabling probabilistic reasoning over the concept space.

## Extracting and Processing News Headlines

Before semantic mapping can occur, the notebook ingests live news data and extracts surface forms. The implementation uses the NewsAPI service to pull top headlines from multiple countries.

```python
newsapi_key = '<your API key here>'

def get_news(country='us'):
    url = f"https://newsapi.org/v2/top-headlines?country={country}&apiKey={newsapi_key}"
    res = json.loads(http(url))
    return res['articles']

all_titles = [a['title'] for a in get_news('us') + get_news('gb')]

```

Once headlines are collected, the **TextBlob** library processes each title to extract `noun_phrases`. This step identifies the key entities and topics that serve as input nodes for the concept graph traversal, bridging the gap between raw text and structured semantic nodes.

## Weight-Based Concept Filtering and Aggregation

The aggregation logic applies a threshold filter to ensure only meaningful conceptual relationships survive. The notebook implements this through a dictionary comprehension that checks the normalized weight property.

```python
from textblob import TextBlob

w = {}
for title in all_titles:
    for noun in TextBlob(title).noun_phrases:
        # Get parent concepts from ConceptNet

        terms = query(noun)
        # Keep only strong parents (weight > 0.1)

        for term in [c for c, wt in terms.items() if wt > 0.1]:
            w.setdefault(term, []).append(title)

```

This filtering mechanism (line 10's `wt > 0.1` condition) eliminates weak semantic associations that could introduce noise into the knowledge representation. The resulting clusters demonstrate **explainable AI**—each grouping traces directly to explicit edges in the ConceptNet graph, making the reasoning process transparent and auditable.

The final presentation displays headlines organized by their abstract categories:

```python
print('\nECONOMY:\n' + '\n'.join(w['economy']))
print('\nNATION:\n' + '\n'.join(w['nation']))
print('\nPERSON:\n' + '\n'.join(w['person']))

```

## Why ConceptNet Replaces the Microsoft Concept Graph

The original **Microsoft Concept Graph** API is no longer publicly available. According to the notebook's opening markdown (lines 9-13) and the lesson README ([`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md), lines 13-21), the implementation substitutes **ConceptNet** as a drop-in replacement.

ConceptNet provides analogous `IsA` relationships across a multilingual knowledge graph containing millions of nodes and edges. This substitution maintains the pedagogical goal of teaching semantic network-based knowledge representation while ensuring the code remains runnable without access to proprietary Microsoft services. The open-source nature of ConceptNet also provides transparency into the underlying graph structure and edge weighting mechanisms.

## Summary

The MSConceptGraph notebook demonstrates practical knowledge representation through a symbolic AI pipeline that transforms unstructured news text into hierarchical concept clusters.

- **Semantic Network Architecture** — The implementation leverages ConceptNet's "is-a" relations to lift lexical items from raw text into abstract conceptual nodes.
- **Probabilistic Weighting** — Edge weights from ConceptNet are normalized to create probability distributions, enabling the selection of the most statistically likely parent concepts.
- **Explainable Categorization** — The 0.1 weight threshold and explicit graph traversal make the clustering decisions transparent and traceable to specific semantic relationships.
- **Practical Implementation** — Located in `lessons/2-Symbolic/MSConceptGraph.ipynb`, the notebook combines NewsAPI data ingestion, TextBlob NLP processing, and RESTful graph queries into a complete knowledge representation system.

## Frequently Asked Questions

### How does the MSConceptGraph notebook handle API authentication for data sources?

The notebook requires two API keys to function: a **NewsAPI.org** key for retrieving headlines and implicit access to the public **ConceptNet** REST endpoint. The NewsAPI key must be inserted into the `newsapi_key` variable at line 82 before execution. ConceptNet requires no authentication for its query endpoint at `https://api.conceptnet.io/query`.

### What determines which parent concepts the notebook retains for categorization?

The notebook applies a **normalized weight threshold of 0.1** to filter ConceptNet results. After calculating the sum of all edge weights returned for a noun phrase, each parent concept's weight is divided by this total to produce a probability score. Only concepts exceeding the 0.1 threshold (10% probability) are retained for final aggregation, ensuring that weak or tangential semantic relationships do not clutter the results.

### Why does the notebook use ConceptNet instead of the original Microsoft Concept Graph service?

The **Microsoft Concept Graph API has been discontinued** and is no longer publicly available. According to the repository documentation in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md), the notebook substitutes **ConceptNet**—an open-source, multilingual semantic network that provides equivalent `IsA` hierarchical relationships. This substitution preserves the educational value of the lesson while ensuring the code remains executable without proprietary API access.

### How does the query function normalize terms for the ConceptNet API?

The `query` function defined at lines 43-62 performs **lexical normalization** by converting input nouns to lowercase and replacing whitespace with underscores (`noun.lower().replace(' ', '_')`). This transformation matches ConceptNet's URI encoding scheme for English concepts (`/c/en/{concept}`), ensuring that multi-word phrases like "artificial intelligence" become `artificial_intelligence` for proper API querying.