# Basic Concepts of Symbolic AI: Knowledge Representation and Reasoning in Microsoft's AI-For-Beginners

> Explore core Symbolic AI concepts like semantic networks, frames, and rules. Understand knowledge representation and reasoning in Microsoft's AI-For-Beginners course.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-29

---

**Symbolic AI represents knowledge through explicit structures like semantic networks, frames, and production rules, then uses inference engines to reason over that knowledge—contrasting sharply with data-driven machine learning approaches.**

The `microsoft/AI-For-Beginners` repository dedicates its second lesson to the foundational ideas of symbolic artificial intelligence, illustrating how machines can encode and manipulate knowledge in a human-readable form. Unlike neural networks that discover patterns from training data, symbolic AI relies on manually encoded rules and structured representations. This article examines the basic concepts of symbolic AI covered in the curriculum, tracing how knowledge moves from raw data to actionable decisions through explicit representation and logical inference.

## Knowledge vs. Information: The DIKW Pyramid

Before building AI systems, the curriculum establishes a hierarchy of understanding. As defined in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L18-L27](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L18-L27)), the **DIKW pyramid** distinguishes four levels:

- **Data**: Raw, unprocessed symbols and signals
- **Information**: Data given context and meaning
- **Knowledge**: A mental model or framework derived from information
- **Wisdom**: Knowledge applied with judgment and ethics

Symbolic AI focuses on the upper layers—explicitly encoding **knowledge** rather than processing raw data. This distinction explains why symbolic systems require domain expertise to build: humans must translate their understanding into machine-readable structures.

## The Knowledge Representation Spectrum

In [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L33-L41](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L33-L41)), the curriculum presents a continuum of representation methods. At one extreme, **algorithms** encode knowledge procedurally—rigid and non-flexible. At the other extreme, **natural language** is expressive but computationally opaque. Practical symbolic AI occupies the middle ground, using structured formats that balance human readability with machine computability.

## Semantic Networks and OAV Triplets

The most intuitive representation method models knowledge as a graph. **Semantic networks** store facts as **object-attribute-value (OAV) triplets**, where relationships form edges between nodes.

As illustrated in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L50-L57](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L50-L57)), programming languages can be represented as triplets:

```python

# Define a list of (object, attribute, value) triplets

triplets = [
    ("Python", "is", "Untyped-Language"),
    ("Python", "invented-by", "Guido van Rossum"),
    ("Python", "block-syntax", "indentation"),
    ("Untyped-Language", "doesn't have", "type definitions"),
]

# Query function: retrieve all values for a given object-attribute pair

def query(subject, predicate):
    return [obj for s, p, obj in triplets if s == subject and p == predicate]

print("Python is:", query("Python", "is"))
print("Python invented by:", query("Python", "invented-by"))

```

This graph structure enables intuitive traversal: asking "What is Python?" retrieves connected attributes by following edges from the "Python" node.

## Hierarchical Representations with Frames

For complex domains requiring inheritance, **frames** organize knowledge into hierarchical structures. Defined in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L61-L65](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L61-L65)), frames mimic object-oriented programming—each frame contains **slots** (attributes) with default values, constraints, or attached procedures.

In a biological taxonomy:
- The `Animal` frame defines slots for `respiration` and `habitat`
- The `Bird` frame inherits from `Animal` but adds `wing-type` and `flight-capable`
- The `Canary` frame inherits from `Bird` with specific values like `color: yellow`

This **inheritance mechanism** eliminates redundant data storage while maintaining consistency across the knowledge base.

## Procedural Knowledge via Production Rules

When knowledge takes the form of conditional actions, **production rules** encode it as **if-then** statements. As described in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L76-L78](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L76-L78)), these rules trigger when their antecedent conditions match working memory facts.

## Logical Representations

For formal verification, predicate logic (and subsets like Horn clauses) provides mathematical rigor. The curriculum notes that **description logics**—lighter-weight variants—power the Semantic Web by defining formal ontologies with decidable reasoning properties.

## Forward and Backward Inference

Reasoning mechanisms determine how symbolic AI derives new facts. The [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) file ([L19-L40](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L19-L40)) details two strategies:

**Forward chaining** starts with known facts and applies rules until reaching a goal. The following Python snippet implements a minimal forward-chaining engine for animal classification:

```python

# Working memory holds currently known OAV facts

working_memory = {
    ("animal", "eats", "meat"),
    ("animal", "has", "sharp teeth"),
    ("animal", "has", "claws"),
    ("animal", "has", "forward-looking eyes")
}

# Production rule for carnivores (simplified)

def carnivore_rule(facts):
    required = {("animal", "eats", "meat"),
                ("animal", "has", "sharp teeth"),
                ("animal", "has", "claws"),
                ("animal", "has", "forward-looking eyes")}
    return required.issubset(facts)

# Forward inference loop

if carnivore_rule(working_memory):
    working_memory.add(("animal", "is", "carnivore"))
    print("Inference result:", ("animal", "is", "carnivore"))
else:
    print("Not enough evidence")

```

**Backward chaining** works in reverse: starting from a hypothesis (e.g., "Is this animal a carnivore?"), the system seeks rules that could prove it, recursively verifying premises until reaching known facts or failing. The `Animals.ipynb` notebook demonstrates both approaches in the animal classification domain.

## Expert Systems Architecture

Expert systems combine the previous concepts into functional applications. According to [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L84-L99](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L84-L99)), these systems contain three core components:

1. **Knowledge base**: Static domain knowledge encoded as rules and facts
2. **Inference engine**: Dynamic reasoning component applying forward or backward chaining
3. **Working memory**: Temporary storage for problem-specific facts

This architecture mimics human experts by isolating domain knowledge (which changes slowly) from reasoning procedures (which remain constant across domains).

## Ontologies and the Semantic Web

Moving beyond simple hierarchies, **ontologies** formally specify domain concepts using RDF/OWL triples. As explained in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L57-L67](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L57-L67)), these specifications enable distributed reasoning across knowledge bases like WikiData and DBpedia.

### The Microsoft Concept Graph

The curriculum also covers automatically extracted knowledge through the **Microsoft Concept Graph** ([L13-L21](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L13-L21)). This large-scale ontology mines "is-a" relationships from web text, containing millions of entities. The `MSConceptGraph.ipynb` notebook demonstrates using this resource for tasks like news article clustering, bridging hand-crafted and automatically extracted knowledge.

## Querying Ontologies in Practice

The `FamilyOntology.ipynb` notebook implements RDF-style querying over genealogical data. The following pattern demonstrates ontology traversal:

```python

# Minimal RDF-style triples for a family ontology

family_triplets = [
    ("Alice", "type", "Person"),
    ("Alice", "hasParent", "Bob"),
    ("Bob", "type", "Person"),
    ("Bob", "hasSibling", "Carol"),
]

def get_relations(subject, predicate):
    return [obj for s, p, obj in family_triplets if s == subject and p == predicate]

print("Alice's parent:", get_relations("Alice", "hasParent"))
print("Bob's sibling:", get_relations("Bob", "hasSibling"))

```

This approach powers the Semantic Web, where SPARQL queries traverse billions of triples across distributed databases.

## Summary

- **Symbolic AI** encodes knowledge explicitly through structures like OAV triplets, frames, and production rules rather than learning from data.
- **Semantic networks** represent facts as graph edges, enabling intuitive querying of object-attribute-value relationships as shown in [`lessons/2-Symbolic/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md) ([L50-L57](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/2-Symbolic/README.md#L50-L57)).
- **Frames** provide hierarchical organization with inheritance, reducing redundancy while maintaining consistency across class hierarchies.
- **Inference engines** apply either forward chaining (data-driven) or backward chaining (goal-driven) to derive conclusions from rule bases.
- **Expert systems** combine static knowledge bases with dynamic inference engines, separating domain expertise from reasoning algorithms.
- **Ontologies** formalize domain knowledge using RDF/OWL standards, enabling machines to reason over distributed knowledge graphs like the Microsoft Concept Graph.

## Frequently Asked Questions

### What is the difference between symbolic AI and machine learning?

Symbolic AI relies on explicit, human-encoded rules and logical representations to manipulate knowledge, while machine learning extracts patterns from training data without explicit programming. The `microsoft/AI-For-Beginners` repository contrasts these approaches in lesson two, showing that symbolic systems are interpretable but require manual knowledge engineering, whereas neural networks learn automatically but operate as "black boxes."

### How do forward and backward chaining differ in expert systems?

**Forward chaining** starts from known facts and applies rules iteratively until reaching a conclusion, making it data-driven and suitable for monitoring systems. **Backward chaining** begins with a hypothesis and works backward to verify supporting facts, making it goal-driven and efficient for diagnostic applications. The `Animals.ipynb` notebook in the repository provides executable implementations of both strategies using animal classification rules.

### What are OAV triplets and why are they important?

**Object-Attribute-Value (OAV) triplets** represent individual facts as three-element tuples (e.g., `("Python", "invented-by", "Guido van Rossum")`). They form the foundation of semantic networks and RDF databases, enabling graph-based querying where relationships are explicit edges between nodes. This format balances human readability with machine processability, serving as the primary data structure in the curriculum's semantic network examples.

### Can symbolic AI systems handle uncertainty?

While the basic concepts in the `AI-For-Beginners` curriculum focus on deterministic logic (where facts are true or false), production rule systems can be extended with **certainty factors** or **fuzzy logic** to handle probabilistic reasoning. However, the fundamental architecture—explicit knowledge representation plus inference engines—remains the same, distinguishing it from statistical machine learning approaches that inherently quantify uncertainty through probability distributions.