# How Named Entity Recognition (NER) Is Implemented in Microsoft's AI-for-Beginners Course

> Learn how Microsoft's AI-for-Beginners course implements Named Entity Recognition NER using TensorFlow and LSTMs. Explore the NER-TF notebook for a practical guide.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: deep-dive
- Published: 2026-08-25

---

**Microsoft's AI-for-Beginners repository implements NER using a TensorFlow-based token-classification pipeline with bidirectional LSTM layers, demonstrated in the notebook `NER-TF.ipynb`.**

The AI-for-Beginners course teaches **named entity recognition** as a sequence-labeling problem using a complete, runnable notebook that processes raw text into BIO-tagged entities. This implementation in `lessons/5-NLP/19-NER/NER-TF.ipynb` follows classical deep learning patterns for production-grade NER systems—embedding layers for token representation, bidirectional LSTMs for context capture, and time-distributed dense layers for per-token classification.

## NER Dataset Preparation and BIO Tagging

The pipeline begins with Kaggle's `ner_dataset.csv`, a flat file where each row contains a word and its entity tag. The **BIO tagging scheme** marks the beginning (B-), inside (I-), or outside (O) of named entities.

Load and inspect the data:

```python
import pandas as pd

df = pd.read_csv('ner_dataset.csv', encoding='unicode-escape')
df.head()

```

### Building Tag and Word Vocabularies

The notebook creates bidirectional lookup dictionaries for both tags and words. For tags, it extracts unique BIO labels and maps them to integers:

```python
tags = df.Tag.unique()
id2tag = dict(enumerate(tags))
tag2id = {v: k for k, v in id2tag.items()}

```

For words, it lowercases tokens, adds an `<UNK>` token for out-of-vocabulary handling, and builds `word2id`/`id2word` mappings:

```python
vocab = set(df['Word'].apply(lambda w: w.lower()))
id2word = {i+1: w for i, w in enumerate(vocab)}
id2word[0] = '<UNK>'
word2id = {w: i for i, w in id2word.items()}

```

## Sentence Reconstruction from Flat Data

The CSV format stores one word per row with "Sentence #" markers indicating new sentences. The code reconstructs per-sentence word and tag lists by detecting `NaN` values in the Sentence # column:

```python
X, Y = [], []
sentence_words, sentence_tags = [], []

for _, row in df[['Sentence #', 'Word', 'Tag']].iterrows():
    if pd.isna(row['Sentence #']):
        sentence_words.append(row['Word'])
        sentence_tags.append(row['Tag'])
    else:
        if sentence_words:
            X.append(sentence_words)
            Y.append(sentence_tags)
        sentence_words, sentence_tags = [row['Word']], [row['Tag']]
X.append(sentence_words)
Y.append(sentence_tags)

```

This produces aligned lists `X` (sentences as word lists) and `Y` (corresponding tag sequences).

## Vectorization and Padding for Neural Network Input

Two helper functions convert text to model-ready tensors. The `vectorize` function maps words to vocabulary IDs; `tagify` converts tags to integer labels:

```python
def vectorize(seq):
    return [word2id[w.lower()] for w in seq]

def tagify(seq):
    return [tag2id[t] for t in seq]

Xv = list(map(vectorize, X))
Yv = list(map(tagify, Y))

```

Keras pads all sequences to uniform length using post-padding:

```python
from keras.preprocessing.sequence import pad_sequences

X_data = pad_sequences(Xv, padding='post')
Y_data = pad_sequences(Yv, padding='post')

```

## NER Model Architecture: Embeddings and Bidirectional LSTMs

The **named entity recognition model** uses a three-layer Sequential architecture defined in `lessons/5-NLP/19-NER/NER-TF.ipynb` (lines 48-54 of the source). This design implements the standard token-classification pattern:

```python
from tensorflow import keras

maxlen = X_data.shape[1]
vocab_size = len(vocab) + 1  # +1 for <UNK>

num_tags = len(tags)

model = keras.models.Sequential([
    # 300-dimensional token embeddings

    keras.layers.Embedding(vocab_size, 300, input_length=maxlen),
    
    # First bidirectional LSTM layer with full sequence output

    keras.layers.Bidirectional(
        keras.layers.LSTM(100, activation='tanh', return_sequences=True)),
    
    # Second bidirectional LSTM for deeper context

    keras.layers.Bidirectional(
        keras.layers.LSTM(100, activation='tanh', return_sequences=True)),
    
    # Per-token classification head

    keras.layers.TimeDistributed(
        keras.layers.Dense(num_tags, activation='softmax'))
])

```

The **Embedding layer** (300 dimensions) converts sparse token IDs to dense vectors. Two stacked **Bidirectional LSTM** layers (100 units each) process context from both directions. The **TimeDistributed Dense** layer applies `softmax` independently to each token position, producing per-token probability distributions over all BIO tags.

The model compiles with **sparse categorical cross-entropy** loss and the **Adam optimizer**:

```python
model.compile(
    loss='sparse_categorical_crossentropy',
    optimizer='adam',
    metrics=['acc']
)

```

## Training and Evaluation

The training loop uses the full padded dataset with minimal epochs for demonstration purposes. The notebook achieves **>98% token-level accuracy** even with a single epoch:

```python
model.fit(X_data, Y_data, epochs=1)

```

This direct fit approach assumes the dataset fits in memory—appropriate for educational purposes and small-to-medium NER corpora.

## Inference: From Raw Text to Named Entities

After training, the pipeline converts new sentences into entity predictions. The inference steps match the training preprocessing exactly: tokenization, vocabulary lookup, padding, and prediction:

```python
import numpy as np

sentence = ["Patient", "reports", "headache", "in", "the", "morning"]

# Vectorize and pad to model input length

seq = keras.preprocessing.sequence.pad_sequences(
    [vectorize(sentence)], maxlen=maxlen, padding='post'
)

# Predict per-token tag probabilities

pred = model.predict(seq)[0]

# Convert probabilities to tag IDs, then to BIO labels

pred_tags = [id2tag[np.argmax(p)] for p in pred[:len(sentence)]]

# Display word-tag pairs

print(list(zip(sentence, pred_tags)))

```

Output shows each token paired with its predicted BIO tag, identifying entity boundaries and types.

## Key Files in the AI-for-Beginners NER Implementation

| Component | Path | Purpose |
|-----------|------|---------|
| **Main notebook** | `lessons/5-NLP/19-NER/NER-TF.ipynb` | Complete TensorFlow implementation with data loading, model definition, training, and inference |
| **Lesson documentation** | [`lessons/5-NLP/19-NER/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/lessons/5-NLP/19-NER/README.md) | Conceptual overview of BIO tagging and NER fundamentals |
| **English translation** | [`translations/en/lessons/5-NLP/19-NER/README.md`](https://github.com/microsoft/AI-For-Beginners/blob/main/translations/en/lessons/5-NLP/19-NER/README.md) | Localized explanatory material |
| **Dataset** | `ner_dataset.csv` (Kaggle) | Annotated training corpus with word-level BIO tags |

## Summary

Microsoft's AI-for-Beginners repository teaches **named entity recognition** through a complete, reproducible pipeline:

- **Data handling**: Flat CSV ingestion with sentence reconstruction and BIO tag indexing
- **Preprocessing**: Vocabulary building with `<UNK>` handling, word/tag vectorization, and Keras sequence padding
- **Model architecture**: 300-dim embeddings → 2× bidirectional LSTM layers → time-distributed softmax classifier
- **Training**: Sparse categorical cross-entropy with Adam optimizer, achieving >98% accuracy
- **Inference**: Identical preprocessing pipeline for new sentences with tag ID-to-label decoding

This implementation demonstrates production-grade patterns including bidirectional context modeling and per-token classification—foundational techniques for any NER system.

## Frequently Asked Questions

### What BIO tagging scheme does the AI-for-Beginners NER use?

The implementation uses the **standard BIO scheme** where `B-` prefixes mark the beginning of an entity, `I-` prefixes mark words inside an entity, and `O` marks non-entity tokens. This encoding allows the model to distinguish adjacent entities of the same type and identify multi-word entity spans.

### Why does the NER model use two bidirectional LSTM layers?

Two stacked **Bidirectional LSTM** layers capture increasingly abstract contextual representations. The first layer learns local token contexts; the second layer composes these into higher-level patterns. Both use `return_sequences=True` to preserve per-token outputs for the final classification layer, as required for sequence labeling.

### How does the pipeline handle words not seen during training?

The vocabulary construction explicitly reserves index 0 for `<UNK>` (unknown token). During `vectorize`, any lowercase word missing from `word2id` would map to 0—though the provided code assumes in-vocabulary inputs. For robust production systems, you would add explicit unknown-token handling with `word2id.get(w.lower(), 0)`.

### Can this NER implementation process sentences of variable length?

Yes. The `pad_sequences` function standardizes all inputs to `maxlen` (the longest training sentence). Shorter sentences receive post-padding with zeros. During inference, the model accepts any length up to `maxlen`; predictions beyond the actual sentence length are discarded using `pred[:len(sentence)]`.