How Named Entity Recognition (NER) Is Implemented in Microsoft's AI-for-Beginners Course
Microsoft's AI-for-Beginners repository implements NER using a TensorFlow-based token-classification pipeline with bidirectional LSTM layers, demonstrated in the notebook NER-TF.ipynb.
The AI-for-Beginners course teaches named entity recognition as a sequence-labeling problem using a complete, runnable notebook that processes raw text into BIO-tagged entities. This implementation in lessons/5-NLP/19-NER/NER-TF.ipynb follows classical deep learning patterns for production-grade NER systems—embedding layers for token representation, bidirectional LSTMs for context capture, and time-distributed dense layers for per-token classification.
NER Dataset Preparation and BIO Tagging
The pipeline begins with Kaggle's ner_dataset.csv, a flat file where each row contains a word and its entity tag. The BIO tagging scheme marks the beginning (B-), inside (I-), or outside (O) of named entities.
Load and inspect the data:
import pandas as pd
df = pd.read_csv('ner_dataset.csv', encoding='unicode-escape')
df.head()
Building Tag and Word Vocabularies
The notebook creates bidirectional lookup dictionaries for both tags and words. For tags, it extracts unique BIO labels and maps them to integers:
tags = df.Tag.unique()
id2tag = dict(enumerate(tags))
tag2id = {v: k for k, v in id2tag.items()}
For words, it lowercases tokens, adds an <UNK> token for out-of-vocabulary handling, and builds word2id/id2word mappings:
vocab = set(df['Word'].apply(lambda w: w.lower()))
id2word = {i+1: w for i, w in enumerate(vocab)}
id2word[0] = '<UNK>'
word2id = {w: i for i, w in id2word.items()}
Sentence Reconstruction from Flat Data
The CSV format stores one word per row with "Sentence #" markers indicating new sentences. The code reconstructs per-sentence word and tag lists by detecting NaN values in the Sentence # column:
X, Y = [], []
sentence_words, sentence_tags = [], []
for _, row in df[['Sentence #', 'Word', 'Tag']].iterrows():
if pd.isna(row['Sentence #']):
sentence_words.append(row['Word'])
sentence_tags.append(row['Tag'])
else:
if sentence_words:
X.append(sentence_words)
Y.append(sentence_tags)
sentence_words, sentence_tags = [row['Word']], [row['Tag']]
X.append(sentence_words)
Y.append(sentence_tags)
This produces aligned lists X (sentences as word lists) and Y (corresponding tag sequences).
Vectorization and Padding for Neural Network Input
Two helper functions convert text to model-ready tensors. The vectorize function maps words to vocabulary IDs; tagify converts tags to integer labels:
def vectorize(seq):
return [word2id[w.lower()] for w in seq]
def tagify(seq):
return [tag2id[t] for t in seq]
Xv = list(map(vectorize, X))
Yv = list(map(tagify, Y))
Keras pads all sequences to uniform length using post-padding:
from keras.preprocessing.sequence import pad_sequences
X_data = pad_sequences(Xv, padding='post')
Y_data = pad_sequences(Yv, padding='post')
NER Model Architecture: Embeddings and Bidirectional LSTMs
The named entity recognition model uses a three-layer Sequential architecture defined in lessons/5-NLP/19-NER/NER-TF.ipynb (lines 48-54 of the source). This design implements the standard token-classification pattern:
from tensorflow import keras
maxlen = X_data.shape[1]
vocab_size = len(vocab) + 1 # +1 for <UNK>
num_tags = len(tags)
model = keras.models.Sequential([
# 300-dimensional token embeddings
keras.layers.Embedding(vocab_size, 300, input_length=maxlen),
# First bidirectional LSTM layer with full sequence output
keras.layers.Bidirectional(
keras.layers.LSTM(100, activation='tanh', return_sequences=True)),
# Second bidirectional LSTM for deeper context
keras.layers.Bidirectional(
keras.layers.LSTM(100, activation='tanh', return_sequences=True)),
# Per-token classification head
keras.layers.TimeDistributed(
keras.layers.Dense(num_tags, activation='softmax'))
])
The Embedding layer (300 dimensions) converts sparse token IDs to dense vectors. Two stacked Bidirectional LSTM layers (100 units each) process context from both directions. The TimeDistributed Dense layer applies softmax independently to each token position, producing per-token probability distributions over all BIO tags.
The model compiles with sparse categorical cross-entropy loss and the Adam optimizer:
model.compile(
loss='sparse_categorical_crossentropy',
optimizer='adam',
metrics=['acc']
)
Training and Evaluation
The training loop uses the full padded dataset with minimal epochs for demonstration purposes. The notebook achieves >98% token-level accuracy even with a single epoch:
model.fit(X_data, Y_data, epochs=1)
This direct fit approach assumes the dataset fits in memory—appropriate for educational purposes and small-to-medium NER corpora.
Inference: From Raw Text to Named Entities
After training, the pipeline converts new sentences into entity predictions. The inference steps match the training preprocessing exactly: tokenization, vocabulary lookup, padding, and prediction:
import numpy as np
sentence = ["Patient", "reports", "headache", "in", "the", "morning"]
# Vectorize and pad to model input length
seq = keras.preprocessing.sequence.pad_sequences(
[vectorize(sentence)], maxlen=maxlen, padding='post'
)
# Predict per-token tag probabilities
pred = model.predict(seq)[0]
# Convert probabilities to tag IDs, then to BIO labels
pred_tags = [id2tag[np.argmax(p)] for p in pred[:len(sentence)]]
# Display word-tag pairs
print(list(zip(sentence, pred_tags)))
Output shows each token paired with its predicted BIO tag, identifying entity boundaries and types.
Key Files in the AI-for-Beginners NER Implementation
| Component | Path | Purpose |
|---|---|---|
| Main notebook | lessons/5-NLP/19-NER/NER-TF.ipynb |
Complete TensorFlow implementation with data loading, model definition, training, and inference |
| Lesson documentation | lessons/5-NLP/19-NER/README.md |
Conceptual overview of BIO tagging and NER fundamentals |
| English translation | translations/en/lessons/5-NLP/19-NER/README.md |
Localized explanatory material |
| Dataset | ner_dataset.csv (Kaggle) |
Annotated training corpus with word-level BIO tags |
Summary
Microsoft's AI-for-Beginners repository teaches named entity recognition through a complete, reproducible pipeline:
- Data handling: Flat CSV ingestion with sentence reconstruction and BIO tag indexing
- Preprocessing: Vocabulary building with
<UNK>handling, word/tag vectorization, and Keras sequence padding - Model architecture: 300-dim embeddings → 2× bidirectional LSTM layers → time-distributed softmax classifier
- Training: Sparse categorical cross-entropy with Adam optimizer, achieving >98% accuracy
- Inference: Identical preprocessing pipeline for new sentences with tag ID-to-label decoding
This implementation demonstrates production-grade patterns including bidirectional context modeling and per-token classification—foundational techniques for any NER system.
Frequently Asked Questions
What BIO tagging scheme does the AI-for-Beginners NER use?
The implementation uses the standard BIO scheme where B- prefixes mark the beginning of an entity, I- prefixes mark words inside an entity, and O marks non-entity tokens. This encoding allows the model to distinguish adjacent entities of the same type and identify multi-word entity spans.
Why does the NER model use two bidirectional LSTM layers?
Two stacked Bidirectional LSTM layers capture increasingly abstract contextual representations. The first layer learns local token contexts; the second layer composes these into higher-level patterns. Both use return_sequences=True to preserve per-token outputs for the final classification layer, as required for sequence labeling.
How does the pipeline handle words not seen during training?
The vocabulary construction explicitly reserves index 0 for <UNK> (unknown token). During vectorize, any lowercase word missing from word2id would map to 0—though the provided code assumes in-vocabulary inputs. For robust production systems, you would add explicit unknown-token handling with word2id.get(w.lower(), 0).
Can this NER implementation process sentences of variable length?
Yes. The pad_sequences function standardizes all inputs to maxlen (the longest training sentence). Shorter sentences receive post-padding with zeros. During inference, the model accepts any length up to maxlen; predictions beyond the actual sentence length are discarded using pred[:len(sentence)].
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →