How to Perform Named Entity Recognition (NER) Using TensorFlow in AI-For-Beginners

The AI-For-Beginners repository provides a complete TensorFlow 2.x implementation of Named Entity Recognition in lessons/5-NLP/19-NER/NER-TF.ipynb, utilizing a Bidirectional LSTM architecture with TimeDistributed dense layers to classify tokens into entity categories.

The Microsoft AI-For-Beginners curriculum offers a practical entry point into natural language processing through hands-on TensorFlow implementations. The official NER tutorial demonstrates how to build an end-to-end sequence labeling model using the Keras API, processing the Kaggle Annotated Corpus to identify persons, organizations, and locations. This guide breaks down the architectural patterns and preprocessing steps defined in the source notebook to help you implement robust NER pipelines.

Understanding the NER Pipeline Architecture

The lessons/5-NLP/19-NER/NER-TF.ipynb notebook implements a classic sequence-to-sequence tagging approach where each input token receives a classification label. The pipeline handles variable-length sentences through padding and masking, ensuring efficient batch processing while preserving bidirectional semantic context.

Dataset Acquisition and Preprocessing

The workflow begins with the Annotated Corpus for Named Entity Recognition from Kaggle, referenced as ner_dataset.csv within the notebook. The implementation uses pandas to load the CSV, extracting parallel lists of tokens and their corresponding IOB-style entity tags (e.g., B-PER, I-ORG, O).

Label encoding converts string tags into integer indices using scikit-learn's LabelEncoder, a necessary step for TensorFlow's categorical loss functions. The sequences are then padded to uniform length using tf.keras.preprocessing.sequence.pad_sequences with padding="post" to maintain consistent tensor shapes across batches.

Tokenization and Embedding Strategy

Text vectorization relies on Keras Tokenizer (tf.keras.preprocessing.text.Tokenizer) configured with oov_token="<OOV>" to handle out-of-vocabulary words during inference. This tokenizer generates integer word indices that feed into an Embedding layer initialized with mask_zero=True, ensuring the model ignores padding tokens during gradient computation.

The embedding layer projects sparse token indices into 128-dimensional dense vectors, providing distributional representations that capture semantic relationships required for entity disambiguation.

Bidirectional LSTM Encoder

The core feature extractor is a Bidirectional LSTM (tf.keras.layers.Bidirectional(tf.keras.layers.LSTM(...))) configured with return_sequences=True. By processing sequences in both forward and backward directions and concatenating the hidden states, this architecture captures contextual dependencies from both preceding and following tokens—critical for identifying entity boundaries that depend on future context.

Time-Distributed Classification

Rather than producing a single output per sequence, the model employs a TimeDistributed Dense layer with softmax activation applied at every time step. This generates independent probability distributions over entity tags for each token position, enabling fine-grained, per-token predictions during the inference phase.

Implementing the TensorFlow NER Model

The following implementation reflects the exact patterns found in lessons/5-NLP/19-NER/NER-TF.ipynb, covering data preparation through model inference.

Data Loading and Preparation

import pandas as pd
import tensorflow as tf
from tensorflow.keras.preprocessing.text import Tokenizer
from tensorflow.keras.preprocessing.sequence import pad_sequences
from sklearn.preprocessing import LabelEncoder

# Load the Kaggle NER dataset

df = pd.read_csv("ner_dataset.csv")
sentences = df["Sentence"].astype(str).tolist()
labels = df["Tag"].astype(str).tolist()

# Tokenize text sequences

tokenizer = Tokenizer(oov_token="<OOV>")
tokenizer.fit_on_texts(sentences)
X = tokenizer.texts_to_sequences(sentences)
X = pad_sequences(X, padding="post")

# Encode entity labels as integers

label_encoder = LabelEncoder()
y = label_encoder.fit_transform(labels)
y = pad_sequences(y, padding="post")

Model Architecture Definition

from tensorflow.keras import Model
from tensorflow.keras.layers import Input, Embedding, Bidirectional, LSTM, TimeDistributed, Dense

vocab_size = len(tokenizer.word_index) + 1
num_tags = len(label_encoder.classes_)
max_len = X.shape[1]

# Construct the Bi-LSTM tagger

input_layer = Input(shape=(max_len,), dtype="int32")
embedding = Embedding(input_dim=vocab_size, output_dim=128, mask_zero=True)(input_layer)
bilstm = Bidirectional(LSTM(64, return_sequences=True))(embedding)
output = TimeDistributed(Dense(num_tags, activation="softmax"))(bilstm)

model = Model(input_layer, output)
model.compile(
    optimizer="adam",
    loss="sparse_categorical_crossentropy",
    metrics=["accuracy"]
)
model.summary()

Training and Inference

import numpy as np

# Reshape labels for sparse_categorical_crossentropy

y_train = np.expand_dims(y, axis=-1)

# Train the sequence tagger

model.fit(X, y_train, batch_size=32, epochs=8, validation_split=0.1)

# Perform inference on new text

test_sentence = ["Google", "was", "founded", "in", "California"]
test_seq = tokenizer.texts_to_sequences([test_sentence])
test_seq = pad_sequences(test_seq, maxlen=max_len, padding="post")

predictions = model.predict(test_seq)
predicted_indices = np.argmax(predictions[0], axis=-1)
predicted_tags = label_encoder.inverse_transform(predicted_indices[:len(test_sentence)])

print(list(zip(test_sentence, predicted_tags)))

Key Files and Resources

  • lessons/5-NLP/19-NER/NER-TF.ipynb – The primary TensorFlow implementation notebook containing data preprocessing, Bi-LSTM model definition, training configuration using sparse_categorical_crossentropy, and inference utilities.

  • ner_dataset.csv (Kaggle) – External dataset imported within the notebook, providing token-level IOB annotations for training and evaluation.

  • translations/ – Localized versions of the NER notebook maintain identical TensorFlow code implementations while offering region-specific instructional text.

Summary

  • The AI-For-Beginners NER tutorial utilizes a Bidirectional LSTM sequence tagger implemented in lessons/5-NLP/19-NER/NER-TF.ipynb to classify individual tokens into entity categories.
  • Keras Tokenizer and scikit-learn LabelEncoder bridge the gap between raw text/strings and integer indices required by TensorFlow operations.
  • The TimeDistributed Dense layer with softmax enables per-token predictions across variable-length sequences, optimized using sparse_categorical_crossentropy loss for integer-encoded labels.
  • The pipeline processes the Kaggle Annotated Corpus through pandas, with pad_sequences ensuring uniform input dimensions for efficient batched training on both CPU and GPU.

Frequently Asked Questions

What dataset does the AI-For-Beginners NER tutorial use?

The tutorial utilizes the Annotated Corpus for Named Entity Recognition available on Kaggle, referenced as ner_dataset.csv. This dataset contains sentences annotated with IOB format tags marking entities such as persons, organizations, and locations.

Why does the TensorFlow NER model use a Bidirectional LSTM?

The Bidirectional LSTM processes sequences in both forward and backward directions, allowing the model to incorporate future context when predicting the current token's entity tag. This bidirectional context is essential for NER because entity boundaries often depend on words that appear after the entity itself.

How are entity labels encoded in the NER-TF.ipynb notebook?

The notebook uses scikit-learn's LabelEncoder to convert string entity tags (e.g., B-PER, I-LOC) into integer indices. The encoded labels are padded to match sequence lengths and reshaped to support TensorFlow's sparse_categorical_crossentropy loss function during training.

Can I replace the random embeddings with pre-trained vectors like GloVe or Word2Vec?

Yes, the Embedding layer accepts pre-trained weight matrices through the weights parameter. You can load GloVe or Word2Vec vectors into a matrix indexed by the Tokenizer's word index and set trainable=False to freeze the embeddings or allow fine-tuning during NER training.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →