How to Set Up Named Entity Recognition (NER) Pipelines Using the NER-TF.ipynb Lesson

The NER-TF.ipynb notebook in the microsoft/AI-For-Beginners repository provides a complete TensorFlow workflow for building named entity recognition pipelines using bidirectional LSTMs, CoNLL-2003 data formatting, and entity-level evaluation metrics.

The microsoft/AI-For-Beginners curriculum offers a production-ready implementation for constructing named entity recognition (NER) pipelines from scratch. Located at lessons/5-NLP/19-NER/NER-TF.ipynb, this lesson demonstrates how to process CoNLL-formatted datasets, implement bidirectional recurrent architectures, and evaluate models using the seqeval library for precise entity-level scoring.

Preparing the CoNLL-2003 Dataset

The pipeline begins by loading data in CoNLL-2003 format, which structures text as tab-separated tokens and their corresponding IOB (Inside-Outside-Beginning) tags.

First, split your dataset into training and validation sets to ensure robust model evaluation:

import pandas as pd
from sklearn.model_selection import train_test_split

data = pd.read_csv('data/conll2003.tsv', sep='\t')
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)

The lessons/5-NLP/19-NER/data/conll2003.tsv file in the repository provides sample data in the correct format for testing the pipeline immediately.

Tokenizing Text and Encoding Tags

Convert raw sentences into numerical sequences using tf.keras.preprocessing.text.Tokenizer. Set lower=False to preserve case sensitivity, which often carries semantic meaning in entity recognition (e.g., distinguishing between "apple" the fruit and "Apple" the company).

from tensorflow.keras.preprocessing.text import Tokenizer

tokenizer = Tokenizer(lower=False, oov_token='<OOV>')
tokenizer.fit_on_texts(train_df['sentence'])
train_seq = tokenizer.texts_to_sequences(train_df['sentence'])
val_seq   = tokenizer.texts_to_sequences(test_df['sentence'])

Map entity tags to integer indices for model compatibility:

tag2idx = {t: i for i, t in enumerate(sorted(set(data['tag'])))}
idx2tag = {i: t for t, i in tag2idx.items()}

Padding Sequences for Batch Processing

Neural networks require uniform input dimensions. Use tf.keras.preprocessing.sequence.pad_sequences to standardize sequence lengths, applying post-padding to preserve the initial tokens where entity information typically appears.

from tensorflow.keras.preprocessing.sequence import pad_sequences

MAX_LEN = 128
X_train = pad_sequences(train_seq, maxlen=MAX_LEN, padding='post')
X_val   = pad_sequences(val_seq,   maxlen=MAX_LEN, padding='post')

y_train = pad_sequences(
    train_df['tag'].apply(lambda tags: [tag2idx[t] for t in tags]), 
    maxlen=MAX_LEN, padding='post'
)
y_val = pad_sequences(
    test_df['tag'].apply(lambda tags: [tag2idx[t] for t in tags]), 
    maxlen=MAX_LEN, padding='post'
)

Building the TensorFlow NER Architecture

The model architecture follows a standard Bi-LSTM design optimized for sequence tagging tasks.

Embedding Layer with Masking

Initialize an Embedding layer that handles variable-length sequences by setting mask_zero=True. This tells the model to ignore padded positions during computation:

import tensorflow as tf

embedding_layer = tf.keras.layers.Embedding(
    input_dim=len(tokenizer.word_index) + 1,
    output_dim=200,
    mask_zero=True
)

Bidirectional LSTM for Context Awareness

Stack a Bidirectional wrapper around an LSTM layer to capture both past and future context for each token. Setting return_sequences=True ensures the layer outputs predictions for every position in the sequence:

bilstm = tf.keras.layers.Bidirectional(
    tf.keras.layers.LSTM(128, return_sequences=True)
)

TimeDistributed Output Layer

Apply a dense layer with softmax activation to each time step independently using TimeDistributed. This produces probability distributions over all possible entity tags for every token:

model = tf.keras.Sequential([
    embedding_layer,
    bilstm,
    tf.keras.layers.TimeDistributed(
        tf.keras.layers.Dense(len(tag2idx), activation='softmax')
    )
])

Compile the model with sparse_categorical_crossentropy loss and accuracy metrics:

model.compile(
    optimizer='adam',
    loss='sparse_categorical_crossentropy',
    metrics=['accuracy']
)

Training with Early Stopping and Checkpointing

Configure callbacks to prevent overfitting and preserve the best model weights. The EarlyStopping callback monitors validation loss with a patience of 2 epochs, while ModelCheckpoint saves the optimal state to disk.

callbacks = [
    tf.keras.callbacks.EarlyStopping(patience=2, restore_best_weights=True),
    tf.keras.callbacks.ModelCheckpoint('ner_tf_best.h5', save_best_only=True)
]

model.fit(
    X_train, y_train,
    validation_data=(X_val, y_val),
    epochs=10,
    batch_size=32,
    callbacks=callbacks
)

Evaluating Entity-Level Performance with seqeval

Token-level accuracy can be misleading in NER tasks. Install and import the seqeval library (listed in the repository's requirements.txt) to compute precision, recall, and F1-score at the entity level, properly handling multi-token entities defined by IOB tags.

from seqeval.metrics import classification_report

# seqeval expects lists of lists containing string labels

print(classification_report(y_true_tags, y_pred_tags))

Decoding Predictions and Extracting Entities

After training, convert model outputs back to human-readable entity labels. Use the idx2tag dictionary to map predicted indices to tag strings, then merge B-I sequences into complete entity spans:

pred = model.predict(X_val)
pred_tags = pred.argmax(axis=-1)

decoded = [[idx2tag[idx] for idx in seq] for seq in pred_tags]

Post-processing logic can then merge consecutive B-PER and I-PER tags into single "Person" entity spans for downstream applications.

Summary

  • Data Format: The pipeline expects CoNLL-2003 formatted data available at lessons/5-NLP/19-NER/data/conll2003.tsv.
  • Preprocessing: Use tf.keras.preprocessing.text.Tokenizer for vocabulary creation and pad_sequences for uniform input dimensions.
  • Architecture: A Bidirectional(LSTM) layer with TimeDistributed(Dense) output provides strong baseline performance for named entity recognition tasks.
  • Training: Implement EarlyStopping and ModelCheckpoint callbacks to optimize training efficiency and model selection.
  • Evaluation: Leverage the seqeval library for entity-level metrics that account for multi-token entity boundaries.

Frequently Asked Questions

What hardware requirements are needed to run the NER-TF.ipynb notebook?

The notebook runs on standard CPU configurations, though training accelerates significantly with GPU support. The repository's environment.yml file specifies TensorFlow versions compatible with CUDA for GPU acceleration. For educational purposes, the small dataset size in conll2003.tsv allows complete training cycles on CPU within minutes.

How does the mask_zero parameter in the Embedding layer improve NER performance?

Setting mask_zero=True in tf.keras.layers.Embedding instructs the model to ignore padded positions (zeros) during backpropagation. This prevents the network from learning patterns from artificial padding tokens, ensuring that loss calculations and gradient updates only occur on actual text tokens, which improves convergence on variable-length sentences.

Can I replace the LSTM with a GRU or Transformer architecture in this pipeline?

Yes. According to the source code in NER-TF.ipynb, you can substitute tf.keras.layers.LSTM with tf.keras.layers.GRU within the Bidirectional wrapper for faster training with minimal accuracy trade-off. For Transformer architectures, you would replace the recurrent layers with tf.keras.layers.MultiHeadAttention and positional encoding, though this requires modifying the padding and masking logic to handle attention mechanisms.

Why does the pipeline use sparse_categorical_crossentropy instead of categorical_crossentropy?

The sparse_categorical_crossentropy loss function expects integer labels rather than one-hot encoded vectors, which reduces memory usage when dealing with large tag vocabularies (PER, ORG, LOC, etc.). Since the CoNLL data provides integer-encoded tags via the tag2idx mapping, this loss function eliminates the need to convert labels to one-hot format, streamlining the data pipeline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →