# How to Set Up Named Entity Recognition (NER) Pipelines Using the NER-TF.ipynb Lesson

> Learn to set up named entity recognition NER pipelines with the NER-TF.ipynb lesson from Microsoft AI For Beginners. Explore TensorFlow workflows, LSTMs, and data formatting.

- Repository: [Microsoft/AI-For-Beginners](https://github.com/microsoft/AI-For-Beginners)
- Tags: how-to-guide
- Published: 2026-08-27

---

**The NER-TF.ipynb notebook in the microsoft/AI-For-Beginners repository provides a complete TensorFlow workflow for building named entity recognition pipelines using bidirectional LSTMs, CoNLL-2003 data formatting, and entity-level evaluation metrics.**

The microsoft/AI-For-Beginners curriculum offers a production-ready implementation for constructing named entity recognition (NER) pipelines from scratch. Located at `lessons/5-NLP/19-NER/NER-TF.ipynb`, this lesson demonstrates how to process CoNLL-formatted datasets, implement bidirectional recurrent architectures, and evaluate models using the `seqeval` library for precise entity-level scoring.

## Preparing the CoNLL-2003 Dataset

The pipeline begins by loading data in CoNLL-2003 format, which structures text as tab-separated tokens and their corresponding IOB (Inside-Outside-Beginning) tags.

First, split your dataset into training and validation sets to ensure robust model evaluation:

```python
import pandas as pd
from sklearn.model_selection import train_test_split

data = pd.read_csv('data/conll2003.tsv', sep='\t')
train_df, test_df = train_test_split(data, test_size=0.2, random_state=42)

```

The `lessons/5-NLP/19-NER/data/conll2003.tsv` file in the repository provides sample data in the correct format for testing the pipeline immediately.

## Tokenizing Text and Encoding Tags

Convert raw sentences into numerical sequences using `tf.keras.preprocessing.text.Tokenizer`. Set `lower=False` to preserve case sensitivity, which often carries semantic meaning in entity recognition (e.g., distinguishing between "apple" the fruit and "Apple" the company).

```python
from tensorflow.keras.preprocessing.text import Tokenizer

tokenizer = Tokenizer(lower=False, oov_token='<OOV>')
tokenizer.fit_on_texts(train_df['sentence'])
train_seq = tokenizer.texts_to_sequences(train_df['sentence'])
val_seq   = tokenizer.texts_to_sequences(test_df['sentence'])

```

Map entity tags to integer indices for model compatibility:

```python
tag2idx = {t: i for i, t in enumerate(sorted(set(data['tag'])))}
idx2tag = {i: t for t, i in tag2idx.items()}

```

## Padding Sequences for Batch Processing

Neural networks require uniform input dimensions. Use `tf.keras.preprocessing.sequence.pad_sequences` to standardize sequence lengths, applying post-padding to preserve the initial tokens where entity information typically appears.

```python
from tensorflow.keras.preprocessing.sequence import pad_sequences

MAX_LEN = 128
X_train = pad_sequences(train_seq, maxlen=MAX_LEN, padding='post')
X_val   = pad_sequences(val_seq,   maxlen=MAX_LEN, padding='post')

y_train = pad_sequences(
    train_df['tag'].apply(lambda tags: [tag2idx[t] for t in tags]), 
    maxlen=MAX_LEN, padding='post'
)
y_val = pad_sequences(
    test_df['tag'].apply(lambda tags: [tag2idx[t] for t in tags]), 
    maxlen=MAX_LEN, padding='post'
)

```

## Building the TensorFlow NER Architecture

The model architecture follows a standard Bi-LSTM design optimized for sequence tagging tasks.

### Embedding Layer with Masking

Initialize an `Embedding` layer that handles variable-length sequences by setting `mask_zero=True`. This tells the model to ignore padded positions during computation:

```python
import tensorflow as tf

embedding_layer = tf.keras.layers.Embedding(
    input_dim=len(tokenizer.word_index) + 1,
    output_dim=200,
    mask_zero=True
)

```

### Bidirectional LSTM for Context Awareness

Stack a `Bidirectional` wrapper around an `LSTM` layer to capture both past and future context for each token. Setting `return_sequences=True` ensures the layer outputs predictions for every position in the sequence:

```python
bilstm = tf.keras.layers.Bidirectional(
    tf.keras.layers.LSTM(128, return_sequences=True)
)

```

### TimeDistributed Output Layer

Apply a dense layer with softmax activation to each time step independently using `TimeDistributed`. This produces probability distributions over all possible entity tags for every token:

```python
model = tf.keras.Sequential([
    embedding_layer,
    bilstm,
    tf.keras.layers.TimeDistributed(
        tf.keras.layers.Dense(len(tag2idx), activation='softmax')
    )
])

```

Compile the model with `sparse_categorical_crossentropy` loss and accuracy metrics:

```python
model.compile(
    optimizer='adam',
    loss='sparse_categorical_crossentropy',
    metrics=['accuracy']
)

```

## Training with Early Stopping and Checkpointing

Configure callbacks to prevent overfitting and preserve the best model weights. The `EarlyStopping` callback monitors validation loss with a patience of 2 epochs, while `ModelCheckpoint` saves the optimal state to disk.

```python
callbacks = [
    tf.keras.callbacks.EarlyStopping(patience=2, restore_best_weights=True),
    tf.keras.callbacks.ModelCheckpoint('ner_tf_best.h5', save_best_only=True)
]

model.fit(
    X_train, y_train,
    validation_data=(X_val, y_val),
    epochs=10,
    batch_size=32,
    callbacks=callbacks
)

```

## Evaluating Entity-Level Performance with seqeval

Token-level accuracy can be misleading in NER tasks. Install and import the `seqeval` library (listed in the repository's [`requirements.txt`](https://github.com/microsoft/AI-For-Beginners/blob/main/requirements.txt)) to compute precision, recall, and F1-score at the entity level, properly handling multi-token entities defined by IOB tags.

```python
from seqeval.metrics import classification_report

# seqeval expects lists of lists containing string labels

print(classification_report(y_true_tags, y_pred_tags))

```

## Decoding Predictions and Extracting Entities

After training, convert model outputs back to human-readable entity labels. Use the `idx2tag` dictionary to map predicted indices to tag strings, then merge B-I sequences into complete entity spans:

```python
pred = model.predict(X_val)
pred_tags = pred.argmax(axis=-1)

decoded = [[idx2tag[idx] for idx in seq] for seq in pred_tags]

```

Post-processing logic can then merge consecutive `B-PER` and `I-PER` tags into single "Person" entity spans for downstream applications.

## Summary

- **Data Format**: The pipeline expects CoNLL-2003 formatted data available at `lessons/5-NLP/19-NER/data/conll2003.tsv`.
- **Preprocessing**: Use `tf.keras.preprocessing.text.Tokenizer` for vocabulary creation and `pad_sequences` for uniform input dimensions.
- **Architecture**: A `Bidirectional(LSTM)` layer with `TimeDistributed(Dense)` output provides strong baseline performance for named entity recognition tasks.
- **Training**: Implement `EarlyStopping` and `ModelCheckpoint` callbacks to optimize training efficiency and model selection.
- **Evaluation**: Leverage the `seqeval` library for entity-level metrics that account for multi-token entity boundaries.

## Frequently Asked Questions

### What hardware requirements are needed to run the NER-TF.ipynb notebook?

The notebook runs on standard CPU configurations, though training accelerates significantly with GPU support. The repository's [`environment.yml`](https://github.com/microsoft/AI-For-Beginners/blob/main/environment.yml) file specifies TensorFlow versions compatible with CUDA for GPU acceleration. For educational purposes, the small dataset size in `conll2003.tsv` allows complete training cycles on CPU within minutes.

### How does the mask_zero parameter in the Embedding layer improve NER performance?

Setting `mask_zero=True` in `tf.keras.layers.Embedding` instructs the model to ignore padded positions (zeros) during backpropagation. This prevents the network from learning patterns from artificial padding tokens, ensuring that loss calculations and gradient updates only occur on actual text tokens, which improves convergence on variable-length sentences.

### Can I replace the LSTM with a GRU or Transformer architecture in this pipeline?

Yes. According to the source code in `NER-TF.ipynb`, you can substitute `tf.keras.layers.LSTM` with `tf.keras.layers.GRU` within the `Bidirectional` wrapper for faster training with minimal accuracy trade-off. For Transformer architectures, you would replace the recurrent layers with `tf.keras.layers.MultiHeadAttention` and positional encoding, though this requires modifying the padding and masking logic to handle attention mechanisms.

### Why does the pipeline use sparse_categorical_crossentropy instead of categorical_crossentropy?

The `sparse_categorical_crossentropy` loss function expects integer labels rather than one-hot encoded vectors, which reduces memory usage when dealing with large tag vocabularies (PER, ORG, LOC, etc.). Since the CoNLL data provides integer-encoded tags via the `tag2idx` mapping, this loss function eliminates the need to convert labels to one-hot format, streamlining the data pipeline.