How to Perform Image Processing with OpenCV in Microsoft AI for Beginners

The Microsoft AI for Beginners curriculum teaches fundamental computer vision through hands-on OpenCV image processing lessons covering color space conversion, geometric transformations, and motion analysis in the lessons/4-ComputerVision/06-IntroCV/ directory.

The microsoft/AI-For-Beginners repository introduces computer vision fundamentals using OpenCV—the industry-standard C++ library with Python bindings. Learners progress from basic image loading to advanced motion detection through interactive Jupyter notebooks and structured assignments.

Getting Started with OpenCV in the Curriculum

The introductory computer vision lesson is located at lessons/4-ComputerVision/06-IntroCV/ and consists of three core components:

  • README.md – Conceptual overview, installation notes, and workflow guidance
  • OpenCV.ipynb – Executable notebook with runnable code examples
  • lab/README.md – Practical assignment requiring optical flow implementation

All dependencies are managed through the root environment.yml, which includes OpenCV via Conda. The curriculum supports local execution, VS Code dev containers, Binder, or GitHub Codespaces.

Loading Images and Managing Color Spaces

OpenCV loads images in BGR (Blue-Green-Red) order by default, which differs from Matplotlib's expected RGB format. According to the lesson documentation in lessons/4-ComputerVision/06-IntroCV/README.md, you must convert color spaces before displaying images with standard Python visualization libraries.

import cv2
import matplotlib.pyplot as plt

# Load returns a NumPy array in BGR order

img_bgr = cv2.imread('image.jpeg')

# Convert to RGB for correct display

img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB)

plt.imshow(img_rgb)
plt.title('Original RGB')
plt.show()

Essential Pre-processing Operations

Before feeding images to neural networks, the curriculum demonstrates four fundamental transformations using OpenCV functions.

Resizing and Interpolation

Use cv2.resize with specific interpolation algorithms to standardize input dimensions. The lesson recommends INTER_LANCZOS for high-quality downscaling.

img_resized = cv2.resize(img_rgb, (320, 200), interpolation=cv2.INTER_LANCZOS)

Noise Reduction

OpenCV provides multiple blurring techniques to reduce sensor noise:

  • Median Blur – cv2.medianBlur effective for salt-and-pepper noise
  • Gaussian Blur – cv2.GaussianBlur for general smoothing
img_blurred = cv2.GaussianBlur(img_resized, (5, 5), 0)

Brightness and Contrast Adjustment

The curriculum teaches direct array manipulation using cv2.convertScaleAbs, which applies the formula output = alpha * input + beta where alpha controls contrast and beta controls brightness.

import numpy as np

alpha = 1.3  # 30% more contrast

beta = 20    # Increase brightness

img_adjusted = cv2.convertScaleAbs(img_blurred, alpha=alpha, beta=beta)

Thresholding for Segmentation

Binary segmentation uses cv2.threshold for global thresholds or cv2.adaptiveThreshold for varying lighting conditions.

gray = cv2.cvtColor(img_adjusted, cv2.COLOR_RGB2GRAY)
_, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)

Geometric Transformations

The lesson covers two matrix-based transformation types in lessons/4-ComputerVision/06-IntroCV/README.md:

Affine Transformations (cv2.warpAffine) – Preserve parallel lines, enabling rotation, scaling, and translation using 2×3 matrices.

Perspective Transformations (cv2.warpPerspective) – Handle 3D viewpoint changes using 3×3 homography matrices for document rectification or viewpoint correction.


# Define three point correspondences for affine transform

pts_src = np.float32([[0, 0], [200, 0], [0, 200]])
pts_dst = np.float32([[20, 30], [180, 20], [30, 210]])

# Calculate and apply transformation matrix

M = cv2.getAffineTransform(pts_src, pts_dst)
img_affine = cv2.warpAffine(thresh, M, (200, 200))

Motion Analysis with Optical Flow

The curriculum extends into video processing with optical flow techniques for motion detection, distinguishing between:

  • Dense Optical Flow – Calculates motion vectors for every pixel using cv2.calcOpticalFlowFarneback
  • Sparse Optical Flow – Tracks specific feature points (mentioned in lab assignments)

The implementation in OpenCV.ipynb demonstrates Farneback's algorithm to compute motion between video frames:

cap = cv2.VideoCapture('video.mp4')
ret, prev = cap.read()
prev_gray = cv2.cvtColor(prev, cv2.COLOR_BGR2GRAY)

while True:
    ret, frame = cap.read()
    if not ret:
        break
    
    gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
    
    # Calculate dense optical flow

    flow = cv2.calcOpticalFlowFarneBack(
        prev_gray, gray, None, 
        0.5, 3, 15, 3, 5, 1.2, 0
    )
    
    # Visualize flow magnitude and direction

    mag, ang = cv2.cartToPolar(flow[..., 0], flow[..., 1])
    
    cap.release()

Running the Code Examples

The OpenCV.ipynb notebook in lessons/4-ComputerVision/06-IntroCV/ contains all executable examples referenced above. To run locally:

  1. Clone the microsoft/AI-For-Beginners repository
  2. Install the Conda environment: conda env create -f environment.yml
  3. Activate and launch Jupyter: conda activate ai4beg && jupyter notebook
  4. Navigate to lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb

For cloud execution, the repository supports one-click deployment to Binder or GitHub Codespaces with pre-installed OpenCV dependencies.

Summary

  • OpenCV loads images in BGR format requiring cv2.cvtColor conversion for RGB display libraries
  • Pre-processing pipelines use cv2.resize, cv2.GaussianBlur, and cv2.convertScaleAbs for network-ready inputs
  • Binary segmentation applies cv2.threshold or adaptive variants for foreground extraction
  • Geometric corrections utilize cv2.warpAffine and cv2.warpPerspective with transformation matrices
  • Motion analysis leverages cv2.calcOpticalFlowFarneback for dense optical flow in video sequences
  • All implementations are provided in the executable OpenCV.ipynb notebook within the AI for Beginners curriculum

Frequently Asked Questions

How do I install OpenCV for the AI for Beginners curriculum?

The repository uses Conda for dependency management. OpenCV is included in the root environment.yml file. Run conda env create -f environment.yml to install all required packages including OpenCV, NumPy, and Matplotlib. Alternatively, the curriculum runs in GitHub Codespaces or Binder without local installation.

Why does my OpenCV image look blue when displayed with Matplotlib?

OpenCV stores images in BGR (Blue-Green-Red) order while Matplotlib expects RGB. According to the lesson at lessons/4-ComputerVision/06-IntroCV/README.md, you must convert using cv2.cvtColor(img, cv2.COLOR_BGR2RGB) before displaying with plt.imshow().

What is the difference between affine and perspective transformations in OpenCV?

Affine transformations (cv2.warpAffine) preserve parallel lines and require three point pairs to create a 2×3 matrix, suitable for rotation and scaling. Perspective transformations (cv2.warpPerspective) use four point pairs to create a 3×3 matrix, correcting for 3D viewpoint changes like document scanning or camera angle adjustments.

Which optical flow method does the AI for Beginners curriculum recommend?

The lesson demonstrates dense optical flow using cv2.calcOpticalFlowFarneback for computing motion vectors across entire frames. The accompanying lab assignment in lessons/4-ComputerVision/06-IntroCV/lab/README.md challenges students to extract specific motion directions from these flow calculations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →