How to Perform Image Processing with OpenCV in Microsoft AI for Beginners
The Microsoft AI for Beginners curriculum teaches fundamental computer vision through hands-on OpenCV image processing lessons covering color space conversion, geometric transformations, and motion analysis in the lessons/4-ComputerVision/06-IntroCV/ directory.
The microsoft/AI-For-Beginners repository introduces computer vision fundamentals using OpenCV—the industry-standard C++ library with Python bindings. Learners progress from basic image loading to advanced motion detection through interactive Jupyter notebooks and structured assignments.
Getting Started with OpenCV in the Curriculum
The introductory computer vision lesson is located at lessons/4-ComputerVision/06-IntroCV/ and consists of three core components:
README.md– Conceptual overview, installation notes, and workflow guidanceOpenCV.ipynb– Executable notebook with runnable code exampleslab/README.md– Practical assignment requiring optical flow implementation
All dependencies are managed through the root environment.yml, which includes OpenCV via Conda. The curriculum supports local execution, VS Code dev containers, Binder, or GitHub Codespaces.
Loading Images and Managing Color Spaces
OpenCV loads images in BGR (Blue-Green-Red) order by default, which differs from Matplotlib's expected RGB format. According to the lesson documentation in lessons/4-ComputerVision/06-IntroCV/README.md, you must convert color spaces before displaying images with standard Python visualization libraries.
import cv2
import matplotlib.pyplot as plt
# Load returns a NumPy array in BGR order
img_bgr = cv2.imread('image.jpeg')
# Convert to RGB for correct display
img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB)
plt.imshow(img_rgb)
plt.title('Original RGB')
plt.show()
Essential Pre-processing Operations
Before feeding images to neural networks, the curriculum demonstrates four fundamental transformations using OpenCV functions.
Resizing and Interpolation
Use cv2.resize with specific interpolation algorithms to standardize input dimensions. The lesson recommends INTER_LANCZOS for high-quality downscaling.
img_resized = cv2.resize(img_rgb, (320, 200), interpolation=cv2.INTER_LANCZOS)
Noise Reduction
OpenCV provides multiple blurring techniques to reduce sensor noise:
- Median Blur –
cv2.medianBlureffective for salt-and-pepper noise - Gaussian Blur –
cv2.GaussianBlurfor general smoothing
img_blurred = cv2.GaussianBlur(img_resized, (5, 5), 0)
Brightness and Contrast Adjustment
The curriculum teaches direct array manipulation using cv2.convertScaleAbs, which applies the formula output = alpha * input + beta where alpha controls contrast and beta controls brightness.
import numpy as np
alpha = 1.3 # 30% more contrast
beta = 20 # Increase brightness
img_adjusted = cv2.convertScaleAbs(img_blurred, alpha=alpha, beta=beta)
Thresholding for Segmentation
Binary segmentation uses cv2.threshold for global thresholds or cv2.adaptiveThreshold for varying lighting conditions.
gray = cv2.cvtColor(img_adjusted, cv2.COLOR_RGB2GRAY)
_, thresh = cv2.threshold(gray, 127, 255, cv2.THRESH_BINARY)
Geometric Transformations
The lesson covers two matrix-based transformation types in lessons/4-ComputerVision/06-IntroCV/README.md:
Affine Transformations (cv2.warpAffine) – Preserve parallel lines, enabling rotation, scaling, and translation using 2×3 matrices.
Perspective Transformations (cv2.warpPerspective) – Handle 3D viewpoint changes using 3×3 homography matrices for document rectification or viewpoint correction.
# Define three point correspondences for affine transform
pts_src = np.float32([[0, 0], [200, 0], [0, 200]])
pts_dst = np.float32([[20, 30], [180, 20], [30, 210]])
# Calculate and apply transformation matrix
M = cv2.getAffineTransform(pts_src, pts_dst)
img_affine = cv2.warpAffine(thresh, M, (200, 200))
Motion Analysis with Optical Flow
The curriculum extends into video processing with optical flow techniques for motion detection, distinguishing between:
- Dense Optical Flow – Calculates motion vectors for every pixel using
cv2.calcOpticalFlowFarneback - Sparse Optical Flow – Tracks specific feature points (mentioned in lab assignments)
The implementation in OpenCV.ipynb demonstrates Farneback's algorithm to compute motion between video frames:
cap = cv2.VideoCapture('video.mp4')
ret, prev = cap.read()
prev_gray = cv2.cvtColor(prev, cv2.COLOR_BGR2GRAY)
while True:
ret, frame = cap.read()
if not ret:
break
gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)
# Calculate dense optical flow
flow = cv2.calcOpticalFlowFarneBack(
prev_gray, gray, None,
0.5, 3, 15, 3, 5, 1.2, 0
)
# Visualize flow magnitude and direction
mag, ang = cv2.cartToPolar(flow[..., 0], flow[..., 1])
cap.release()
Running the Code Examples
The OpenCV.ipynb notebook in lessons/4-ComputerVision/06-IntroCV/ contains all executable examples referenced above. To run locally:
- Clone the
microsoft/AI-For-Beginnersrepository - Install the Conda environment:
conda env create -f environment.yml - Activate and launch Jupyter:
conda activate ai4beg && jupyter notebook - Navigate to
lessons/4-ComputerVision/06-IntroCV/OpenCV.ipynb
For cloud execution, the repository supports one-click deployment to Binder or GitHub Codespaces with pre-installed OpenCV dependencies.
Summary
- OpenCV loads images in BGR format requiring
cv2.cvtColorconversion for RGB display libraries - Pre-processing pipelines use
cv2.resize,cv2.GaussianBlur, andcv2.convertScaleAbsfor network-ready inputs - Binary segmentation applies
cv2.thresholdor adaptive variants for foreground extraction - Geometric corrections utilize
cv2.warpAffineandcv2.warpPerspectivewith transformation matrices - Motion analysis leverages
cv2.calcOpticalFlowFarnebackfor dense optical flow in video sequences - All implementations are provided in the executable
OpenCV.ipynbnotebook within the AI for Beginners curriculum
Frequently Asked Questions
How do I install OpenCV for the AI for Beginners curriculum?
The repository uses Conda for dependency management. OpenCV is included in the root environment.yml file. Run conda env create -f environment.yml to install all required packages including OpenCV, NumPy, and Matplotlib. Alternatively, the curriculum runs in GitHub Codespaces or Binder without local installation.
Why does my OpenCV image look blue when displayed with Matplotlib?
OpenCV stores images in BGR (Blue-Green-Red) order while Matplotlib expects RGB. According to the lesson at lessons/4-ComputerVision/06-IntroCV/README.md, you must convert using cv2.cvtColor(img, cv2.COLOR_BGR2RGB) before displaying with plt.imshow().
What is the difference between affine and perspective transformations in OpenCV?
Affine transformations (cv2.warpAffine) preserve parallel lines and require three point pairs to create a 2×3 matrix, suitable for rotation and scaling. Perspective transformations (cv2.warpPerspective) use four point pairs to create a 3×3 matrix, correcting for 3D viewpoint changes like document scanning or camera angle adjustments.
Which optical flow method does the AI for Beginners curriculum recommend?
The lesson demonstrates dense optical flow using cv2.calcOpticalFlowFarneback for computing motion vectors across entire frames. The accompanying lab assignment in lessons/4-ComputerVision/06-IntroCV/lab/README.md challenges students to extract specific motion directions from these flow calculations.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →