# What AI Models and Libraries Are Integrated into ai-job-search?

> Discover the AI models and libraries powering ai-job-search. Explore integrations with GPT-4, Claude, LangChain, BeautifulSoup, Selenium, PyPDF2, and FastAPI.

- Repository: [Mads Lorentzen/ai-job-search](https://github.com/MadsLorentzen/ai-job-search)
- Tags: internals
- Published: 2026-09-02

---

**The ai-job-search repository integrates OpenAI GPT-4, Anthropic Claude, and LangChain for LLM orchestration, alongside BeautifulSoup4 and Selenium for web scraping, PyPDF2 for PDF parsing, and FastAPI for the optional web interface.**

The `ai-job-search` project is an open-source Python toolkit designed to automate job searching, application generation, and interview preparation through artificial intelligence. According to the source code in the `MadsLorentzen/ai-job-search` repository, the project leverages multiple large language model providers and specialized libraries to power its ranking engine and document processing pipeline. Understanding what AI models and libraries are integrated into ai-job-search is essential for developers seeking to extend functionality or deploy custom workflows.

## Core AI Models and LLM Providers

### OpenAI SDK for GPT-4 and ChatGPT

The repository utilizes the **OpenAI SDK** (`openai>=0.27.0`) as specified in [`requirements.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/requirements.txt) to access GPT-4 and ChatGPT models for text generation tasks. In [`tools/check_framework_version.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/check_framework_version.py), the project programmatically verifies the installed OpenAI package using pip introspection to ensure compatibility.

```python
import sys
import subprocess

def get_version(package_name: str) -> str:
    try:
        result = subprocess.run([sys.executable, '-m', 'pip', 'show', package_name], capture_output=True, text=True)
        for line in result.stdout.splitlines():
            if line.startswith('Version:'):
                return line.split(':')[1].strip()
    except Exception:
        return 'unknown'

def main():
    packages = ['openai', 'anthropic', 'langchain']
    versions = {pkg: get_version(pkg) for pkg in packages}
    print(json.dumps(versions, indent=2))

```

### Anthropic Claude SDK

As an alternative LLM provider, the **Anthropic Claude SDK** (`anthropic>=0.3.0`) is included in [`requirements.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/requirements.txt) and checked alongside OpenAI in [`tools/check_framework_version.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/check_framework_version.py). This dual-provider architecture allows users to switch between GPT-4 and Claude models for job ranking and cover letter generation without refactoring the underlying orchestration logic.

### LangChain for Prompt Orchestration

The **LangChain** library (`langchain>=0.0.200`) serves as the orchestration layer for prompt chains and memory management. According to the repository documentation, LangChain coordinates the LLM interactions for the ranking engine and cover letter workflows, enabling complex multi-step AI operations that maintain context across job screening and generation tasks.

## Data Processing and Web Scraping Libraries

### BeautifulSoup4 and Selenium

For pulling job listings from portals like LinkedIn and Indeed, the repository depends on **BeautifulSoup4** (`beautifulsoup4>=4.12.2`) and **Selenium** (`selenium>=4.9.0`). These libraries handle HTML parsing and browser automation respectively, forming the foundation of the job scraper component that normalizes data before AI processing.

### PyPDF2 for CV Parsing

The [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py) file implements PDF text extraction using **PyPDF2** (`PyPDF2>=3.0.1`) to parse CV documents. The `extract_text` function demonstrates direct integration with the PDF reader to extract candidate information for later processing by the LLM components.

```python
import sys
import PyPDF2

def extract_text(pdf_path: str) -> str:
    reader = PyPDF2.PdfReader(pdf_path)
    text = ''
    for page in reader.pages:
        text += page.extract_text() or ''
    return text

def main():
    if len(sys.argv) < 2:
        print('Usage: verify_pdf.py <pdf_path>')
        sys.exit(1)
    pdf_path = sys.argv[1]
    print(extract_text(pdf_path))

```

### Pandas and NumPy

Data normalization and analysis rely on **Pandas** (`pandas>=2.0.3`) and **NumPy** (`numpy>=1.24.3`). These libraries process scraped job data, handle tabular storage of listings, and calculate preliminary ranking metrics before LLM evaluation in the filtering pipeline.

## Automation and Infrastructure

### Google API Client for Gmail Integration

The **Google API Python Client** (`google-api-python-client>=2.92.0`) enables automated email composition and sending through Gmail. This integration supports the application executor component that automates submission workflows, allowing the system to send tailored applications directly from the user's inbox.

### FastAPI and Uvicorn

For optional web-based workflow management, the repository includes **FastAPI** (`fastapi>=0.95.2`) and **Uvicorn** (`uvicorn>=0.22.0`). These provide the asynchronous server capabilities for the management interface, exposed via ASGI through Uvicorn.

## Summary

- The **OpenAI** and **Anthropic** SDKs provide dual LLM provider support for text generation and job ranking tasks.
- **LangChain** orchestrates complex prompt chains and memory management across the AI workflow components.
- **BeautifulSoup4** and **Selenium** power the job scraping functionality for external recruitment portals.
- **PyPDF2** handles CV parsing in [`tools/verify_pdf.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/verify_pdf.py), while **Pandas** and **NumPy** manage data structures and analysis.
- **Google API Client**, **FastAPI**, and **Uvicorn** enable email automation and web interface capabilities.

## Frequently Asked Questions

### Does ai-job-search support local or open-source LLMs?

The current implementation in [`requirements.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/requirements.txt) specifies cloud-based providers (OpenAI and Anthropic). While the architecture using LangChain could theoretically support local models, the source code analysis shows no explicit integration with local LLM servers like Ollama or text-generation-webui in the current codebase.

### How does the repository handle API authentication?

The [`tools/lint_skills.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/lint_skills.py) file includes a security check for API keys using regex patterns (`re.search(r'(?i)api[_-]?key', ...)`), indicating the project enforces key security through linting rules rather than hardcoded credentials. Users must configure API keys through environment variables or external configuration files not tracked in version control.

### Is LangChain required for basic job scraping?

No. According to [`requirements.txt`](https://github.com/MadsLorentzen/ai-job-search/blob/main/requirements.txt), web scraping depends solely on BeautifulSoup4 and Selenium. LangChain is specifically utilized for LLM orchestration in the ranking engine and cover letter generation components, making it necessary only for AI-powered features, not for data collection.

### What Python version is required to run ai-job-search?

While the repository does not explicitly specify a Python version constraint in the analyzed files, the dependency versions—particularly Pandas 2.0+ and FastAPI 0.95+—imply compatibility with Python 3.8 or newer. The subprocess-based version checking in [`tools/check_framework_version.py`](https://github.com/MadsLorentzen/ai-job-search/blob/main/tools/check_framework_version.py) uses standard library features available in Python 3.6 and above.