# How to Run Career-Ops Cost-Effectively Using Local or Custom AI Models

> Run Career-Ops cost-effectively with local or custom AI models. Process job descriptions for under $0.001 each or entirely offline thanks to model-agnostic evaluation via Ollama or OpenAI compatible endpoints.

- Repository: [Santiago Fernández de Valderrama/career-ops](https://github.com/santifer/career-ops)
- Tags: how-to-guide
- Published: 2026-07-03

---

**Career-Ops supports model-agnostic evaluation through OpenAI-compatible endpoints or local Ollama servers, allowing you to process job descriptions for less than $0.001 per evaluation or entirely offline.**

The **santifer/career-ops** repository separates prompt assembly from AI inference, storing evaluation logic in local Markdown files while delegating LLM calls to configurable backends. This design allows you to route requests to budget providers like DeepSeek, Qwen-2.5, or GLM-4 via OpenRouter, or to a self-hosted Ollama server, without modifying any evaluation logic. By keeping the heavy lifting in static files and only using tokens for the final inference step, you maintain full functionality while controlling costs.

## Architecture: Model-Agnostic Design

At the core of Career-Ops is a strict separation between **prompt templates** and **inference backends**. The repository stores system instructions in [`modes/_shared.md`](https://github.com/santifer/career-ops/blob/main/modes/_shared.md) and evaluation blocks in [`modes/oferta.md`](https://github.com/santifer/career-ops/blob/main/modes/oferta.md), which combine with your [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md) to form a ~10,000-token context.

Only the **evaluator scripts** consume LLM tokens. The `openai-eval.mjs` module sends assembled prompts to any OpenAI-compatible endpoint, while `ollama-eval.mjs` targets local servers. This means switching from Claude to a free local model requires only environment variable changes, not code modifications.

## Zero-Token Pipeline Stages

Two major components operate without any LLM token consumption:

- **Job-board scanning**: The `scan.mjs` script performs zero-token HTTP queries to Greenhouse, Lever, and Ashby APIs, populating [`data/pipeline.md`](https://github.com/santifer/career-ops/blob/main/data/pipeline.md) with fresh job URLs. This stage involves only network I/O.
- **PDF generation**: The `generate-pdf.mjs` script uses Playwright to convert tailored CVs into ATS-optimized PDFs entirely client-side, eliminating token costs for document creation.

These stages ensure that token usage is confined strictly to the evaluation step.

## Connecting to Budget LLM Providers

To reduce costs, route evaluations through `openai-eval.mjs` using cheap hosted APIs. The script accepts custom base URLs, models, and API keys via environment variables.

Set your provider details and run an evaluation:

```bash
export OPENAI_BASE_URL=https://openrouter.ai/api/v1
export OPENAI_MODEL=deepseek/deepseek-chat
export OPENAI_API_KEY=sk-...

node openai-eval.mjs --file ./jds/target.txt

```

This configuration processes approximately 4,500 tokens for **less than $0.001** per evaluation. The script reads your CV from [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md), merges it with the job description and prompt templates, then writes a Markdown report to `reports/` and updates the TSV tracker.

Alternatively, use the convenience wrapper:

```bash
node openrouter-runner.mjs evaluate <url-or-paste>

```

This automatically selects free or low-cost models from OpenRouter's catalog while maintaining the same evaluation flow.

## Running Fully Local with Ollama

For zero recurring costs, use **ollama-eval.mjs** to target a local Ollama server. As implemented in `santifer/career-ops`, this script includes security checks to ensure the endpoint is loopback-only (preventing accidental data leaks) and probes `/api/tags` before sending prompts.

Configure local evaluation:

```bash
export OLLAMA_BASE_URL=http://localhost:11434
export OLLAMA_MODEL=llama3.3  # or qwen2.5:72b for higher quality

node ollama-eval.mjs --file ./jds/target.txt

```

The local path supports models like **Llama 3.3** and **Qwen 2.5**, allowing completely offline operation without API fees.

## Batch Processing and Cost Controls

When processing multiple job descriptions, use [`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh) to enforce spending limits and prevent runaway token consumption.

Available safety flags:

- `--dry-run` – Preview which offers will run without calling the LLM
- `--limit 5` – Cap evaluation to a specific number of offers
- `--resume-paused` – Continue after a rate-limit pause without reprocessing completed items

These flags, documented in [`docs/RUNNING_ON_A_BUDGET.md`](https://github.com/santifer/career-ops/blob/main/docs/RUNNING_ON_A_BUDGET.md), let you validate batches before incurring costs and resume interrupted runs efficiently.

## Key Configuration Files

Understanding these files helps optimize your setup:

| File | Purpose | Cost Impact |
|------|---------|-------------|
| [`modes/_shared.md`](https://github.com/santifer/career-ops/blob/main/modes/_shared.md) | System prompt template (~10K tokens context) | Static file, zero cost |
| [`modes/oferta.md`](https://github.com/santifer/career-ops/blob/main/modes/oferta.md) | Evaluation criteria blocks | Static file, zero cost |
| `openai-eval.mjs` | Generic OpenAI-compatible evaluator | Token usage depends on selected provider |
| `ollama-eval.mjs` | Local Ollama evaluator | Zero API cost, local compute only |
| [`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh) | Parallel processing orchestrator | Caps token spend via `--limit` flag |
| [`docs/RUNNING_ON_A_BUDGET.md`](https://github.com/santifer/career-ops/blob/main/docs/RUNNING_ON_A_BUDGET.md) | Budget optimization guide | Reference documentation |

Configuration occurs through environment variables (`OPENAI_*`, `OLLAMA_*`) or CLI settings in [`.opencode/config.json`](https://github.com/santifer/career-ops/blob/main/.opencode/config.json), allowing you to switch providers without touching the core evaluation logic.

## Summary

- **Career-Ops** separates prompt logic from inference, enabling cheap or local LLM usage via `openai-eval.mjs` or `ollama-eval.mjs`.
- **Zero-token stages** (`scan.mjs`, `generate-pdf.mjs`) handle job discovery and PDF creation without API costs.
- **Budget providers** like DeepSeek via OpenRouter cost under $0.001 per evaluation for ~4,500 tokens.
- **Local execution** through Ollama (`ollama-eval.mjs`) supports offline operation using models like Llama 3.3 or Qwen 2.5.
- **Batch controls** (`--limit`, `--dry-run`, `--resume-paused`) in [`batch/batch-runner.sh`](https://github.com/santifer/career-ops/blob/main/batch/batch-runner.sh) prevent unexpected token expenditure.

## Frequently Asked Questions

### Can I run Career-Ops without an OpenAI API key?

Yes. By using `ollama-eval.mjs` with a local Ollama server, you can evaluate job descriptions entirely offline. The script checks that your endpoint is localhost-only before transmitting data, ensuring no external API calls or costs occur.

### Is it safe to use Ollama on a remote server with Career-Ops?

The `ollama-eval.mjs` script explicitly validates that the Ollama endpoint is loopback-only (127.0.0.1 or localhost) to prevent accidental data leakage. For remote Ollama instances, you would need to modify the security check in the source or use a VPN/tunnel to present the remote server as local, as the default configuration blocks non-local endpoints to protect your CV data.

### How much does it cost to evaluate 100 job descriptions using cheap providers?

Using DeepSeek via OpenRouter at approximately 4,500 tokens per evaluation, 100 job descriptions would cost roughly **$0.10** total (under $0.001 per evaluation). This assumes the standard ~10K token context assembled from [`modes/_shared.md`](https://github.com/santifer/career-ops/blob/main/modes/_shared.md), [`modes/oferta.md`](https://github.com/santifer/career-ops/blob/main/modes/oferta.md), and [`cv.md`](https://github.com/santifer/career-ops/blob/main/cv.md), with most providers charging only a few cents per million tokens.

### What is the cheapest model that works well with Career-Ops?

According to the [`docs/RUNNING_ON_A_BUDGET.md`](https://github.com/santifer/career-ops/blob/main/docs/RUNNING_ON_A_BUDGET.md) guidelines, **DeepSeek V3** via OpenRouter offers an excellent balance of cost and quality at under $0.001 per evaluation. For local execution, **Qwen 2.5** (72B parameter version) or **Llama 3.3** provide strong evaluation capabilities at zero API cost, requiring only sufficient local GPU or CPU resources to run the inference.