# How to Optimize LLM Costs During Spider Generation in SpiderCreator

> Optimize LLM costs in SpiderCreator with cheaper models, caching, prompt trimming, batching, and structured outputs. Reduce expenses during spider generation effectively.

- Repository: [Carlos A. Planchón/spidercreator](https://github.com/carlosplanchon/spidercreator)
- Tags: optimization
- Published: 2026-02-26

---

**You can reduce LLM costs by switching to cheaper models for early-stage generation, caching repeated calls, trimming prompt payloads, batching expensive verification steps, and leveraging structured outputs throughout the SpiderCreator pipeline.**

SpiderCreator is an open-source project that transforms user recordings into functional web-scraping spiders using multiple LLM calls. Optimizing LLM costs during the spider generation process is essential because the pipeline invokes expensive models like GPT-4.1 and GPT-4o across six distinct stages, from initial draft creation to final execution verification.

## Understanding the LLM Cost Drivers in SpiderCreator

The pipeline defined in [`shared.py`](https://github.com/carlosplanchon/spidercreator/blob/main/shared.py) ([lines 3‑13](https://github.com/carlosplanchon/spidercreator/blob/main/shared.py#L3-L13)) orchestrates several model instances that drive costs at different phases:

| Stage | LLM Instance | Purpose | Cost Impact |
|-------|--------------|---------|-------------|
| 1. Recording → Scrapy draft | `o3_llm` | Generates initial spider code from raw recordings | Low (cheap model) |
| 2. Mind‑map → XPath planning (text) | `o3_llm` | Produces textual extraction plans | Low |
| 3. Mind‑map → XPath planning (structured) | `o3_llm` with structured output | Returns JSON schema (`Planning`) | Low |
| 4. Spider combination | `gpt41_llm` | Merges multiple drafts into one script | **High** |
| 5. Address remapping | `gpt4o_llm` | Rewrites URLs for local testing | **High** |
| 6. Execution verification | `gpt4o_llm` | Scores spider run against criteria | **High** |

Steps 4‑6 in [`pipeline/sp_combination.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_combination.py) ([lines 99‑106](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_combination.py#L99-L106)), [`pipeline/sp_addr_remapping.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_addr_remapping.py) ([lines 45‑68](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_addr_remapping.py#L45-L68)), and [`pipeline/verify_sp_execution.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/verify_sp_execution.py) ([lines 54‑78](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/verify_sp_execution.py#L54-L78)) are the primary cost drivers because they use high‑capability models.

## Six Strategies to Optimize LLM Costs

### 1. Consolidate Model Choices for Early Stages

Most early‑stage steps only need a fast, cheap model. The default `model="gpt-4.1"` for `gpt41_llm` and `gpt4o_llm` is overkill for drafting. Replace them with lower‑priced models (e.g., `gpt-3.5-turbo` or an open‑source LLM) **only for the expensive stages**.

```python

# shared.py – choose cheaper model for intermediate steps

from langchain.chat_models import init_chat_model

# cheap model for drafts & planning

o3_llm = init_chat_model(model="gpt-3.5-turbo", model_provider="openai")

# high‑quality model only where needed (e.g., combination)

gpt41_llm = init_chat_model(model="gpt-4.1", model_provider="openai")
gpt4o_llm = init_chat_model(model="gpt-4.1", model_provider="openai")

```

### 2. Cache Repeated LLM Calls

During a single run the same prompt can be regenerated (e.g., when retrying or when multiple spiders share the same mind‑map). A simple memoisation wrapper avoids duplicate API requests.

```python

# utils/cache_llm.py

from functools import lru_cache
from typing import Callable
from langchain_core.messages import HumanMessage

def cache_llm_invoke(llm: Callable, prompt: str):
    @lru_cache(maxsize=128)
    def _cached(content: str):
        return llm.invoke([HumanMessage(content=content)])

    return _cached(prompt)

# Usage in xpath_builder_planning.py

from utils.cache_llm import cache_llm_invoke
from shared import o3_llm

def make_non_structured_planning(mermaid_code: str, scrapy_spider: str) -> str:
    prompt = XPATH_BUILDER_PLANNING_PROMPT.format(
        mindmap=mermaid_code, scrapy_spider=scrapy_spider
    )
    result = cache_llm_invoke(o3_llm, prompt)
    return result.content

```

### 3. Trim and Compact Prompt Payloads

LLM pricing is token‑based. The prompts contain large JSON blobs (`recordings`, `mindmap`, etc.). Before sending them:

1. **Serialize compactly** – `json.dumps(..., separators=(",", ":"))` removes whitespace.
2. **Truncate** non‑essential sections (e.g., keep only the last N recordings).

```python
import json

def compact_recordings(recordings: list[dict]) -> str:
    # Keep only the most recent 20 recordings

    recent = recordings[-20:]
    return json.dumps(recent, separators=(",", ":"))

```

Apply this in `make_scrapy_spider_draft` (see **[`pipeline/spider_draft.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/spider_draft.py)** [lines 31‑42](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/spider_draft.py#L31-L42)).

### 4. Batch Multiple Verification Requests

When several spiders need the same verification step, batch them into a single LLM call that returns a list of `XPathExecutionVerificationResult`. Update the structured output schema accordingly and invoke once instead of many times.

```python

# verification_pipeline.py (new helper)

from pydantic import BaseModel
from typing import List
from langchain_core.messages import HumanMessage

class BatchVerificationResult(BaseModel):
    results: List[XPathExecutionVerificationResult]

structured_batch_verifier = gpt4o_llm.with_structured_output(BatchVerificationResult)

def batch_verify(spider_infos: list[dict]) -> BatchVerificationResult:
    # Build a combined prompt listing each spider with its data

    combined_prompt = "\n---\n".join(
        f"Spider {i}:\n{XPATH_EXECUTION_VERIFICATION_PROMPT.format(**info)}"
        for i, info in enumerate(spider_infos, 1)
    )
    return structured_batch_verifier.invoke([HumanMessage(content=combined_prompt)])

```

### 5. Prefer Structured Outputs When Possible

Structured output (`with_structured_output`) reduces post‑processing overhead and often yields *shorter* responses because the model follows a strict schema. Use it for all steps that already have a Pydantic model (e.g., `Planning`, `XPathExecutionVerificationResult`). The current code already does this in [`pipeline/xpath_builder_planning.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/xpath_builder_planning.py) ([lines 27‑90](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/xpath_builder_planning.py#L27-L90)); keep it and avoid free‑form text generation where structured alternatives exist.

### 6. Monitor Token Usage

Insert lightweight logging around every `invoke` to record `prompt_tokens` and `completion_tokens`. This data enables you to spot unexpectedly large prompts and tune them.

```python
def log_usage(result):
    usage = getattr(result, "usage", None)
    if usage:
        print(f"Prompt tokens: {usage.prompt_tokens}, Completion tokens: {usage.completion_tokens}")

# Example

resp = o3_llm.invoke([...])
log_usage(resp)

```

## Summary

- **Consolidate model choices** – Use cheap models like `gpt-3.5-turbo` for drafting and planning stages defined in [`shared.py`](https://github.com/carlosplanchon/spidercreator/blob/main/shared.py), reserving `gpt-4.1` only for combination and verification.
- **Cache repeated calls** – Implement `lru_cache` wrappers around `invoke` methods to eliminate redundant API requests during retries or shared mind-map processing.
- **Trim prompt payloads** – Compact JSON with `separators=(",", ":")` and truncate recordings to the last 20 entries before sending to [`pipeline/spider_draft.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/spider_draft.py).
- **Batch expensive operations** – Combine multiple verification requests into a single structured output call returning `BatchVerificationResult` instead of individual `gpt4o_llm` invocations.
- **Prefer structured outputs** – Leverage `with_structured_output` with Pydantic models to reduce token waste and post-processing overhead.
- **Monitor token usage** – Log `prompt_tokens` and `completion_tokens` after every invoke to identify cost spikes early.

## Frequently Asked Questions

### How does SpiderCreator use LLMs to generate spiders?

SpiderCreator orchestrates a six-stage pipeline that transforms raw user recordings into executable Scrapy spiders. According to the source code in [`shared.py`](https://github.com/carlosplanchon/spidercreator/blob/main/shared.py), the system uses `o3_llm` for initial drafting and planning, `gpt41_llm` for combining multiple spider drafts, and `gpt4o_llm` for address remapping and execution verification. Each stage is defined in separate pipeline files such as [`pipeline/spider_draft.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/spider_draft.py) and [`pipeline/sp_combination.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_combination.py).

### Which LLM calls in SpiderCreator are the most expensive?

The costliest invocations occur in stages 4, 5, and 6. Stage 4 uses `gpt41_llm` in [`pipeline/sp_combination.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_combination.py) to merge multiple drafts into a single script. Stage 5 employs `gpt4o_llm` in [`pipeline/sp_addr_remapping.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/sp_addr_remapping.py) to rewrite URLs for local testing. Stage 6 utilizes `gpt4o_llm` again in [`pipeline/verify_sp_execution.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/verify_sp_execution.py) to score spider runs against verification criteria. These high-capability models charge significantly more per token than the `o3_llm` used in early stages.

### Can I use open-source models instead of OpenAI in SpiderCreator?

Yes. The `init_chat_model` function from LangChain used in [`shared.py`](https://github.com/carlosplanchon/spidercreator/blob/main/shared.py) supports multiple providers. You can replace `gpt-3.5-turbo` or `gpt-4.1` with open-source alternatives such as `llama3-70b` via Ollama or `mixtral-8x7b` via Groq by changing the `model` and `model_provider` parameters. However, ensure the replacement model supports structured output schemas if you are modifying stages that rely on Pydantic models like `Planning` or `XPathExecutionVerificationResult`.

### What is the fastest way to reduce token usage without changing models?

Implement prompt compaction and caching. In [`pipeline/spider_draft.py`](https://github.com/carlosplanchon/spidercreator/blob/main/pipeline/spider_draft.py), replace standard JSON serialization with `json.dumps(recordings, separators=(",", ":"))` to remove whitespace, and truncate recordings to the last 20 entries before sending. Additionally, wrap LLM invocations with `functools.lru_cache` to prevent duplicate API calls when the same mind-map or recording set appears multiple times during a run. These changes reduce both input tokens and API request volume immediately.