# Can CreativeMath Be Used with Local LLMs? Complete Setup Guide

> Explore how CreativeMath integrates with local LLMs using the ModelWrapper class for seamless Hugging Face Transformers inference. Get started with your local models today.

- Repository: [Junyi Ye/creativemath](https://github.com/junyiye/creativemath)
- Tags: how-to-guide
- Published: 2026-03-05

---

**Yes, CreativeMath fully supports local large language models through the `ModelWrapper` class, which automatically routes non-API model names to local inference pipelines using Hugging Face Transformers.**

The junyiye/creativemath repository provides a flexible mathematics reasoning framework that operates with both cloud-based APIs and local LLMs. For researchers and developers concerned with data privacy, API costs, or offline availability, CreativeMath local LLMs integration enables you to run models like DeepSeek-Math-7B-RL and Llama-3-70B entirely on your own hardware without external dependencies.

## How CreativeMath Detects and Routes Local Models

### Automatic Model Type Detection

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) serves as the central router for all model interactions. When you instantiate the wrapper with a model name, it checks the identifier against an internal registry of known API models such as `gpt-4` or `claude-3-opus`. If the provided name does not match any API entry, the wrapper automatically treats the request as a local deployment and bypasses remote authentication.

### Local Inference Architecture

For local execution, CreativeMath invokes two critical functions defined in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py):

- **`load_local_model`**: Downloads model weights via the 🤗 Transformers library and initializes both the model and tokenizer. Weights are cached locally after first download.
- **`generate_local_response`**: Handles prompt formatting, tokenization, and text generation with architecture-specific adaptations for supported models including DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B.

All generation parameters, model identifiers, and hardware acceleration settings are defined in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json), which is parsed at runtime by [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py) and made available throughout the inference pipeline.

## Running CreativeMath with Local LLMs

### Basic Local Model Usage

To execute mathematical reasoning with a local model, instantiate `ModelWrapper` using the exact identifier listed in your configuration:

```python
from models.model_loader import ModelWrapper

# Initialize with a supported local model

wrapper = ModelWrapper("Deepseek-math-7b-rl")

prompt = "Explain the derivative of sin(x)."
response = wrapper.generate_response(prompt)

print(response)

```

When executing this code, `ModelWrapper` automatically triggers `load_local_model` to initialize the pipeline and routes your prompt through `generate_local_response`. The model downloads automatically on first use if not present in the local Hugging Face cache.

### Switching Between API and Local Models

CreativeMath allows seamless mixing of API-based and local models within the same workflow, determined solely by the model name provided:

```python

# Remote API model (requires API key in config.json)

api_wrapper = ModelWrapper("gpt-4")
api_answer = api_wrapper.generate_response("What is 2+2?")

# Local model (runs offline, no API key needed)

local_wrapper = ModelWrapper("Llama-3-70B")
local_answer = local_wrapper.generate_response("Summarize Euler's formula.")

```

The architecture automatically handles the underlying differences, using [`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py) for remote calls and [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py) for offline inference.

## Configuring and Extending Local Model Support

### Configuration via config.json

The framework's flexibility relies on [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json), which maps human-readable model names to Hugging Face repository identifiers and stores generation hyperparameters such as temperature, max tokens, and device mapping. The [`config.py`](https://github.com/junyiye/creativemath/blob/main/config.py) module loads these mappings at import time, ensuring consistent behavior across the `ModelWrapper` interface.

### Adding Custom Local Models

To integrate additional local models beyond the pre-configured set:

1. Add the model's Hugging Face identifier to [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) under the `model_version` section
2. Extend `load_local_model` in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py) to handle specific architecture requirements (e.g., quantization settings or trust_remote_code flags)
3. Update `generate_local_response` if the model requires unique prompt templates or generation parameters, following the existing pattern for DeepSeek or Llama models

## Summary

- **CreativeMath local LLMs** are supported through automatic detection in `ModelWrapper` when model names don't match known API entries in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py)
- Local inference uses `load_local_model` and `generate_local_response` in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py) with the 🤗 Transformers library
- Requirements include `torch` and `transformers`; model weights download automatically on first use and cache for offline operation
- Configuration is centralized in [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) and loaded via [`src/config.py`](https://github.com/junyiye/creativemath/blob/main/src/config.py), allowing easy extension to new models
- You can run fully offline workflows with supported architectures like DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B

## Frequently Asked Questions

### What local models are currently supported by CreativeMath?

CreativeMath officially supports DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B through dedicated loading logic in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py). You can extend support to any Hugging Face Transformers-compatible model by updating [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json) and the loading functions to recognize the new architecture.

### Do I need an internet connection to use CreativeMath with local LLMs?

No. After the initial model weight download (cached via Hugging Face's standard `cache_dir` mechanism), CreativeMath runs entirely offline. No API keys or network connectivity are required for local models, making this configuration suitable for air-gapped environments or sensitive mathematical computations.

### How does CreativeMath decide whether to use an API or local model?

The `ModelWrapper` class in [`src/models/model_loader.py`](https://github.com/junyiye/creativemath/blob/main/src/models/model_loader.py) checks the provided model name against a hardcoded list of API providers. If the name matches entries like `gpt-4` or `claude-3-opus`, it routes to the API handlers in [`src/models/api_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/api_models.py); otherwise, it triggers local inference via `load_local_model` in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py).

### Can I use my own fine-tuned models with CreativeMath?

Yes. Add your model's Hugging Face repository identifier or local filesystem path to [`config.json`](https://github.com/junyiye/creativemath/blob/main/config.json), then ensure `load_local_model` in [`src/models/local_models.py`](https://github.com/junyiye/creativemath/blob/main/src/models/local_models.py) can resolve the architecture configuration. As long as your model uses standard Transformers generation APIs, it will work with the existing `generate_local_response` pipeline without additional modifications.