Can CreativeMath Be Used with Local LLMs? Complete Setup Guide
Yes, CreativeMath fully supports local large language models through the ModelWrapper class, which automatically routes non-API model names to local inference pipelines using Hugging Face Transformers.
The junyiye/creativemath repository provides a flexible mathematics reasoning framework that operates with both cloud-based APIs and local LLMs. For researchers and developers concerned with data privacy, API costs, or offline availability, CreativeMath local LLMs integration enables you to run models like DeepSeek-Math-7B-RL and Llama-3-70B entirely on your own hardware without external dependencies.
How CreativeMath Detects and Routes Local Models
Automatic Model Type Detection
The ModelWrapper class in src/models/model_loader.py serves as the central router for all model interactions. When you instantiate the wrapper with a model name, it checks the identifier against an internal registry of known API models such as gpt-4 or claude-3-opus. If the provided name does not match any API entry, the wrapper automatically treats the request as a local deployment and bypasses remote authentication.
Local Inference Architecture
For local execution, CreativeMath invokes two critical functions defined in src/models/local_models.py:
load_local_model: Downloads model weights via the 🤗 Transformers library and initializes both the model and tokenizer. Weights are cached locally after first download.generate_local_response: Handles prompt formatting, tokenization, and text generation with architecture-specific adaptations for supported models including DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B.
All generation parameters, model identifiers, and hardware acceleration settings are defined in config.json, which is parsed at runtime by src/config.py and made available throughout the inference pipeline.
Running CreativeMath with Local LLMs
Basic Local Model Usage
To execute mathematical reasoning with a local model, instantiate ModelWrapper using the exact identifier listed in your configuration:
from models.model_loader import ModelWrapper
# Initialize with a supported local model
wrapper = ModelWrapper("Deepseek-math-7b-rl")
prompt = "Explain the derivative of sin(x)."
response = wrapper.generate_response(prompt)
print(response)
When executing this code, ModelWrapper automatically triggers load_local_model to initialize the pipeline and routes your prompt through generate_local_response. The model downloads automatically on first use if not present in the local Hugging Face cache.
Switching Between API and Local Models
CreativeMath allows seamless mixing of API-based and local models within the same workflow, determined solely by the model name provided:
# Remote API model (requires API key in config.json)
api_wrapper = ModelWrapper("gpt-4")
api_answer = api_wrapper.generate_response("What is 2+2?")
# Local model (runs offline, no API key needed)
local_wrapper = ModelWrapper("Llama-3-70B")
local_answer = local_wrapper.generate_response("Summarize Euler's formula.")
The architecture automatically handles the underlying differences, using src/models/api_models.py for remote calls and src/models/local_models.py for offline inference.
Configuring and Extending Local Model Support
Configuration via config.json
The framework's flexibility relies on config.json, which maps human-readable model names to Hugging Face repository identifiers and stores generation hyperparameters such as temperature, max tokens, and device mapping. The config.py module loads these mappings at import time, ensuring consistent behavior across the ModelWrapper interface.
Adding Custom Local Models
To integrate additional local models beyond the pre-configured set:
- Add the model's Hugging Face identifier to
config.jsonunder themodel_versionsection - Extend
load_local_modelinsrc/models/local_models.pyto handle specific architecture requirements (e.g., quantization settings or trust_remote_code flags) - Update
generate_local_responseif the model requires unique prompt templates or generation parameters, following the existing pattern for DeepSeek or Llama models
Summary
- CreativeMath local LLMs are supported through automatic detection in
ModelWrapperwhen model names don't match known API entries insrc/models/model_loader.py - Local inference uses
load_local_modelandgenerate_local_responseinsrc/models/local_models.pywith the 🤗 Transformers library - Requirements include
torchandtransformers; model weights download automatically on first use and cache for offline operation - Configuration is centralized in
config.jsonand loaded viasrc/config.py, allowing easy extension to new models - You can run fully offline workflows with supported architectures like DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B
Frequently Asked Questions
What local models are currently supported by CreativeMath?
CreativeMath officially supports DeepSeek-Math-7B-RL, Llama-3-70B, and Mixtral-8x22B through dedicated loading logic in src/models/local_models.py. You can extend support to any Hugging Face Transformers-compatible model by updating config.json and the loading functions to recognize the new architecture.
Do I need an internet connection to use CreativeMath with local LLMs?
No. After the initial model weight download (cached via Hugging Face's standard cache_dir mechanism), CreativeMath runs entirely offline. No API keys or network connectivity are required for local models, making this configuration suitable for air-gapped environments or sensitive mathematical computations.
How does CreativeMath decide whether to use an API or local model?
The ModelWrapper class in src/models/model_loader.py checks the provided model name against a hardcoded list of API providers. If the name matches entries like gpt-4 or claude-3-opus, it routes to the API handlers in src/models/api_models.py; otherwise, it triggers local inference via load_local_model in src/models/local_models.py.
Can I use my own fine-tuned models with CreativeMath?
Yes. Add your model's Hugging Face repository identifier or local filesystem path to config.json, then ensure load_local_model in src/models/local_models.py can resolve the architecture configuration. As long as your model uses standard Transformers generation APIs, it will work with the existing generate_local_response pipeline without additional modifications.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →