# How to Configure Local Models with LM Studio in G0DM0D3: A Complete Setup Guide

> Learn how to configure local models with LM Studio for G0DM0D3. Route all LLM operations to your private instance for enhanced control and privacy. A complete setup guide.

- Repository: [pliny/G0DM0D3](https://github.com/elder-plinius/G0DM0D3)
- Tags: how-to-guide
- Published: 2026-07-19

---

**G0DM0D3 supports local inference through any OpenAI-compatible server, allowing you to route all LLM operations—including classification, judging, and refinement—to a private LM Studio instance running on localhost.**

The G0DM0D3 repository by elder-plinius provides a privacy-focused interface that can operate entirely offline by connecting to local inference engines. When you configure local models with LM Studio, the application bypasses cloud providers like OpenRouter and Venice, ensuring sensitive data never leaves your machine while maintaining full access to features like TASTEMAKER judging and **Liquid** refinement.

## Prerequisites: Launching the LM Studio Server

Before configuring G0DM0D3, start LM Studio and load your desired model. LM Studio exposes an HTTP server at `http://localhost:1234/v1` by default, implementing the standard OpenAI-compatible `/v1/models` and `/v1/chat/completions` endpoints. Verify the server is running by testing `http://localhost:1234/v1/models` in your browser or via `curl`.

## Configuring the Local Endpoint in G0DM0D3

### Step 1: Access the Settings Modal

Open the Settings modal in the G0DM0D3 web interface. The component [`SettingsModal.tsx`](https://github.com/elder-plinius/G0DM0D3/blob/main/SettingsModal.tsx) initializes state hooks for your local configuration at lines 1648–1654, reading the current `localUrl` and `localKey` values from the global store. These fields control where the application sends inference requests when operating in private mode.

### Step 2: Enter the LM Studio Base URL

Locate the *Local Model URL* input field and enter `http://localhost:1234/v1` (or your custom LM Studio URL). In [`src/components/SettingsModal.tsx`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/components/SettingsModal.tsx) at line 1703, this input binds directly to the `localUrl` state via `setLocalUrl`, updating the target endpoint in real time.

### Step 3: Test and Discover Available Models

Click the *Test & Discover Models* button to verify connectivity. This triggers a request to `${baseUrl}/models` handled by [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts) at line 45. The function parses the JSON response and populates the model selector with available local models, ensuring you can choose specific variants loaded in LM Studio.

### Step 4: Enable Local-Only Mode

Toggle the *Local-only mode* switch to force all traffic to your private server. This setting updates the `localOnly` boolean flag defined in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) at line 12. When active, the request layer excludes OpenRouter and Venice from every "ULTRAPLINIAN" race, disables telemetry to remote G0DM0D3 services, and routes all internal LLM calls—including query classification, refusal checks, and coaching—to your local endpoint.

## How Request Routing Works Under the Hood

The `chatCompletion` function in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts) dynamically selects the appropriate base URL based on your configuration. When `localOnly` is true, it uses `store.getState().localUrl` instead of the cloud provider URL.

```typescript
export async function chatCompletion(payload: ChatPayload) {
  const base = store.getState().localOnly
    ? store.getState().localUrl      // ← LM Studio endpoint
    : store.getState().ultraplinianApiUrl; // ← cloud provider

  const resp = await fetch(`${base}/chat/completions`, {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify(payload),
  });
  return resp.json();
}

```

This conditional routing ensures that enabling local mode affects every downstream operation, from initial query classification to final **Liquid** refinement passes, without requiring code changes or restarts.

## Handling CORS for Hosted vs. Local Deployments

If you access G0DM0D3 from the hosted page at `https://godmod3.ai`, you must configure LM Studio to allow cross-origin requests from that domain. Start LM Studio with the `--origin` flag:

```bash
lm-studio start --origin https://godmod3.ai

```

Alternatively, run the UI locally to avoid CORS restrictions entirely:

```bash
python3 -m http.server 8080

```

Then navigate to `http://localhost:8080` and point the settings to `http://localhost:1234/v1`. Local deployments require no additional CORS configuration since both the UI and inference server share the same origin.

## Troubleshooting Common Connection Issues

**Connection failed / Failed to fetch** – Verify LM Studio is running and accessible at your configured URL. If using the hosted UI, confirm you started LM Studio with the appropriate `--origin` flag to enable CORS for `https://godmod3.ai`.

**No models returned** – Ensure a `GET` request to `<your-base-url>/models` returns a JSON object containing a `data` array with objects structured as `{ "id": "model-name" }`. LM Studio must have a model actively loaded to populate this list.

**Slow responses or timeouts** – Reduce the *Max Tokens* setting in G0DM0D3 or unload unused models from LM Studio to free up VRAM. Local inference speed depends heavily on your hardware configuration and model quantization level.

## Summary

- G0DM0D3 connects to LM Studio via OpenAI-compatible endpoints at `/v1/models` and `/v1/chat/completions`.
- Configure the connection in [`SettingsModal.tsx`](https://github.com/elder-plinius/G0DM0D3/blob/main/SettingsModal.tsx) by setting the base URL and enabling *Local-only mode* stored in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts).
- The `localOnly` flag forces all inference requests to bypass cloud providers, disabling external telemetry while preserving full functionality.
- Persistence is handled automatically through `localStorage` serialization in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) at line 78.
- CORS configuration is only required when accessing the UI from `https://godmod3.ai`; local deployments work without additional flags.

## Frequently Asked Questions

### What endpoints must LM Studio expose for G0DM0D3 compatibility?

LM Studio must provide OpenAI-compatible `/v1/models` and `/v1/chat/completions` endpoints. The application queries `/v1/models` to populate the model selector and sends all inference requests to `/v1/chat/completions` with standard ChatML formatting.

### Does enabling local-only mode disable cloud providers entirely?

Yes. When the `localOnly` flag in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) is true, the request layer in [`src/lib/openrouter.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/lib/openrouter.ts) excludes OpenRouter and Venice from all operations. This prevents any data from being sent to external APIs and disables telemetry to the G0DM0D3 service.

### How does G0DM0D3 persist my LM Studio configuration?

The global Zustand store defined in [`src/store/index.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/src/store/index.ts) serializes the `localUrl` and `localOnly` states into `localStorage` whenever they change. This ensures your LM Studio endpoint and mode preference persist across browser sessions and page reloads without requiring reconfiguration.

### Can I use G0DM0D3 with other local inference servers besides LM Studio?

Yes. Any inference server implementing the OpenAI-compatible `/v1/models` and `/v1/chat/completions` specifications will work. Simply enter the server's base URL (e.g., `http://localhost:5000/v1`) into the *Local Model URL* field. The application treats all compatible endpoints identically, whether they come from LM Studio, Ollama, or custom implementations.