# How to Opt Into Dataset Collection for Research with G0DM0D3

> Easily opt into G0DM0D3 dataset collection for research. Simply add contribute_to_dataset: true to your JSON payload when using G0DM0D3 completion endpoints.

- Repository: [pliny/G0DM0D3](https://github.com/elder-plinius/G0DM0D3)
- Tags: how-to-guide
- Published: 2026-07-19

---

**To opt into dataset collection for research, include `"contribute_to_dataset": true` in the JSON payload when calling any of G0DM0D3's primary completion endpoints.**

G0DM0D3 by elder-plinius offers an optional dataset collection mechanism designed for research purposes. This feature is strictly opt-in per request, allowing you to contribute conversation data to a public Hugging Face dataset only when explicitly desired. Understanding how to opt into dataset collection for research ensures you maintain full control over your data privacy while supporting open-source AI development.

## How the Dataset Collection Opt-In Works

The dataset collection system in G0DM0D3 operates on a per-request basis. When you include the Boolean field **`contribute_to_dataset: true`** in your API request, the server evaluates this flag and captures the full non-system conversation—including user messages, assistant responses, and pipeline metadata—into an in-memory buffer. This data is later flushed to a public Hugging Face dataset repository specified by the `HF_DATASET_REPO` environment variable.

Privacy safeguards are built into the implementation. According to the source code in [`api/lib/dataset.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/dataset.ts), the system explicitly excludes PII such as API keys, IP addresses, and authentication tokens from collection. By default, the feature is disabled, meaning no data is collected unless you explicitly set the flag to `true`.

## Supported Endpoints

You can opt into dataset collection for research across three primary endpoints in the G0DM0D3 API:

- **`/v1/chat/completions`** – Standard chat completion interface handled in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts)
- **`/v1/ultraplinian/completions`** – Race-model completion endpoint for advanced inference
- **`/v1/consortium/completions`** – Multi-model consortium aggregation endpoint

Each route handler checks for the `contribute_to_dataset` parameter. For example, in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts) at line 231, the code reads this flag and conditionally triggers dataset entry creation via the helper functions defined in [`api/lib/dataset.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/dataset.ts).

## Implementation Details

The request body schema defining the `contribute_to_dataset` field is documented in [`API.md`](https://github.com/elder-plinius/G0DM0D3/blob/main/API.md) at line 93. When the flag is detected as `true`, the system:

1. Sanitizes the conversation by removing system messages and authentication headers
2. Creates a structured dataset entry via the helper in [`api/lib/dataset.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/dataset.ts) (line 5)
3. Queues the entry for export to the configured Hugging Face repository

Collected data becomes accessible through the `/v1/dataset/export` endpoint and related dataset management routes documented in [`API.md`](https://github.com/elder-plinius/G0DM0D3/blob/main/API.md) (lines 90-94). Note that if you are using the hosted UI at [`index.html`](https://github.com/elder-plinius/G0DM0D3/blob/main/index.html), there is no interface control for this flag; you must interact with the backend API directly to enable contribution.

## Code Examples

Below are practical implementations showing how to opt into dataset collection for research across different programming languages.

### Python with Requests

```python
import requests

BASE = "https://your-space.hf.space"
HEADERS = {
    "Authorization": "Bearer YOUR_KEY",
    "Content-Type": "application/json"
}

# ULTRAPLINIAN – race models and opt-in to dataset collection

resp = requests.post(
    f"{BASE}/v1/ultraplinian/completions",
    headers=HEADERS,
    json={
        "messages": [{"role": "user", "content": "Explain buffer overflow"}],
        "openrouter_api_key": "sk-or-v1-…",
        "tier": "fast",
        "contribute_to_dataset": True   # <-- opt-in

    },
)
print(resp.json()["dataset"]["contributed"])   # → True

```

### cURL

```bash
curl -X POST https://your-space.hf.space/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{"role":"user","content":"Write a poem about recursion"}],
        "model": "nousresearch/hermes-3-llama-3.1-70b",
        "openrouter_api_key":"sk-or-v1-…",
        "contribute_to_dataset": true
      }'

```

### Node.js with OpenAI SDK

```javascript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://your-space.hf.space/v1",
  apiKey: "YOUR_KEY",
});

const result = await client.chat.completions.create({
  model: "nousresearch/hermes-3-llama-3.1-70b",
  messages: [{ role: "user", content: "What is a hash collision?" }],
  openrouter_api_key: "sk-or-v1-…",
  contribute_to_dataset: true,   // <-- opt-in
});
console.log(result.choices[0].message.content);

```

## Summary

- **Opt-in mechanism**: Set `contribute_to_dataset: true` in your JSON payload to enable research data collection
- **Supported endpoints**: `/v1/chat/completions`, `/v1/ultraplinian/completions`, and `/v1/consortium/completions`
- **Privacy protection**: No API keys, IP addresses, or tokens are stored; only non-system conversation data is collected
- **Implementation**: Handled server-side in [`api/routes/chat.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/routes/chat.ts) and [`api/lib/dataset.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/dataset.ts), with schema defined in [`API.md`](https://github.com/elder-plinius/G0DM0D3/blob/main/API.md)
- **Access**: Collected data is exported to a Hugging Face dataset and accessible via `/v1/dataset/export`

## Frequently Asked Questions

### What data is collected when I opt into dataset collection?

When you set `contribute_to_dataset: true`, the system stores user messages, assistant responses, and pipeline metadata from the conversation. According to the source code in [`api/lib/dataset.ts`](https://github.com/elder-plinius/G0DM0D3/blob/main/api/lib/dataset.ts), the implementation explicitly filters out system messages, API keys, IP addresses, and authentication tokens to protect your privacy.

### Can I opt in using the web interface?

No. The hosted UI ([`index.html`](https://github.com/elder-plinius/G0DM0D3/blob/main/index.html)) does not provide a control for the `contribute_to_dataset` flag. To opt into dataset collection for research, you must interact directly with the backend API using tools like cURL, Python requests, or the OpenAI SDK, as shown in the code examples above.

### Where is the collected dataset stored?

Collected data is flushed to a public Hugging Face dataset repository specified by the `HF_DATASET_REPO` environment variable. You can access and export this data programmatically through the `/v1/dataset/export` endpoint, as documented in [`API.md`](https://github.com/elder-plinius/G0DM0D3/blob/main/API.md) lines 90-94.

### Is dataset collection enabled by default?

No. The feature is strictly opt-in per request. If you do not include `contribute_to_dataset: true` in your JSON payload, no conversation data is stored or exported. The default value is `false`, ensuring complete privacy unless you explicitly choose to contribute to research.