How to Opt Into Dataset Collection for Research with G0DM0D3

To opt into dataset collection for research, include "contribute_to_dataset": true in the JSON payload when calling any of G0DM0D3's primary completion endpoints.

G0DM0D3 by elder-plinius offers an optional dataset collection mechanism designed for research purposes. This feature is strictly opt-in per request, allowing you to contribute conversation data to a public Hugging Face dataset only when explicitly desired. Understanding how to opt into dataset collection for research ensures you maintain full control over your data privacy while supporting open-source AI development.

How the Dataset Collection Opt-In Works

The dataset collection system in G0DM0D3 operates on a per-request basis. When you include the Boolean field contribute_to_dataset: true in your API request, the server evaluates this flag and captures the full non-system conversation—including user messages, assistant responses, and pipeline metadata—into an in-memory buffer. This data is later flushed to a public Hugging Face dataset repository specified by the HF_DATASET_REPO environment variable.

Privacy safeguards are built into the implementation. According to the source code in api/lib/dataset.ts, the system explicitly excludes PII such as API keys, IP addresses, and authentication tokens from collection. By default, the feature is disabled, meaning no data is collected unless you explicitly set the flag to true.

Supported Endpoints

You can opt into dataset collection for research across three primary endpoints in the G0DM0D3 API:

  • /v1/chat/completions – Standard chat completion interface handled in api/routes/chat.ts
  • /v1/ultraplinian/completions – Race-model completion endpoint for advanced inference
  • /v1/consortium/completions – Multi-model consortium aggregation endpoint

Each route handler checks for the contribute_to_dataset parameter. For example, in api/routes/chat.ts at line 231, the code reads this flag and conditionally triggers dataset entry creation via the helper functions defined in api/lib/dataset.ts.

Implementation Details

The request body schema defining the contribute_to_dataset field is documented in API.md at line 93. When the flag is detected as true, the system:

  1. Sanitizes the conversation by removing system messages and authentication headers
  2. Creates a structured dataset entry via the helper in api/lib/dataset.ts (line 5)
  3. Queues the entry for export to the configured Hugging Face repository

Collected data becomes accessible through the /v1/dataset/export endpoint and related dataset management routes documented in API.md (lines 90-94). Note that if you are using the hosted UI at index.html, there is no interface control for this flag; you must interact with the backend API directly to enable contribution.

Code Examples

Below are practical implementations showing how to opt into dataset collection for research across different programming languages.

Python with Requests

import requests

BASE = "https://your-space.hf.space"
HEADERS = {
    "Authorization": "Bearer YOUR_KEY",
    "Content-Type": "application/json"
}

# ULTRAPLINIAN – race models and opt-in to dataset collection

resp = requests.post(
    f"{BASE}/v1/ultraplinian/completions",
    headers=HEADERS,
    json={
        "messages": [{"role": "user", "content": "Explain buffer overflow"}],
        "openrouter_api_key": "sk-or-v1-…",
        "tier": "fast",
        "contribute_to_dataset": True   # <-- opt-in

    },
)
print(resp.json()["dataset"]["contributed"])   # → True

cURL

curl -X POST https://your-space.hf.space/v1/chat/completions \
  -H "Authorization: Bearer YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "messages": [{"role":"user","content":"Write a poem about recursion"}],
        "model": "nousresearch/hermes-3-llama-3.1-70b",
        "openrouter_api_key":"sk-or-v1-…",
        "contribute_to_dataset": true
      }'

Node.js with OpenAI SDK

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://your-space.hf.space/v1",
  apiKey: "YOUR_KEY",
});

const result = await client.chat.completions.create({
  model: "nousresearch/hermes-3-llama-3.1-70b",
  messages: [{ role: "user", content: "What is a hash collision?" }],
  openrouter_api_key: "sk-or-v1-…",
  contribute_to_dataset: true,   // <-- opt-in
});
console.log(result.choices[0].message.content);

Summary

  • Opt-in mechanism: Set contribute_to_dataset: true in your JSON payload to enable research data collection
  • Supported endpoints: /v1/chat/completions, /v1/ultraplinian/completions, and /v1/consortium/completions
  • Privacy protection: No API keys, IP addresses, or tokens are stored; only non-system conversation data is collected
  • Implementation: Handled server-side in api/routes/chat.ts and api/lib/dataset.ts, with schema defined in API.md
  • Access: Collected data is exported to a Hugging Face dataset and accessible via /v1/dataset/export

Frequently Asked Questions

What data is collected when I opt into dataset collection?

When you set contribute_to_dataset: true, the system stores user messages, assistant responses, and pipeline metadata from the conversation. According to the source code in api/lib/dataset.ts, the implementation explicitly filters out system messages, API keys, IP addresses, and authentication tokens to protect your privacy.

Can I opt in using the web interface?

No. The hosted UI (index.html) does not provide a control for the contribute_to_dataset flag. To opt into dataset collection for research, you must interact directly with the backend API using tools like cURL, Python requests, or the OpenAI SDK, as shown in the code examples above.

Where is the collected dataset stored?

Collected data is flushed to a public Hugging Face dataset repository specified by the HF_DATASET_REPO environment variable. You can access and export this data programmatically through the /v1/dataset/export endpoint, as documented in API.md lines 90-94.

Is dataset collection enabled by default?

No. The feature is strictly opt-in per request. If you do not include contribute_to_dataset: true in your JSON payload, no conversation data is stored or exported. The default value is false, ensuring complete privacy unless you explicitly choose to contribute to research.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →