# Creating Post-Training Datasets for LLMs: 5 Proven Strategies for High-Quality Data Curation

> Discover 5 proven strategies for creating high-quality post-training datasets for LLMs. Learn to standardize, generate, enrich, deduplicate, and filter data effectively.

- Repository: [Maxime Labonne/llm-course](https://github.com/mlabonne/llm-course)
- Tags: best-practices
- Published: 2026-03-01

---

**Effective post-training datasets require a five-stage pipeline: standardize chat templates, generate synthetic instructions with frontier models, enrich with chain-of-thought reasoning, deduplicate via embedding similarity, and filter with reward models to ensure quality and diversity.**

The **mlabonne/llm-course** repository provides a comprehensive knowledge base for building high-quality post-training datasets used in supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning from human feedback (RLHF). According to the [Post-Training Datasets section](https://github.com/mlabonne/llm-course/blob/main/README.md#post‑training-datasets) of the [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md), successful dataset curation depends on structured storage formats, synthetic generation strategies, and rigorous quality filtering.

## Standardize on a Chat-Template Schema

Before generating or processing data, define a unified conversational schema that downstream tokenizers can parse correctly.

**ChatML**, **Alpaca**, and **ShareGPT** are the dominant formats in modern LLM training. Each structures the interaction between system, user, and assistant roles differently. The repository emphasizes storing transformed data as **JSONL**, where each line represents a complete conversation turn. This format enables streaming reads during training and simplifies version control.

When converting raw dumps like ShareGPT to ChatML, map the legacy `from` and `value` fields to explicit `role` and `content` keys:

```python
import json
from pathlib import Path

def sharegpt_to_chatml(sharegpt_path: str, output_path: str):
    with open(sharegpt_path, "r", encoding="utf-8") as f:
        data = json.load(f)

    chatml = []
    for conv in data:
        messages = [{"role": "system", "content": ""}]
        for turn in conv["conversations"]:
            role = "assistant" if turn["from"] == "gpt" else "user"
            messages.append({"role": role, "content": turn["value"]})
        chatml.append(messages)

    with open(output_path, "w", encoding="utf-8") as out:
        for conv in chatml:
            out.write(json.dumps(conv, ensure_ascii=False) + "\n")

sharegpt_to_chatml("sharegpt_raw.json", "sharegpt_chatml.jsonl")

```

The resulting `sharegpt_chatml.jsonl` follows the [Hugging Face Chat Template](https://huggingface.co/docs/transformers/main/en/chat_templating) specification, ensuring compatibility with `transformers` tokenizers.

## Generate Synthetic Data with Frontier Models

Synthetic data generation closes the gap between pre-training corpora and task-specific instructions. Use powerful LLMs like **GPT-4o** to convert seed prompts into high-fidelity instruction-response pairs.

**Diverse system prompts** are critical for stylistic variety. Vary the system message to generate different personas, tones, and domain expertise within the same task family. For batch generation, implement parallel API calls with rate-limiting:

```python
import os, json, time
import openai

openai.api_key = os.getenv("OPENAI_API_KEY")

def generate_instructions(seed_prompt: str, n: int = 10) -> list[dict]:
    samples = []
    for _ in range(n):
        resp = openai.ChatCompletion.create(
            model="gpt-4o",
            messages=[
                {"role": "system", "content": "You are a data‑generation assistant."},
                {"role": "user", "content": seed_prompt},
            ],
            temperature=0.9,
        )
        instruction = resp.choices[0].message["content"]
        samples.append({"instruction": instruction.strip()})
        time.sleep(0.5)
    return samples

seed = "Create a concise instruction to summarize a news article in three bullet points."
synthetic = generate_instructions(seed, n=5)

with open("synthetic_instructions.jsonl", "w", encoding="utf-8") as f:
    for s in synthetic:
        f.write(json.dumps(s, ensure_ascii=False) + "\n")

```

For a managed UI-based approach, the repository references the [Argilla Synthetic Data Generator](https://huggingface.co/spaces/argilla/synthetic-data-generator) on Hugging Face Spaces, which provides a no-code interface for seed-to-dataset pipelines.

## Enrich Data with Enhancement Techniques

Raw synthetic generations require augmentation to maximize training signal. **Verified outputs** ensure correctness—execute generated code through unit tests or verify mathematical answers before inclusion.

For preference alignment methods like DPO, store **chosen** and **rejected** sample pairs within the same JSON object. This structure enables direct loading into alignment frameworks like **TRL** or **OpenRLHF**.

Enhance reasoning capabilities by injecting **Chain-of-Thought (CoT)** prompts:

```python
def cot_answer(instruction: str) -> str:
    prompt = f"""You are an expert AI assistant. Answer the following instruction step‑by‑step.

Instruction:
{instruction}

Answer (include reasoning):"""
    resp = openai.ChatCompletion.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": prompt}],
        temperature=0.7,
    )
    return resp.choices[0].message["content"]

```

The repository also cites **Auto-Evol** methodologies, which use **Branch-Solve-Merge** strategies and persona-based evolution to iteratively increase dataset complexity and coverage.

## Implement Multi-Stage Quality Filtering

Quality filtering transforms large synthetic dumps into clean, trainable corpora. Apply filters in cascading order of computational cost:

1. **Rule-based filters**: Use regex to remove profanity, off-topic content, and length outliers (extremely short or long samples).
2. **Deduplication**: Eliminate exact and near-duplicate examples using **MinHash** or embedding-based similarity. The repository recommends **semhash** from MinishLab for efficient semantic deduplication.
3. **N-gram decontamination**: Compute n-gram overlap with held-out benchmark test sets to prevent data leakage.
4. **Reward model scoring**: Fine-tune a lightweight reward model to score samples for relevance, factuality, and style.

Implement embedding-based deduplication with `sentence-transformers`:

```python
from sentence_transformers import SentenceTransformer, util
import json

model = SentenceTransformer("all-MiniLM-L6-v2")

def deduplicate(jsonl_path: str, threshold: float = 0.9):
    with open(jsonl_path, "r", encoding="utf-8") as f:
        lines = [json.loads(l) for l in f]

    embeddings = model.encode(
        [" ".join(m["content"] for m in conv) for conv in lines],
        convert_to_tensor=True,
    )
    keep = []
    for i, emb in enumerate(embeddings):
        if any(util.cos_sim(emb, embeddings[j]) > threshold for j in keep):
            continue
        keep.append(i)

    with open(jsonl_path.replace(".jsonl", "_dedup.jsonl"), "w", encoding="utf-8") as out:
        for idx in keep:
            out.write(json.dumps(lines[idx], ensure_ascii=False) + "\n")

deduplicate("sharegpt_chatml.jsonl", threshold=0.92)

```

For final quality scoring, use **TRL** to train or load a reward model:

```python
from trl import RewardModelTrainer, RewardConfig
from datasets import load_dataset

ds = load_dataset("json", data_files="rated_responses.jsonl")["train"]

config = RewardConfig(
    model_name="OpenAssistant/reward-model-deberta-v3-base",
    learning_rate=5e-5,
    per_device_train_batch_size=8,
)

trainer = RewardModelTrainer(
    model_name_or_path=config.model_name,
    reward_config=config,
    train_dataset=ds,
)

trainer.train()

```

After training, inference the reward model on your dataset to assign quality scores, then filter entries below your chosen threshold.

## Assemble the End-to-End Pipeline

Combine these strategies into a reproducible workflow: **Raw seed corpora → Cleaning/deduplication → Chat-template conversion → Synthetic generation → Data enhancement → Quality filtering → Final JSONL dataset**.

This architecture ensures post-training data is **structured**, **high-quality**, and **diverse**. As documented in [`README.md`](https://github.com/mlabonne/llm-course/blob/main/README.md), structured JSONL storage with chat-template awareness enables seamless integration with fine-tuning frameworks like **TRL**, **Axolotl**, and **Llama-Factory**.

## Summary

- Store conversations in **JSONL** format using **ChatML** or **ShareGPT** schemas for tokenizer compatibility.
- Generate diverse synthetic instructions using **GPT-4o** with varied system prompts and batch API calls.
- Enrich datasets with **Chain-of-Thought** reasoning and **chosen/rejected** pairs for preference alignment.
- Filter through **rule-based** cleaning, **MinHash** deduplication, n-gram decontamination, and **reward model** scoring.
- Reference the `mlabonne/llm-course` [Post-Training Datasets section](https://github.com/mlabonne/llm-course/blob/main/README.md#post‑training-datasets) for tool recommendations like Argilla and semhash.

## Frequently Asked Questions

### What is the best format for storing post-training datasets?

**JSONL (JSON Lines)** is the optimal format because it enables streaming reads during training and simplifies git diffs. Each line should contain a complete conversation following **ChatML**, **ShareGPT**, or **Alpaca** schemas with explicit `role` and `content` keys. This structure ensures compatibility with Hugging Face tokenizers and chat templating systems.

### How much synthetic data versus real data should I use?

The ideal ratio depends on your domain, but the repository suggests synthetic data is most effective for **instruction-following** and **reasoning** tasks where real annotated data is scarce. Use frontier models like GPT-4o for seed generation, then apply verification layers (code execution, math checks) to ensure synthetic samples meet quality bars comparable to human-written examples.

### Which tools handle deduplication at scale?

For semantic deduplication, use **semhash** (MinishLab) which implements optimized MinHash algorithms for near-duplicate detection. For embedding-based similarity, **sentence-transformers** with cosine similarity thresholds (typically 0.90–0.95) provides fine-grained control. Both approaches prevent benchmark contamination and reduce training redundancy.

### Do I need a large reward model for quality filtering?

No. Lightweight models like `OpenAssistant/reward-model-deberta-v3-base` (approximately 300M parameters) are sufficient for initial quality scoring. Fine-tune these with **TRL** on a small set of human ratings, then apply them as filters before expensive training runs. This approach balances computational cost with filtering efficacy.