Creating Post-Training Datasets for LLMs: 5 Proven Strategies for High-Quality Data Curation
Effective post-training datasets require a five-stage pipeline: standardize chat templates, generate synthetic instructions with frontier models, enrich with chain-of-thought reasoning, deduplicate via embedding similarity, and filter with reward models to ensure quality and diversity.
The mlabonne/llm-course repository provides a comprehensive knowledge base for building high-quality post-training datasets used in supervised fine-tuning (SFT), direct preference optimization (DPO), and reinforcement learning from human feedback (RLHF). According to the Post-Training Datasets section of the README.md, successful dataset curation depends on structured storage formats, synthetic generation strategies, and rigorous quality filtering.
Standardize on a Chat-Template Schema
Before generating or processing data, define a unified conversational schema that downstream tokenizers can parse correctly.
ChatML, Alpaca, and ShareGPT are the dominant formats in modern LLM training. Each structures the interaction between system, user, and assistant roles differently. The repository emphasizes storing transformed data as JSONL, where each line represents a complete conversation turn. This format enables streaming reads during training and simplifies version control.
When converting raw dumps like ShareGPT to ChatML, map the legacy from and value fields to explicit role and content keys:
import json
from pathlib import Path
def sharegpt_to_chatml(sharegpt_path: str, output_path: str):
with open(sharegpt_path, "r", encoding="utf-8") as f:
data = json.load(f)
chatml = []
for conv in data:
messages = [{"role": "system", "content": ""}]
for turn in conv["conversations"]:
role = "assistant" if turn["from"] == "gpt" else "user"
messages.append({"role": role, "content": turn["value"]})
chatml.append(messages)
with open(output_path, "w", encoding="utf-8") as out:
for conv in chatml:
out.write(json.dumps(conv, ensure_ascii=False) + "\n")
sharegpt_to_chatml("sharegpt_raw.json", "sharegpt_chatml.jsonl")
The resulting sharegpt_chatml.jsonl follows the Hugging Face Chat Template specification, ensuring compatibility with transformers tokenizers.
Generate Synthetic Data with Frontier Models
Synthetic data generation closes the gap between pre-training corpora and task-specific instructions. Use powerful LLMs like GPT-4o to convert seed prompts into high-fidelity instruction-response pairs.
Diverse system prompts are critical for stylistic variety. Vary the system message to generate different personas, tones, and domain expertise within the same task family. For batch generation, implement parallel API calls with rate-limiting:
import os, json, time
import openai
openai.api_key = os.getenv("OPENAI_API_KEY")
def generate_instructions(seed_prompt: str, n: int = 10) -> list[dict]:
samples = []
for _ in range(n):
resp = openai.ChatCompletion.create(
model="gpt-4o",
messages=[
{"role": "system", "content": "You are a data‑generation assistant."},
{"role": "user", "content": seed_prompt},
],
temperature=0.9,
)
instruction = resp.choices[0].message["content"]
samples.append({"instruction": instruction.strip()})
time.sleep(0.5)
return samples
seed = "Create a concise instruction to summarize a news article in three bullet points."
synthetic = generate_instructions(seed, n=5)
with open("synthetic_instructions.jsonl", "w", encoding="utf-8") as f:
for s in synthetic:
f.write(json.dumps(s, ensure_ascii=False) + "\n")
For a managed UI-based approach, the repository references the Argilla Synthetic Data Generator on Hugging Face Spaces, which provides a no-code interface for seed-to-dataset pipelines.
Enrich Data with Enhancement Techniques
Raw synthetic generations require augmentation to maximize training signal. Verified outputs ensure correctness—execute generated code through unit tests or verify mathematical answers before inclusion.
For preference alignment methods like DPO, store chosen and rejected sample pairs within the same JSON object. This structure enables direct loading into alignment frameworks like TRL or OpenRLHF.
Enhance reasoning capabilities by injecting Chain-of-Thought (CoT) prompts:
def cot_answer(instruction: str) -> str:
prompt = f"""You are an expert AI assistant. Answer the following instruction step‑by‑step.
Instruction:
{instruction}
Answer (include reasoning):"""
resp = openai.ChatCompletion.create(
model="gpt-4o",
messages=[{"role": "user", "content": prompt}],
temperature=0.7,
)
return resp.choices[0].message["content"]
The repository also cites Auto-Evol methodologies, which use Branch-Solve-Merge strategies and persona-based evolution to iteratively increase dataset complexity and coverage.
Implement Multi-Stage Quality Filtering
Quality filtering transforms large synthetic dumps into clean, trainable corpora. Apply filters in cascading order of computational cost:
- Rule-based filters: Use regex to remove profanity, off-topic content, and length outliers (extremely short or long samples).
- Deduplication: Eliminate exact and near-duplicate examples using MinHash or embedding-based similarity. The repository recommends semhash from MinishLab for efficient semantic deduplication.
- N-gram decontamination: Compute n-gram overlap with held-out benchmark test sets to prevent data leakage.
- Reward model scoring: Fine-tune a lightweight reward model to score samples for relevance, factuality, and style.
Implement embedding-based deduplication with sentence-transformers:
from sentence_transformers import SentenceTransformer, util
import json
model = SentenceTransformer("all-MiniLM-L6-v2")
def deduplicate(jsonl_path: str, threshold: float = 0.9):
with open(jsonl_path, "r", encoding="utf-8") as f:
lines = [json.loads(l) for l in f]
embeddings = model.encode(
[" ".join(m["content"] for m in conv) for conv in lines],
convert_to_tensor=True,
)
keep = []
for i, emb in enumerate(embeddings):
if any(util.cos_sim(emb, embeddings[j]) > threshold for j in keep):
continue
keep.append(i)
with open(jsonl_path.replace(".jsonl", "_dedup.jsonl"), "w", encoding="utf-8") as out:
for idx in keep:
out.write(json.dumps(lines[idx], ensure_ascii=False) + "\n")
deduplicate("sharegpt_chatml.jsonl", threshold=0.92)
For final quality scoring, use TRL to train or load a reward model:
from trl import RewardModelTrainer, RewardConfig
from datasets import load_dataset
ds = load_dataset("json", data_files="rated_responses.jsonl")["train"]
config = RewardConfig(
model_name="OpenAssistant/reward-model-deberta-v3-base",
learning_rate=5e-5,
per_device_train_batch_size=8,
)
trainer = RewardModelTrainer(
model_name_or_path=config.model_name,
reward_config=config,
train_dataset=ds,
)
trainer.train()
After training, inference the reward model on your dataset to assign quality scores, then filter entries below your chosen threshold.
Assemble the End-to-End Pipeline
Combine these strategies into a reproducible workflow: Raw seed corpora → Cleaning/deduplication → Chat-template conversion → Synthetic generation → Data enhancement → Quality filtering → Final JSONL dataset.
This architecture ensures post-training data is structured, high-quality, and diverse. As documented in README.md, structured JSONL storage with chat-template awareness enables seamless integration with fine-tuning frameworks like TRL, Axolotl, and Llama-Factory.
Summary
- Store conversations in JSONL format using ChatML or ShareGPT schemas for tokenizer compatibility.
- Generate diverse synthetic instructions using GPT-4o with varied system prompts and batch API calls.
- Enrich datasets with Chain-of-Thought reasoning and chosen/rejected pairs for preference alignment.
- Filter through rule-based cleaning, MinHash deduplication, n-gram decontamination, and reward model scoring.
- Reference the
mlabonne/llm-coursePost-Training Datasets section for tool recommendations like Argilla and semhash.
Frequently Asked Questions
What is the best format for storing post-training datasets?
JSONL (JSON Lines) is the optimal format because it enables streaming reads during training and simplifies git diffs. Each line should contain a complete conversation following ChatML, ShareGPT, or Alpaca schemas with explicit role and content keys. This structure ensures compatibility with Hugging Face tokenizers and chat templating systems.
How much synthetic data versus real data should I use?
The ideal ratio depends on your domain, but the repository suggests synthetic data is most effective for instruction-following and reasoning tasks where real annotated data is scarce. Use frontier models like GPT-4o for seed generation, then apply verification layers (code execution, math checks) to ensure synthetic samples meet quality bars comparable to human-written examples.
Which tools handle deduplication at scale?
For semantic deduplication, use semhash (MinishLab) which implements optimized MinHash algorithms for near-duplicate detection. For embedding-based similarity, sentence-transformers with cosine similarity thresholds (typically 0.90–0.95) provides fine-grained control. Both approaches prevent benchmark contamination and reduce training redundancy.
Do I need a large reward model for quality filtering?
No. Lightweight models like OpenAssistant/reward-model-deberta-v3-base (approximately 300M parameters) are sufficient for initial quality scoring. Fine-tune these with TRL on a small set of human ratings, then apply them as filters before expensive training runs. This approach balances computational cost with filtering efficacy.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →