How to Debug Issues with Loading Prompt Datasets in Heretic
To debug prompt dataset loading issues in Heretic, verify your DatasetSpecification path, split syntax, and column name using the load_prompts function in src/heretic/utils.py, and isolate failures by running the loading logic interactively with explicit error handling.
When working with the p-e-w/heretic repository, you may encounter errors while loading prompt datasets from Hugging Face or local sources. Understanding how to debug issues with loading prompt datasets in Heretic requires familiarity with the load_prompts helper and the DatasetSpecification configuration structure.
Understanding the Prompt Loading Pipeline in Heretic
Heretic loads prompts via the load_prompts function defined in src/heretic/utils.py. This function accepts a global Settings object and a DatasetSpecification (defined in src/heretic/config.py) to resolve, load, and format prompt data.
The loading process follows these eight distinct steps:
- Resolve the dataset source – The
pathis extracted fromspecification.dataset(line 70), which may be a Hugging Face identifier (e.g.,mlabonne/harmless_alpaca) or a local filesystem path. - Select the split – The
split_stris taken fromspecification.split(line 71), using Hugging Face datasets syntax liketrain[:400]ortest. - Load the dataset – The code branches (lines 76-105) to handle three cases: saved dataset directories via
load_from_disk, local folders viaload_datasetwithverification_mode=NO_CHECKS, or remote identifiers viaload_dataset. - Apply split instructions – For locally saved datasets,
ReadInstruction.from_specconverts split strings into absolute indices (lines 84-93). - Extract the prompt column – Raw text is extracted via
list(dataset[specification.column])(line 107). - Apply prefix/suffix – Optional formatting is added (lines 109-113) without modifying the source dataset.
- Determine system prompt – The function checks for a specification-level override or falls back to
settings.system_prompt(lines 115-119). - Build
Promptobjects – Returns a list ofPromptdataclasses (lines 122-127) consumed bymain.py(lines 324-330).
Common Failure Modes and Diagnostic Steps
FileNotFoundError and Path Resolution
If you encounter a FileNotFoundError stating the path does not exist, verify that specification.dataset points to a valid Hugging Face repository ID or an absolute local path.
Debugging steps:
- Insert a print statement before loading:
print(f"Loading dataset from {specification.dataset}") - Verify local paths with
ls <path>or check the repository name on the Hugging Face Hub
DatasetDict Not Supported Errors
Heretic expects a single dataset split, but load_from_disk may return a DatasetDict containing multiple splits (train, test, validation).
Solution: Ensure you saved the dataset after selecting a specific split, or manually select the split after loading:
dataset = load_from_disk(path)
dataset = dataset["train"] # Select specific split
Empty Prompt Lists and Column Mismatches
An empty list usually indicates a wrong column name or a split containing zero rows (e.g., train[:0]).
Debugging steps:
- Print dataset info after loading:
print(dataset)to see available columns and length - Confirm
specification.columnmatches one of the printed column names - Verify the split string yields data:
print(len(dataset))
Split Syntax Errors
Invalid split expressions like train[:a] or missing brackets cause immediate failures.
Fix: Use valid Hugging Face datasets split syntax. Test in a REPL:
from datasets import load_dataset
load_dataset("mlabonne/harmless_alpaca", split="train[:400]")
Network and Authentication Failures
Remote datasets may fail due to missing internet connectivity or required authentication.
Debugging steps:
- Run with
download_mode=DownloadMode.FORCE_REDOWNLOADto see detailed errors - Check if the dataset requires a Hugging Face token:
huggingface-cli login
Memory and Performance Issues
Large datasets may cause out-of-memory errors, especially with verification_mode=NO_CHECKS loading everything into memory.
Solution: Reduce the split size or enable streaming mode:
load_dataset(specification.dataset, split=specification.split, streaming=True)
Interactive Debugging Example
Use this standalone script to isolate loading issues outside of the full Heretic pipeline:
from heretic.utils import load_prompts
from heretic.config import Settings
# Load settings (reads config.toml, env vars, CLI, etc.)
settings = Settings()
# Inspect the good-prompt specification
spec = settings.good_prompts
print("Dataset spec →", spec)
try:
prompts = load_prompts(settings, spec)
print(f"✅ Loaded {len(prompts)} prompts")
# Show a few examples
for p in prompts[:3]:
print("- System:", p.system)
print(" User:", p.user[:80], "…")
except Exception as exc:
print("❌ Error while loading prompts:", exc)
raise
What this reveals:
- Exact spec values – Confirm the dataset ID, split, column, and any prefix/suffix defined in
config.toml - Successful loading – Verify the split produced rows and the column exists
- Prompt samples – Spot malformed prefixes, missing system prompts, or unexpected text formatting
Key Source Files for Reference
src/heretic/utils.py– Contains theload_promptsfunction,Promptdataclass, and dataset loading logic (lines 70-127)src/heretic/config.py– DefinesDatasetSpecificationand the globalSettingsclass that configures prompt datasetssrc/heretic/main.py– Entry point that invokesload_promptsfor both good and bad prompt sets (lines 324-330)README.md– High-level documentation of Heretic's workflow and dataset configuration options
Summary
- Verify the three pillars of
DatasetSpecification: the datasetpath(Hugging Face ID or local directory), thesplitsyntax (e.g.,train[:400]), and thecolumnname containing prompt text. - Use interactive debugging by importing
load_promptsandSettingsdirectly to isolate errors outside the full pipeline. - Handle
DatasetDicterrors by ensuring you load a single split, not a dictionary of splits. - Check prefix/suffix configurations in your
config.tomlif prompts appear malformed after loading. - Monitor memory usage with large datasets by using streaming mode or smaller split ranges.
Frequently Asked Questions
Why does Heretic raise a FileNotFoundError for a valid Hugging Face dataset ID?
Heretic treats the dataset field in your DatasetSpecification as a filesystem path first. If the string does not match a local directory, it attempts to load from Hugging Face. However, if you have a local folder with the same name as a Hugging Face repo, or if the repo ID contains characters that look like relative paths, the resolution may fail. Always verify the exact string in your config.toml and test with print(settings.good_prompts.dataset) before loading.
How do I fix the "DatasetDict not supported" error when loading prompts?
This error occurs when load_from_disk returns a DatasetDict (containing multiple splits like train and test) rather than a single Dataset object. Heretic's load_prompts function expects a single split. To fix this, either save only the specific split you need using dataset["train"].save_to_disk(path), or modify your local copy of the dataset to select a split before Heretic processes it.
Why is my prompt list empty even though the dataset loads successfully?
An empty list indicates that the dataset loaded but contains zero rows for your specified configuration. Check three things: first, verify your split syntax (e.g., train[:400] vs train[400:]); second, confirm the column name in your DatasetSpecification exactly matches a column in the dataset (check available columns with print(dataset.column_names)); third, ensure you are not applying a prefix or suffix filter that excludes all entries.
Can I use streaming mode to debug large datasets without running out of memory?
Yes, though Heretic's default load_prompts implementation loads the full split into memory to extract the prompt column. For debugging large datasets, temporarily modify the loading call in src/heretic/utils.py to pass streaming=True to load_dataset, or reduce your split size to a small subset (e.g., train[:10]) while validating your configuration. Once verified, you can process the full dataset with standard loading parameters.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →