How to Run DFlash Benchmarks on Custom Datasets: A Complete Guide

To run DFlash benchmarks on custom datasets, register a new dataset entry in the DATASETS dictionary inside dflash/benchmark.py with a Hugging Face loader and a format lambda, or provide a pre-cached JSONL file in the cache/ directory following the {"turns": [...]} schema.

The z-lab/dflash repository provides a high-performance speculative decoding framework. Its benchmark driver, located in dflash/benchmark.py, is designed to evaluate speedups on standardized datasets but can be extended to evaluate custom workloads. Understanding the dataset registration and caching pipeline is essential for integrating your own data.

Understanding the DFlash Benchmark Architecture

The benchmark system relies on a central registry and a cache-based loading pipeline to minimize network overhead during repeated runs.

The Dataset Registration System

At the core of the system is the DATASETS dictionary defined at lines 28–55 of dflash/benchmark.py. This mapping associates a string name with a configuration dict containing:

  • load_args: Tuple passed to datasets.load_dataset()
  • load_kwargs: Additional keyword arguments for the loader
  • format: A lambda that transforms a raw record into a list of user-turn strings

The benchmark driver validates the --dataset CLI argument against the keys in this dictionary. If the name is not present, load_and_process_dataset (lines 84–94) raises a KeyError.

The Cache and Loading Pipeline

To avoid repeated downloads, the _prepare_dataset function (lines 58–80) checks for a pre-existing JSONL file in CACHE_DIR (default: cache/). If the file exists, it skips the download and loads directly from disk. If not, it downloads the Hugging Face dataset, applies the format lambda to each row, and writes a line-delimited JSON file where each line is a JSON object with a "turns" key.

Method 1: Registering a Custom Hugging Face Dataset

The most flexible approach is to extend the DATASETS dictionary with your own Hugging Face dataset or a local dataset exposed via the datasets library.

First, modify dflash/benchmark.py to add your entry:


# Add near the existing DATASETS definition (line 28)

CUSTOM_DATASETS = {
    "my_medical_qa": {
        "load_args": ("my-org/medical-qa",),
        "load_kwargs": {"split": "validation"},
        # Transform each record into a list of user turns

        "format": lambda x: [x["question"], x["context"]],
    },
    "local_json": {
        "load_args": ("json",),
        "load_kwargs": {"data_files": "data/my_data.json"},
        "format": lambda x: x["messages"],
    },
}
DATASETS.update(CUSTOM_DATASETS)

Then run the benchmark using your registered name:

python -m dflash.benchmark \
    --backend transformers \
    --model meta-llama/Meta-Llama-3.1-8B-Instruct \
    --draft-model meta-llama/Meta-Llama-3.1-8B-Instruct \
    --dataset my_medical_qa \
    --max-samples 200 \
    --block-size 8 \
    --max-new-tokens 512

The format lambda must return a list of strings, where each string represents a user turn in the conversation. This matches the schema used by the built-in GSM8K and MT-Bench datasets.

Method 2: Using a Pre-Cached JSONL File

If you prefer not to modify the source code or have data that is not available on Hugging Face, you can provide a pre-cached JSONL file directly.

Create a file at cache/<dataset_name>.jsonl where <dataset_name> exists as a key in the DATASETS dictionary (or add a dummy entry). Each line must be a valid JSON object with a "turns" key containing a list of strings:

{"turns": ["Explain the theory of relativity in simple terms.", "The theory of relativity, developed by Albert Einstein..."]}
{"turns": ["Write a Python function to calculate Fibonacci numbers.", "def fibonacci(n):\n    if n <= 1:\n        return n\n    return fibonacci(n-1) + fibonacci(n-2)"]}
{"turns": ["What is the capital of France?", "The capital of France is Paris."]}

Place this file in the cache directory:

mkdir -p cache
cp my_custom_data.jsonl cache/gsm8k.jsonl

When you run the benchmark, the _prepare_dataset function detects the existing cache file and skips the Hugging Face download:

python -m dflash.benchmark \
    --backend mlx \
    --model meta-llama/Meta-Llama-3.1-8B-Instruct \
    --draft-model meta-llama/Meta-Llama-3.1-8B-Instruct \
    --dataset gsm8k \
    --max-samples 100

Key Implementation Details

Understanding the internal flow helps debug issues when integrating custom data.

Dataset Validation and Loading

The entry point load_and_process_dataset (lines 84–94) performs a strict lookup:

def load_and_process_dataset(name: str) -> List[Dict]:
    if name not in DATASETS:
        raise ValueError(f"Unknown dataset: {name}")
    dataset_path = _prepare_dataset(name)
    with open(dataset_path, 'r') as f:
        return [json.loads(line) for line in f]

If your custom dataset name is not registered, the benchmark exits immediately with a ValueError.

The Formatting Pipeline

The _prepare_dataset function (lines 58–80) handles the transformation. It checks for the cache file at CACHE_DIR / f"{name}.jsonl". If missing, it loads the Hugging Face dataset using the load_args and load_kwargs, then applies the format lambda to each record. The result is written as JSONL with the schema {"turns": formatted_list}.

This design means the benchmark loop itself (lines 97+) is agnostic to the data source; it only expects a list of dictionaries with a "turns" key.

Summary

  • DFlash benchmarks require registered datasets: The DATASETS dictionary in dflash/benchmark.py (lines 28–55) maps names to Hugging Face loaders and formatting functions.
  • Two integration paths exist: Register a new dataset entry with a format lambda, or provide a pre-cached JSONL file in cache/<name>.jsonl with the {"turns": [...]} schema.
  • Validation is strict: The load_and_process_dataset function raises a ValueError if the dataset name is not found in the registry.
  • Cache bypasses download: If a JSONL file exists in the cache directory, _prepare_dataset skips the Hugging Face download and loads directly from disk.

Frequently Asked Questions

What format does DFlash expect for custom datasets?

DFlash expects either a registered Hugging Face dataset with a format lambda that returns a list of strings, or a JSONL file where each line is a JSON object containing a "turns" key with a list of user-turn strings. The benchmark code in dflash/benchmark.py loads these records in load_and_process_dataset and iterates over them directly.

Can I use local files without Hugging Face?

Yes. You can use local files by either adding a dataset entry that uses datasets.load_dataset with a local JSON or JSONL file path in load_kwargs, or by creating a pre-cached JSONL file in the cache/ directory with the correct schema. The benchmark will read from the cache without attempting a Hugging Face download.

Where does DFlash store cached datasets?

DFlash stores cached datasets in the cache/ directory relative to the working directory, specifically at cache/<dataset_name>.jsonl. This path is constructed in the _prepare_dataset function (lines 58–80) using CACHE_DIR / f"{name}.jsonl". If this file exists, the benchmark skips the download step.

How do I verify my custom dataset is loaded correctly?

To verify loading, check that your dataset name appears in the DATASETS dictionary keys and that the format lambda returns a list of strings for each record. You can also inspect the generated cache file at cache/<name>.jsonl to ensure it contains valid JSON objects with the "turns" key. Running the benchmark with --max-samples 1 and verbose logging will confirm the dataset loads without raising a ValueError in load_and_process_dataset.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →