CreativeMath Output Formats for Generation and Evaluation: JSON Schema Reference

CreativeMath produces UTF-8 encoded JSON arrays where the generation phase outputs simple objects containing problem IDs and LLM responses, and the evaluation phase enriches these objects with multi-model correctness checks and two-tier novelty assessments including majority-vote final decisions.

The junyiye/creativemath repository implements a two-stage pipeline for generating and evaluating novel mathematical solutions using large language models. Understanding the exact output formats for both generation and evaluation is essential for parsing results, conducting downstream analysis, and integrating the tool into broader research workflows. Both phases serialize data as structured JSON arrays written to disk in UTF-8 encoding.

Generation Output Format in CreativeMath

The generation phase, implemented in src/generation.py, produces a JSON array where each element represents a single experiment run. According to the source code at lines 66-70, the script writes this array to <generation_dir>/<model_name>.json after iterating through the dataset and calling the target LLM via prompts built with load_novel_solution_generation_prompt.

Core Schema Fields

Each object in the generation output contains four required fields:

  • problem_id (int): Unique identifier for the mathematical problem.
  • k (int): Number of reference solutions provided in the prompt context.
  • n (int): Total number of available reference solutions for that problem.
  • response (string): The novel solution generated by the LLM.
{
  "problem_id": 12,
  "k": 3,
  "n": 5,
  "response": "The novel solution is ..."
}

Evaluation Output Format in CreativeMath

The evaluation phase, defined in src/evaluation.py, loads the generation JSON and augments each entry with three nested assessment objects. As implemented at lines 65-70, the enriched data is written to <evaluation_dir>/<model_name>.json, preserving all original generation fields while adding multi-model voting results.

Correctness Assessment Structure

The correctness field maps evaluator model names (e.g., claude-3-opus, gpt-4) to "YES" or "NO" strings, plus a final_decision key representing the majority vote across evaluators.

Coarse-Grained Novelty Structure

The coarse_grained_novelty object follows an identical structure, capturing whether the solution is novel at a high level compared to reference solutions, with per-evaluator votes and a majority final_decision.

Fine-Grained Novelty Structure

The fine_grained_novelty field provides granular novelty assessment using the same voting schema, checking for specific methodological differences at the step-by-step level.

{
  "problem_id": 12,
  "k": 3,
  "n": 5,
  "response": "The novel solution is ...",
  "correctness": {
    "claude-3-opus": "YES",
    "gemini-1.5-pro": "YES",
    "gpt-4": "NO",
    "final_decision": "YES"
  },
  "coarse_grained_novelty": {
    "claude-3-opus": "YES",
    "gemini-1.5-pro": "NO",
    "gpt-4": "YES",
    "final_decision": "YES"
  },
  "fine_grained_novelty": {
    "claude-3-opus": "YES",
    "gemini-1.5-pro": "YES",
    "gpt-4": "NO",
    "final_decision": "YES"
  }
}

Implementation Architecture

The output formats are produced by distinct but coordinated components in the repository. The src/generation.py module handles the initial LLM interaction and JSON serialization, while src/evaluation.py orchestrates the three-stage assessment pipeline (correctness, coarse novelty, fine novelty). Supporting infrastructure includes src/prompts/prompts.py for prompt construction, src/models/api_models.py and src/models/local_models.py for evaluator LLM interfaces, and src/utils.py for JSON I/O operations.

Summary

  • CreativeMath uses UTF-8 JSON arrays for all persistent storage.
  • Generation outputs contain problem_id, k, n, and response fields written to <generation_dir>/<model_name>.json.
  • Evaluation enriches generation outputs with correctness, coarse_grained_novelty, and fine_grained_novelty objects, each containing per-model votes and a majority final_decision.
  • All assessments use binary "YES"/"NO" values aggregated through majority voting.

Frequently Asked Questions

What file format does CreativeMath use for output?

CreativeMath serializes all data as UTF-8 encoded JSON arrays. Both generation and evaluation phases write these arrays to disk using standard Python json module operations, making the outputs compatible with any JSON parser.

What fields are added during the evaluation phase?

The evaluation phase adds three nested dictionaries to each generation record: correctness for mathematical accuracy, coarse_grained_novelty for high-level methodological uniqueness, and fine_grained_novelty for step-by-step distinctiveness. Each contains individual evaluator verdicts and a final_decision majority vote.

How is the final decision calculated in evaluation?

The final_decision field in each assessment object represents a majority vote across all evaluator models. For example, if two out of three evaluators mark a solution as "YES" for correctness, the final_decision is set to "YES".

Where are the output files saved in CreativeMath?

Generation outputs are saved to <generation_dir>/<model_name>.json as defined in src/generation.py lines 66-70. Evaluation outputs are written to <evaluation_dir>/<model_name>.json as implemented in src/evaluation.py lines 65-70.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →