What Types of Mathematical Problems Are in the CreativeMath Dataset?
The CreativeMath dataset contains multiple-choice mathematical problems from standardized competitions like AMC 8 and AMC 10, spanning five major categories including Arithmetic & Algebra, Geometry, Number Theory & Combinatorics, Probability & Statistics, and Logic & Word Problems, with difficulty levels ranging from middle-school to Olympiad-level.
The CreativeMath repository (junyiye/creativemath) provides a curated benchmark of competition mathematics designed to test large language model creativity. This collection of mathematical problems serves as a standardized evaluation suite for novel solution generation, covering diverse domains from basic algebra to advanced combinatorics.
Overview of the CreativeMath Dataset
The dataset is stored in data/subset.json and contains problems sourced from major standardized competitions including AMC 8, AMC 10, AMC 10 A, and AMC 10 B. Each problem follows a standardized JSON schema that includes metadata such as competition labels, unique identifiers, difficulty scores, and human-written solutions.
Problems are classified into three difficulty tiers:
- Level 1: Middle-school difficulty
- Level 2: High-school competition level
- Level 3: Olympiad-level difficulty
Categories of Mathematical Problems in CreativeMath
The dataset encompasses five primary problem families, each targeting specific mathematical reasoning skills.
Arithmetic and Algebra
This category includes problems involving linear equations, quadratic equations, ratios, fractions, exponents, and inequalities. Advanced topics cover integer division techniques and the Lifting-the-Exponent Lemma. These problems appear frequently in AMC 8 and AMC 10 competitions.
Geometry
Geometric problems span plane geometry (triangles, circles, polygons, area calculations) and solid geometry (tetrahedra, octahedra). The dataset includes coordinate geometry challenges and transformation problems involving reflections and rotations. Trigonometric relationships and spatial reasoning tasks are also represented.
Number Theory and Combinatorics
This family covers divisibility rules, prime factorization, LCM/GCD calculations, and modular arithmetic. Combinatorial problems include counting selections, permutations, and applications of the pigeonhole principle. The dataset also contains combinatorial identity proofs and probability counting scenarios.
Probability and Statistics
Problems in this category involve simple probability calculations, expected value computations, and counting outcomes in finite sample spaces. Basic statistical reasoning and data interpretation tasks suitable for competition mathematics are included.
Logic and Word Problems
This category encompasses puzzle-style reasoning, sequence interpretation, and problems requiring careful parsing of textual constraints. These problems test logical deduction and systematic problem-solving approaches without requiring advanced computational techniques.
Dataset Structure and Metadata
Each mathematical problem in the dataset follows a rigorous JSON schema defined in data/subset.json. The schema includes:
- competition: Source competition label (e.g., "AMC_8", "AMC_10")
- competition_id: Unique identifier for the competition instance
- problem_id: Unique problem identifier
- difficulty: Numeric score (typically 1.0 to 3.0)
- problem: Raw problem statement including LaTeX formatting
- solutions: Array of human-written solution strings
This structured format enables systematic filtering and analysis of mathematical problems by type, difficulty, or source competition.
Working with the Dataset: Practical Code Examples
The following Python snippets demonstrate how to load and analyze the mathematical problems in CreativeMath.
Loading the Dataset
import json
from pathlib import Path
# Path to the JSON file in the repository
DATA_PATH = Path(__file__).parent.parent / "data" / "subset.json"
with DATA_PATH.open(encoding="utf-8") as f:
problems = json.load(f) # a list of dicts
Analyzing Competition Distribution
from collections import Counter
comp_counter = Counter(p["competition"] for p in problems)
print(comp_counter)
# Example output: Counter({'AMC_8': 70, 'AMC_10': 70})
Filtering by Difficulty Level
difficulty_set = sorted({p["difficulty"] for p in problems})
print("Difficulty levels:", difficulty_set)
# → Difficulty levels: [1, 1.5, 2, 2.5, 3]
Identifying Geometry Problems
geometry_keywords = ["triangle", "circle", "area", "volume", "angle", "segment"]
geom_problems = [
p for p in problems
if any(k in p["problem"].lower() for k in geometry_keywords)
]
print(f"Found {len(geom_problems)} geometry problems.")
Generating Novel Solutions with CreativeMath
The repository includes a generation pipeline for creating novel solutions to these mathematical problems. The core function generate_novel_solution in src/generation.py wraps LLM calls defined in src/models/api_models.py.
from src.generation import generate_novel_solution
# Choose a problem (e.g., the first AMC_8 entry)
sample = problems[0]
model_name = "gpt-4o" # any model listed in config.json
novel = generate_novel_solution(sample, model_name)
print("Original problem:")
print(sample["problem"])
print("\nNovel solution (model output):")
print(novel)
Generated outputs are stored in output/generation/ for subsequent evaluation using src/evaluation.py. The supported models are configured in src/config.json, which includes placeholders for API keys and model identifiers for OpenAI, Anthropic, and other providers.
Summary
- The CreativeMath dataset contains multiple-choice mathematical problems sourced from standardized competitions including AMC 8 and AMC 10.
- Problems span five major categories: Arithmetic & Algebra, Geometry, Number Theory & Combinatorics, Probability & Statistics, and Logic & Word Problems.
- Each problem includes metadata for competition source, difficulty level (1.0–3.0), and human-written solutions, stored in the structured
data/subset.jsonfile. - The repository provides a complete pipeline for novel solution generation via
src/generation.pyand evaluation viasrc/evaluation.py.
Frequently Asked Questions
What competitions are represented in the CreativeMath dataset?
The dataset primarily sources problems from AMC 8 and AMC 10 competitions, including specific variants like AMC 10 A and AMC 10 B. These competitions are visible in the competition field of each record in data/subset.json.
How is the difficulty of mathematical problems rated in CreativeMath?
Each problem carries a difficulty score ranging from 1.0 to 3.0, where level 1 represents middle-school difficulty, level 2 indicates high-school competition standards, and level 3 corresponds to Olympiad-level challenges. These ratings are stored in the difficulty field of the JSON schema.
What file contains the actual mathematical problems in the repository?
The core dataset resides in data/subset.json, which contains a JSON list of problem objects. Each object includes the problem statement (with LaTeX formatting), competition metadata, difficulty rating, and human-written solutions.
Can I use the CreativeMath dataset to test my own language models?
Yes, the repository is designed for benchmarking large language model creativity on mathematical reasoning. You can load problems from data/subset.json and use the provided src/generation.py module to generate novel solutions, or implement your own inference pipeline using the standardized problem format.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →