How to Efficiently Load and Process the Exercises Dataset with Pandas
You can efficiently load and process the exercises dataset with Pandas by ingesting the flat JSON array from data/exercises.json into a DataFrame and applying vectorized filtering, grouping, and multilingual text extraction.
The hasaneyldrm/exercises-dataset repository hosts 1,324 fitness records in a single JSON file, making it straightforward to efficiently load and process the exercises dataset with Pandas. Because the data is a flat array at data/exercises.json, you can ingest it directly without complex ETL pipelines and immediately begin exploratory analysis. This guide covers the exact file structure, loading methods, and processing techniques grounded in the repository source.
Dataset Structure and Key Files
Primary Data Source and Schema
At the core of the repository, data/exercises.json contains the full dataset as a JSON array of 1,324 exercise objects. Each record includes fields such as id, name, category, equipment, multilingual instructions, and media assets including 180 × 180 px thumbnails and GIF animations referenced by image and gif_url. The companion file data/exercises.schema.json provides a JSON Schema (Draft 2020-12) definition that validates the shape of each record before loading.
Supporting Files
The repository includes several lightweight components that demonstrate language-agnostic consumption:
index.html— A client-side browser explorer that reads the same JSON without a server.setup.html— A developer guide that shows how to import the data into SQL, generate API snippets, and scaffold backends.README.md— Central documentation with usage examples and dataset statistics.
Efficiently Load the JSON Data into Pandas
Method 1: Using Python's json Module with pd.DataFrame
For explicit control over encoding and validation, load the raw JSON with Python's standard library and pass the resulting list into pd.DataFrame. As implemented in hasaneyldrm/exercises-dataset, this approach ensures you can inspect records before tabular conversion.
import json
import pandas as pd
# Load raw JSON
with open("data/exercises.json", "r", encoding="utf-8") as f:
exercises = json.load(f)
# Convert to DataFrame
df = pd.DataFrame(exercises)
print(f"Total exercises loaded: {len(df)}")
Method 2: Direct Load with pd.read_json
If you prefer a concise one-liner, pandas can ingest the flat array directly from disk. This is the fastest way to efficiently load and process the exercises dataset with Pandas.
df = pd.read_json("data/exercises.json")
Process and Analyze Exercise Records with Pandas
Aggregating by Category and Equipment
Once loaded, apply vectorized grouping and counting to understand distributions. The repository reports 1,324 total exercises, with categories such as 325 body-weight exercises. Use value_counts() to replicate these statistics in the DataFrame.
# Top 10 categories by number of exercises
print(df["category"].value_counts().head(10))
Filtering by Equipment and Body Part
Combine boolean masks to isolate specific training segments. For example, extract all barbell exercises targeting the upper legs by filtering on the equipment and category columns.
# All barbell exercises that target the upper legs
barbell_quads = df[(df["equipment"] == "barbell") & (df["category"] == "upper legs")]
print(barbell_quads[["name", "target", "equipment"]])
Working with Multilingual Instructions
Each exercise record stores step-by-step instructions in multiple languages inside a nested dictionary. Access English, Spanish, or other locales directly via the instructions column.
# Show English and Spanish instructions for the first exercise
ex = df.iloc[0]
print("English:", ex["instructions"]["en"])
print("Spanish:", ex["instructions"]["es"])
Downstream Analytics and Integration
With the data in a clean tabular format, pandas opens the door to statistical summaries, bar plots of exercise counts per body part, and machine-learning pipelines that require structured inputs. Because data/exercises.json is a flat array and data/exercises.schema.json enforces record consistency, you can trust the schema before feeding the DataFrame into larger fitness-tech stacks.
Summary
- The
hasaneyldrm/exercises-datasetrepository stores 1,324 exercises indata/exercises.jsonas a flat JSON array with a defined schema indata/exercises.schema.json. - You can efficiently load and process the exercises dataset with Pandas using either
pd.DataFrame(json.load(...))for explicit control orpd.read_json("data/exercises.json")for a single-line solution. - Once in a DataFrame, filter by
equipmentandcategory, group withvalue_counts(), and read multilingual instructions from the nestedinstructionscolumn.
Frequently Asked Questions
What fields are included in each exercise record?
Each object in data/exercises.json contains id, name, category, equipment, multilingual instructions, and media paths such as image and gif_url. The data/exercises.schema.json file formally defines these fields using JSON Schema Draft 2020-12.
Can I load the dataset without using Pandas?
Yes. The repository demonstrates language-agnostic usage: index.html consumes the JSON directly in the browser, and the standard json.load() function can process the file in pure Python without installing pandas.
How do I validate the dataset before loading it?
Validate records against data/exercises.schema.json using any JSON Schema validator. This schema enforces type constraints and required fields, enabling automated quality checks before you convert the array into a DataFrame.
How large is the dataset and what are its main categories?
The dataset contains 1,324 exercises. The README.md notes that 325 are body-weight exercises, and you can reproduce exact category distributions by calling df["category"].value_counts() after loading the data into pandas.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →