# How to Use JSON Data from Awesome-Python Programmatically

> Learn to programmatically access awesome-python JSON data. Get package classifications, GitHub stats, and repo details directly without web scraping. Access structured metadata easily.

- Repository: [Dylan Hogg/awesome-python](https://github.com/dylanhogg/awesome-python)
- Tags: how-to-guide
- Published: 2026-03-01

---

**The awesome-python repository maintains a machine-readable [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json) file that contains structured metadata for every listed Python package, enabling programmatic access to category classifications, GitHub statistics, and repository details without scraping the README.**

The `dylanhogg/awesome-python` repository is a curated list of Python frameworks, libraries, and resources. While the README provides a human-friendly view, the project also distributes a machine-readable JSON dataset that allows developers to **use JSON data from awesome-python programmatically** for automation, analysis, and integration into their own tools.

## Understanding the Awesome-Python JSON Schema

The primary data file is [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json) located at the repository root. This file contains a flat array of objects, where each object represents a single Python package entry with over 20 metadata fields extracted from GitHub's API.

Key fields include:

- `category`: The section heading from the README (e.g., "Machine Learning - Deep Learning")
- `githuburl`: Direct link to the repository
- `_stars`, `_forks`, `_watches`: GitHub engagement metrics
- `_repopath` and `_reponame`: Repository identifiers
- `_language`: Primary programming language (typically "Python")
- `_updated_at` and `_created_at`: ISO-8601 timestamps
- `description`: Short summary from the README
- `customtopics`, `customabout`: Optional maintainer-supplied metadata
- `_stars_per_week`: Normalized popularity metric

## Loading and Parsing the JSON Data

You can retrieve the dataset directly from the raw GitHub URL without cloning the repository:

```text
https://raw.githubusercontent.com/dylanhogg/awesome-python/main/github_data.json

```

### Using Python's Standard Library

For lightweight applications, use the built-in `json` module:

```python
import json
from pathlib import Path

# Load from local file or downloaded raw content

json_path = Path("github_data.json")

with json_path.open(encoding="utf-8") as f:
    data = json.load(f)  # List of dictionaries

print(f"Total packages: {len(data)}")
print(f"First entry: {data[0]['_reponame']} in category {data[0]['category']}")

```

### Using Pandas for Data Analysis

For filtering and aggregation, load the JSON into a DataFrame:

```python
import pandas as pd

df = pd.read_json("github_data.json")

# Filter for Machine Learning projects with high star counts

ml_projects = (
    df.query("category == 'Machine Learning'")
    .query("_stars > 1000")
    .sort_values("_stars", ascending=False)
)

print(ml_projects[["_reponame", "_stars", "description"]].head())

```

## Practical Applications and Examples

### Filtering by Category and Popularity

You can build targeted queries using the `category` and `_stars` fields to find relevant tools:

```python
def find_top_packages(category: str, min_stars: int = 500, limit: int = 10):
    """Find the most starred packages in a specific category."""
    matches = [
        entry for entry in data 
        if entry["category"] == category and entry["_stars"] > min_stars
    ]
    matches.sort(key=lambda x: x["_stars"], reverse=True)
    return matches[:limit]

# Example: Find top Web frameworks

for repo in find_top_packages("Web Frameworks", min_stars=10000):
    print(f"{repo['_reponame']}: {repo['_stars']} stars")

```

### Building a Recommendation Engine

The `featured` score and `_stars_per_week` metrics enable popularity-based recommendations:

```python
def recommend_by_popularity(category: str, top_n: int = 5) -> list[dict]:
    """Return the most popular repositories based on stars per week."""
    category_matches = [e for e in data if e["category"] == category]
    # Sort by stars per week (normalized popularity metric)

    category_matches.sort(key=lambda x: x.get("_stars_per_week", 0), reverse=True)
    return category_matches[:top_n]

# Get trending Machine Learning libraries

trending = recommend_by_popularity("Machine Learning - Deep Learning", top_n=3)
for repo in trending:
    print(f"{repo['_reponame']}: {repo['_stars_per_week']:.1f} stars/week")

```

## Summary

- The [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json) file in `dylanhogg/awesome-python` provides structured, machine-readable data for every listed Python package.
- Each entry contains rich metadata including GitHub statistics (`_stars`, `_forks`), categorization (`category`), and temporal data (`_updated_at`, `_age_weeks`).
- You can load the JSON using Python's built-in `json` module for simple scripts or **pandas** for complex filtering and analysis.
- The dataset enables programmatic discovery of tools by category, popularity metrics, or custom metadata fields without parsing the README markdown.

## Frequently Asked Questions

### How often is the github_data.json file updated?

The JSON file is regenerated automatically as part of the repository's continuous integration pipeline. This ensures the GitHub statistics (`_stars`, `_forks`, `_updated_at`) and repository metadata remain synchronized with the current state of the README and the actual GitHub repositories.

### Can I use the awesome-python JSON data in commercial applications?

Yes, the data is provided as part of an open-source repository. While the specific license for the data compilation should be checked in the repository's LICENSE file, the JSON structure itself contains publicly available GitHub metadata. Always verify compliance with GitHub's Terms of Service when accessing repository data programmatically.

### What is the difference between the description and _github_description fields?

The `description` field contains the short summary written by the awesome-python maintainers in the README, while `_github_description` stores the original repository description retrieved from the GitHub API. These may differ when maintainers provide curated context in the README that differs from the repository's official GitHub description.

### How do I filter for packages in multiple categories?

Since the `category` field is a string containing the full section path (e.g., "Machine Learning - Deep Learning"), you can use substring matching or exact equality checks. For pandas queries, use `.str.contains()` for partial matches: `df[df['category'].str.contains('Machine Learning')]`. For standard Python, use list comprehensions with `if "Machine Learning" in entry["category"]`.