# Awesome-Python Metadata: A Complete Guide to the 71-Column Repository Schema

> Explore the awesome-python metadata with our guide to the 71-column github_data.json. Discover project stars, links, and similarity scores for thousands of Python projects.

- Repository: [Dylan Hogg/awesome-python](https://github.com/dylanhogg/awesome-python)
- Tags: deep-dive
- Published: 2026-03-01

---

**The awesome-python repository provides structured metadata across 71 columns in [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json), capturing everything from GitHub stars and bibliographic links to pre-computed similarity scores for every listed Python project.**

The awesome-python curated list by Dylan Hogg goes beyond a simple README by generating rich data files containing detailed metadata for hundreds of Python repositories. Understanding the structure and content of these awesome-python metadata files enables developers to build recommendation engines, conduct ecosystem research, and filter projects programmatically based on popularity metrics, academic citations, and technical attributes.

## Understanding the awesome-python Metadata Structure

The awesome-python metadata schema defines **71 distinct columns** that describe every project in the collection. This structured approach transforms the curated list into a queryable dataset suitable for data science workflows and programmatic discovery.

### Repository Identifiers and Classification

Core identification fields establish how each project is indexed and categorized:

- **index**: Internal row identifier for database operations
- **category**: Top-level classification (e.g., *ml-dl*, *web-frameworks*, *data-viz*)
- **githuburl**: Direct link to the source repository
- **customtopics**: Optional manual tags applied by curators
- **customabout**: Curated description narrative override

### Bibliographic and Academic Metadata

The dataset tracks academic and distribution associations through dedicated paper and package linkage fields:

- **customarxiv**, **_arxiv_links**, **_arxiv_count**: Associations with arXiv research papers
- **custompypi**, **_pypi_links**, **_pypi_count**: References to PyPI package distributions
- **_hf_links**, **_hf_count**: Connections to Hugging Face models and datasets

### GitHub Popularity and Activity Metrics

Quantitative engagement data enables filtering by project maturity and community adoption:

- **_stars**, **_forks**, **_watches**: Raw GitHub statistics as of last update
- **_stars_per_week**: Normalized growth metric accounting for repository age
- **_age_weeks**: Repository maturity calculation
- **_pop_contributor_count**, **_pop_commit_frequency**, **_pop_issue_count**: Development activity indicators
- **_pop_score**: Aggregated popularity metric used for internal ranking

### Repository Technical Details

Core GitHub metadata provides essential context about project implementation:

- **_repopath**, **_reponame**: Full repository path (`owner/repo`) and name
- **_language**: Primary programming language detected by GitHub
- **_homepage**: Associated project website or documentation URL
- **_github_description**: Repository tagline from GitHub API
- **_organization**: Owning organization or user account
- **_updated_at**, **_created_at**, **_last_commit_date**: Temporal tracking fields for freshness analysis

### Content Assets and Configuration

File system metadata captures the location of key repository assets:

- **_github_topics**, **_topics**: GitHub-provided tags and curated topic lists
- **_readme_filename**, **_readme_giturl**, **_readme_localurl**: README file locations across remote and cached sources
- **_requirements_filenames**, **_requirements_giturls**: Dependency specification file paths
- **_avatar_url**, **_avatar_filename**: Repository owner avatar images

### Similarity and Recommendation Data

- **sim**: Pre-computed similarity matrix containing related repositories, similarity scores, shared categories, and recommendation ranks

## Querying awesome-python Metadata with Python

The canonical metadata resides in [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json), where the full JSON schema is declared in lines 4–67 [https://github.com/dylanhogg/awesome-python/blob/main/github_data.json#L4-L67]. Actual repository records live under the `data` key, allowing direct import into pandas or other analytical tools.

### Loading the Full Dataset

```python
import pandas as pd
import json
from pathlib import Path

# Path to the generated JSON file

json_path = Path("github_data.json")

# Load the full JSON structure

with json_path.open() as f:
    payload = json.load(f)

# The actual rows are under the `data` key

df = pd.DataFrame(payload["data"])

# Quick look at essential columns

print(df[["index", "category", "githuburl", "_stars", "_language", "_homepage"]].head())

```

### Filtering by Popularity and Language

```python

# Projects with >10,000 stars written in Python

popular_python = df.query("_stars > 10000 and _language == 'Python'")
print(popular_python[["_reponame", "_stars", "_description"]])

```

### Analyzing Repository Similarities

```python
def top_similar(repo_index: int, n: int = 3):
    row = df.loc[df["index"] == repo_index].iloc[0]
    # `sim` is a list of lists: [repo, score, category, rank]

    similar = row["sim"][:n]
    for sim_repo, score, category, rank in similar:
        print(f"{sim_repo} – score {score:.2f} – cat {category}")

# Example: show similar repos for the 90th entry

top_similar(90)

```

### Exporting for Downstream Processing

```python

# CSV for spreadsheet compatibility

df.to_csv("awesome_python.csv", index=False)

# Parquet for high-performance analytics (Spark, Dask)

df.to_parquet("awesome_python.parquet")

```

## Available Data File Formats

The awesome-python metadata pipeline generates three distribution formats optimized for different use cases:

- **[`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json)**: Master source containing the embedded schema and complete repository records. Best for applications requiring schema validation or full metadata access.
- **`github_data.csv`**: Flattened export suitable for manual inspection in spreadsheet applications and simple data loading.
- **`github_data.parquet`**: Compressed, columnar storage format optimized for analytical query engines and large-scale data science workflows.

## Summary

- The awesome-python repository generates machine-readable metadata files containing **71 columns** of structured data for every listed project.
- **[`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json)** serves as the canonical source, with the complete schema definition located at lines 4–67 of the file.
- Metadata spans repository identifiers, bibliographic links (arXiv, PyPI, Hugging Face), GitHub statistics, content assets, and pre-calculated similarity scores.
- Three export formats (JSON, CSV, Parquet) support different analytical workflows from manual inspection to distributed computing.
- The `sim` column provides immediate access to repository recommendations without requiring additional machine learning infrastructure.

## Frequently Asked Questions

### What columns are available in the awesome-python metadata files?

The awesome-python metadata schema contains 71 columns covering repository identifiers (index, category, githuburl), GitHub statistics (_stars, _forks, _watches), bibliographic references (_arxiv_links, _pypi_links, _hf_links), content metadata (_readme_filename, _requirements_filenames), and pre-computed similarity scores (sim) for recommendation engines.

### How do I load awesome-python metadata into a pandas DataFrame?

Import [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json) using Python's `json` module to extract the `data` array, then pass it to `pd.DataFrame()`. The schema declaration occupies lines 4–67 of the file, while the actual repository records reside under the `data` key, enabling one-line conversion to a queryable DataFrame structure.

### What is the `sim` column in awesome-python metadata?

The `sim` column contains pre-computed similarity recommendations for each repository, stored as a nested list where each entry includes the similar repository name, similarity score (float), shared category (string), and rank (integer). This structure enables immediate implementation of "related projects" features without additional processing or API calls.

### Which awesome-python data file should I use for large-scale analytics?

Use `github_data.parquet` for large-scale analytics since it provides a compressed, columnar storage format optimized for query performance in analytics engines like Spark, Dask, or pandas. Use `github_data.csv` for manual spreadsheet inspection, and [`github_data.json`](https://github.com/dylanhogg/awesome-python/blob/main/github_data.json) when you need the authoritative schema definition or are working in JavaScript ecosystems.