# How to Add Blog or Portfolio Enrichment to Hiring-Agent Beyond GitHub Profiles

> Learn how to enrich candidate profiles in Hiring-Agent beyond GitHub. Add blog and portfolio data by extending Pydantic models, creating a WebEnricher, and modifying score logic.

- Repository: [HackerRank/hiring-agent](https://github.com/interviewstreet/hiring-agent)
- Tags: how-to-guide
- Published: 2026-07-05

---

**Extend the Pydantic models in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py), implement a `WebEnricher` class in a new [`web_enricher.py`](https://github.com/interviewstreet/hiring-agent/blob/main/web_enricher.py) module, and modify [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) to fetch and summarize personal blogs or portfolio websites alongside existing GitHub data.**

The `interviewstreet/hiring-agent` repository currently enriches candidate evaluations by pulling GitHub profiles and repository data via [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py). To capture additional technical signals from personal blogs, Medium articles, or portfolio websites, you can extend the architecture with a modular web enrichment pipeline. This approach keeps the existing workflow intact while adding support for arbitrary web content analysis.

## Understanding the Current Enrichment Architecture

The hiring agent uses a pipeline pattern where [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) orchestrates data extraction. Currently, it processes PDF résumés through [`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py), extracts structured data using Jinja templates in `prompts/templates/`, and enriches results via [`github.py`](https://github.com/interviewstreet/hiring-agent/blob/main/github.py). To add blog or portfolio enrichment, you will mirror the GitHub enrichment pattern: create data models, implement a fetcher/parser class, and wire it into the evaluation flow.

## Step 1: Extend the Data Model in [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)

First, define Pydantic schema objects to represent blog content. Add these classes to **[`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)** after the existing `GitHubProfile` model:

```python

# models.py

from datetime import datetime
from typing import Optional, List
from pydantic import BaseModel

class BlogPost(BaseModel):
    title: str
    url: str
    published: Optional[datetime] = None
    summary: Optional[str] = None

class BlogProfile(BaseModel):
    """Information extracted from a personal blog or portfolio."""
    url: str
    author: Optional[str] = None
    description: Optional[str] = None
    recent_posts: List[BlogPost] = []

```

Update the `Basics` model to capture website URLs from résumés:

```python

# models.py – extend the existing Basics class

class Basics(BaseModel):
    name: str
    email: Optional[str] = None
    phone: Optional[str] = None
    website: Optional[str] = None  # New field for blog/portfolio URL

    summary: Optional[str] = None

```

These models become part of the top-level `EvaluationData` container, allowing downstream code to safely ignore them when absent.

## Step 2: Create the WebEnricher Module

Create a new file **[`web_enricher.py`](https://github.com/interviewstreet/hiring-agent/blob/main/web_enricher.py)** that handles HTTP fetching, HTML parsing, and LLM-powered summarization. This module reuses the existing LLM provider abstraction from **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)**:

```python

# web_enricher.py

import httpx
from bs4 import BeautifulSoup
from typing import List
from .models import BlogProfile, BlogPost
from .llm_utils import LLMProvider, get_llm_response

class WebEnricher:
    def __init__(self, llm: LLMProvider):
        self.llm = llm
        self.client = httpx.AsyncClient(timeout=10)

    async def fetch_html(self, url: str) -> str:
        resp = await self.client.get(url)
        resp.raise_for_status()
        return resp.text

    def extract_metadata(self, html: str) -> dict:
        soup = BeautifulSoup(html, "html.parser")
        title = soup.title.string if soup.title else ""
        description = (
            soup.find("meta", attrs={"name": "description"}) or
            soup.find("meta", attrs={"property": "og:description"})
        )
        return {
            "title": title.strip(),
            "description": (description["content"] if description else "").strip(),
        }

    async def summarize_posts(self, urls: List[str]) -> List[BlogPost]:
        posts = []
        for u in urls[:5]:  # Limit to 5 recent items

            html = await self.fetch_html(u)
            meta = self.extract_metadata(html)
            summary = await get_llm_response(
                self.llm,
                f"Summarize the article at {u} in 2-3 sentences."
            )
            posts.append(
                BlogPost(
                    title=meta["title"],
                    url=u,
                    summary=summary,
                )
            )
        return posts

    async def enrich(self, blog_url: str) -> BlogProfile:
        html = await self.fetch_html(blog_url)
        meta = self.extract_metadata(html)
        
        # Find recent post URLs using common path patterns

        soup = BeautifulSoup(html, "html.parser")
        post_links = [
            a["href"] for a in soup.find_all("a", href=True)
            if "/post" in a["href"] or "/blog" in a["href"]
        ][:5]
        
        recent_posts = await self.summarize_posts(post_links)
        
        return BlogProfile(
            url=blog_url,
            author=meta["title"],  # Often contains author name

            description=meta["description"],
            recent_posts=recent_posts,
        )

```

**Why this pattern works:**

- **Async HTTP**: Uses `httpx.AsyncClient` for non-blocking requests
- **Metadata extraction**: Leverages BeautifulSoup to parse OpenGraph and standard meta tags
- **LLM abstraction**: Reuses the existing `get_llm_response` utility from **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)**, ensuring consistent token handling and provider-agnostic prompts

## Step 3: Detect Blog URLs in Résumé Parsing

Extend the extraction template to capture website fields. Modify **`prompts/templates/basics.jinja`** to include a `website` field:

```jinja

# prompts/templates/basics.jinja

{
  "name": "{{ name }}",
  "email": "{{ email }}",
  "phone": "{{ phone }}",
  "website": "{{ website }}",
  "summary": "{{ summary }}"
}

```

When `PDFHandler` in **[`pdf.py`](https://github.com/interviewstreet/hiring-agent/blob/main/pdf.py)** processes the résumé, the LLM will populate the `website` field if present in the candidate's contact information or header section.

## Step 4: Wire the Enrichment into the Pipeline

Modify **[`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py)** to invoke `WebEnricher` after the existing GitHub enrichment. Add the blog enrichment logic to the orchestration flow:

```python

# score.py

from web_enricher import WebEnricher
import asyncio

async def enrich_blog(resume: JSONResume, llm):
    """Enrich resume with blog/portfolio data if website URL is present."""
    if resume.basics.website and "github.com" not in resume.basics.website:
        enricher = WebEnricher(llm)
        blog_profile = await enricher.enrich(resume.basics.website)
        # Add to evaluation data (ensure EvaluationData model accepts this field)

        resume.evaluation_data.blog = blog_profile
    return resume

async def main(pdf_path: str):
    # Existing PDF parsing flow

    resume = parse_pdf_resume(pdf_path)
    llm = init_llm()  # From existing initialization

    
    # Existing GitHub enrichment

    resume = await enrich_github(resume, llm)
    
    # New blog enrichment

    resume = await enrich_blog(resume, llm)
    
    # Continue with evaluation

    evaluation = evaluator.evaluate(resume)
    return evaluation

```

**Key implementation details:**

- **Conditional execution**: Only runs when `website` exists and is not a GitHub URL
- **Backward compatibility**: The `blog` field is optional; existing tests without website data continue to function
- **Unified LLM interface**: All LLM calls flow through the existing provider abstraction

## Step 5: Update Evaluation Logic (Optional)

To incorporate blog signals into scoring, extend **[`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)** to read the new `blog` field:

```python

# evaluator.py

def calculate_bonus_points(resume):
    bonus = 0
    evidence = []
    
    if resume.evaluation_data.blog:
        # Heuristic: bonus point per substantial recent post

        post_bonus = sum(
            1 for p in resume.evaluation_data.blog.recent_posts 
            if len(p.summary.split()) > 30
        )
        bonus += post_bonus
        evidence.append(f"Blog contributed {post_bonus} bonus point(s) from recent posts.")
    
    return bonus, evidence

```

Because the fields are optional, the logic automatically degrades to original behavior when no blog data is present.

## Practical Usage Example

When a candidate includes their portfolio in their résumé:

```markdown

# Jane Smith

jane@example.com | https://janedev.io

## Summary

Full-stack engineer with 5 years experience...

```

The system automatically:

1. Extracts `https://janedev.io` from the `website` field
2. Fetches the HTML using `WebEnricher.fetch_html`
3. Extracts metadata and recent post URLs via BeautifulSoup
4. Generates summaries using the configured LLM provider
5. Stores structured data in `BlogProfile` and `BlogPost` models

Debug the enrichment manually with:

```python
import asyncio
from llm_utils import get_provider
from web_enricher import WebEnricher

async def debug_enrichment():
    llm = get_provider()  # Uses LLM_PROVIDER env var

    enricher = WebEnricher(llm)
    profile = await enricher.enrich("https://janedev.io")
    print(profile.json(indent=2))

asyncio.run(debug_enrichment())

```

## Summary

Adding blog or portfolio enrichment to Hiring-Agent requires four architectural changes:

- **Extend [`models.py`](https://github.com/interviewstreet/hiring-agent/blob/main/models.py)** with `BlogPost` and `BlogProfile` Pydantic models to type the new data
- **Create [`web_enricher.py`](https://github.com/interviewstreet/hiring-agent/blob/main/web_enricher.py)** to handle async HTTP fetching, HTML parsing, and LLM summarization
- **Update `prompts/templates/basics.jinja`** and [`score.py`](https://github.com/interviewstreet/hiring-agent/blob/main/score.py) to detect and process website URLs from résumés
- **Modify [`evaluator.py`](https://github.com/interviewstreet/hiring-agent/blob/main/evaluator.py)** (optionally) to weight blog content in final scoring

This implementation respects the repository's modular design, maintains backward compatibility, and leverages existing LLM abstractions to minimize configuration overhead.

## Frequently Asked Questions

### What Python dependencies are required for blog enrichment?

The solution requires `httpx` for async HTTP requests and `beautifulsoup4` for HTML parsing. These integrate with the existing LLM provider infrastructure in **[`llm_utils.py`](https://github.com/interviewstreet/hiring-agent/blob/main/llm_utils.py)**. No additional API keys are needed beyond those already configured for the LLM back-end.

### How does the system handle websites that block web scraping?

The `WebEnricher` class uses `httpx.AsyncClient` with a 10-second timeout and standard headers. If a website returns a non-200 status code or blocks the request, `fetch_html` raises an exception that should be caught in `enrich_blog` to skip enrichment gracefully. For production deployments, consider adding retry logic or rotating user agents in [`web_enricher.py`](https://github.com/interviewstreet/hiring-agent/blob/main/web_enricher.py).

### Can I customize how many blog posts are analyzed?

Yes. In [`web_enricher.py`](https://github.com/interviewstreet/hiring-agent/blob/main/web_enricher.py), modify the slice `urls[:5]` in the `summarize_posts` method to adjust the number of posts processed. You can also refine the URL detection logic in the `enrich` method to target specific platforms like Medium (`medium.com`) or Dev.to by adjusting the href pattern matching.

### Does this work with platforms like Medium or Dev.to?

The current implementation detects posts via common URL patterns (`/post`, `/blog`), which covers most personal blogs and Medium publications. For platform-specific extraction (such as Dev.to's API or Medium's RSS feeds), extend `WebEnricher` with platform-specific parsers while keeping the `BlogProfile` output format consistent across all sources.