How to Add Blog or Portfolio Enrichment to Hiring-Agent Beyond GitHub Profiles

Extend the Pydantic models in models.py, implement a WebEnricher class in a new web_enricher.py module, and modify score.py to fetch and summarize personal blogs or portfolio websites alongside existing GitHub data.

The interviewstreet/hiring-agent repository currently enriches candidate evaluations by pulling GitHub profiles and repository data via github.py. To capture additional technical signals from personal blogs, Medium articles, or portfolio websites, you can extend the architecture with a modular web enrichment pipeline. This approach keeps the existing workflow intact while adding support for arbitrary web content analysis.

Understanding the Current Enrichment Architecture

The hiring agent uses a pipeline pattern where score.py orchestrates data extraction. Currently, it processes PDF résumés through pdf.py, extracts structured data using Jinja templates in prompts/templates/, and enriches results via github.py. To add blog or portfolio enrichment, you will mirror the GitHub enrichment pattern: create data models, implement a fetcher/parser class, and wire it into the evaluation flow.

Step 1: Extend the Data Model in models.py

First, define Pydantic schema objects to represent blog content. Add these classes to models.py after the existing GitHubProfile model:


# models.py

from datetime import datetime
from typing import Optional, List
from pydantic import BaseModel

class BlogPost(BaseModel):
    title: str
    url: str
    published: Optional[datetime] = None
    summary: Optional[str] = None

class BlogProfile(BaseModel):
    """Information extracted from a personal blog or portfolio."""
    url: str
    author: Optional[str] = None
    description: Optional[str] = None
    recent_posts: List[BlogPost] = []

Update the Basics model to capture website URLs from résumés:


# models.py – extend the existing Basics class

class Basics(BaseModel):
    name: str
    email: Optional[str] = None
    phone: Optional[str] = None
    website: Optional[str] = None  # New field for blog/portfolio URL

    summary: Optional[str] = None

These models become part of the top-level EvaluationData container, allowing downstream code to safely ignore them when absent.

Step 2: Create the WebEnricher Module

Create a new file web_enricher.py that handles HTTP fetching, HTML parsing, and LLM-powered summarization. This module reuses the existing LLM provider abstraction from llm_utils.py:


# web_enricher.py

import httpx
from bs4 import BeautifulSoup
from typing import List
from .models import BlogProfile, BlogPost
from .llm_utils import LLMProvider, get_llm_response

class WebEnricher:
    def __init__(self, llm: LLMProvider):
        self.llm = llm
        self.client = httpx.AsyncClient(timeout=10)

    async def fetch_html(self, url: str) -> str:
        resp = await self.client.get(url)
        resp.raise_for_status()
        return resp.text

    def extract_metadata(self, html: str) -> dict:
        soup = BeautifulSoup(html, "html.parser")
        title = soup.title.string if soup.title else ""
        description = (
            soup.find("meta", attrs={"name": "description"}) or
            soup.find("meta", attrs={"property": "og:description"})
        )
        return {
            "title": title.strip(),
            "description": (description["content"] if description else "").strip(),
        }

    async def summarize_posts(self, urls: List[str]) -> List[BlogPost]:
        posts = []
        for u in urls[:5]:  # Limit to 5 recent items

            html = await self.fetch_html(u)
            meta = self.extract_metadata(html)
            summary = await get_llm_response(
                self.llm,
                f"Summarize the article at {u} in 2-3 sentences."
            )
            posts.append(
                BlogPost(
                    title=meta["title"],
                    url=u,
                    summary=summary,
                )
            )
        return posts

    async def enrich(self, blog_url: str) -> BlogProfile:
        html = await self.fetch_html(blog_url)
        meta = self.extract_metadata(html)
        
        # Find recent post URLs using common path patterns

        soup = BeautifulSoup(html, "html.parser")
        post_links = [
            a["href"] for a in soup.find_all("a", href=True)
            if "/post" in a["href"] or "/blog" in a["href"]
        ][:5]
        
        recent_posts = await self.summarize_posts(post_links)
        
        return BlogProfile(
            url=blog_url,
            author=meta["title"],  # Often contains author name

            description=meta["description"],
            recent_posts=recent_posts,
        )

Why this pattern works:

  • Async HTTP: Uses httpx.AsyncClient for non-blocking requests
  • Metadata extraction: Leverages BeautifulSoup to parse OpenGraph and standard meta tags
  • LLM abstraction: Reuses the existing get_llm_response utility from llm_utils.py, ensuring consistent token handling and provider-agnostic prompts

Step 3: Detect Blog URLs in Résumé Parsing

Extend the extraction template to capture website fields. Modify prompts/templates/basics.jinja to include a website field:


# prompts/templates/basics.jinja

{
  "name": "{{ name }}",
  "email": "{{ email }}",
  "phone": "{{ phone }}",
  "website": "{{ website }}",
  "summary": "{{ summary }}"
}

When PDFHandler in pdf.py processes the résumé, the LLM will populate the website field if present in the candidate's contact information or header section.

Step 4: Wire the Enrichment into the Pipeline

Modify score.py to invoke WebEnricher after the existing GitHub enrichment. Add the blog enrichment logic to the orchestration flow:


# score.py

from web_enricher import WebEnricher
import asyncio

async def enrich_blog(resume: JSONResume, llm):
    """Enrich resume with blog/portfolio data if website URL is present."""
    if resume.basics.website and "github.com" not in resume.basics.website:
        enricher = WebEnricher(llm)
        blog_profile = await enricher.enrich(resume.basics.website)
        # Add to evaluation data (ensure EvaluationData model accepts this field)

        resume.evaluation_data.blog = blog_profile
    return resume

async def main(pdf_path: str):
    # Existing PDF parsing flow

    resume = parse_pdf_resume(pdf_path)
    llm = init_llm()  # From existing initialization

    
    # Existing GitHub enrichment

    resume = await enrich_github(resume, llm)
    
    # New blog enrichment

    resume = await enrich_blog(resume, llm)
    
    # Continue with evaluation

    evaluation = evaluator.evaluate(resume)
    return evaluation

Key implementation details:

  • Conditional execution: Only runs when website exists and is not a GitHub URL
  • Backward compatibility: The blog field is optional; existing tests without website data continue to function
  • Unified LLM interface: All LLM calls flow through the existing provider abstraction

Step 5: Update Evaluation Logic (Optional)

To incorporate blog signals into scoring, extend evaluator.py to read the new blog field:


# evaluator.py

def calculate_bonus_points(resume):
    bonus = 0
    evidence = []
    
    if resume.evaluation_data.blog:
        # Heuristic: bonus point per substantial recent post

        post_bonus = sum(
            1 for p in resume.evaluation_data.blog.recent_posts 
            if len(p.summary.split()) > 30
        )
        bonus += post_bonus
        evidence.append(f"Blog contributed {post_bonus} bonus point(s) from recent posts.")
    
    return bonus, evidence

Because the fields are optional, the logic automatically degrades to original behavior when no blog data is present.

Practical Usage Example

When a candidate includes their portfolio in their résumé:


# Jane Smith

jane@example.com | https://janedev.io

## Summary

Full-stack engineer with 5 years experience...

The system automatically:

  1. Extracts https://janedev.io from the website field
  2. Fetches the HTML using WebEnricher.fetch_html
  3. Extracts metadata and recent post URLs via BeautifulSoup
  4. Generates summaries using the configured LLM provider
  5. Stores structured data in BlogProfile and BlogPost models

Debug the enrichment manually with:

import asyncio
from llm_utils import get_provider
from web_enricher import WebEnricher

async def debug_enrichment():
    llm = get_provider()  # Uses LLM_PROVIDER env var

    enricher = WebEnricher(llm)
    profile = await enricher.enrich("https://janedev.io")
    print(profile.json(indent=2))

asyncio.run(debug_enrichment())

Summary

Adding blog or portfolio enrichment to Hiring-Agent requires four architectural changes:

  • Extend models.py with BlogPost and BlogProfile Pydantic models to type the new data
  • Create web_enricher.py to handle async HTTP fetching, HTML parsing, and LLM summarization
  • Update prompts/templates/basics.jinja and score.py to detect and process website URLs from résumés
  • Modify evaluator.py (optionally) to weight blog content in final scoring

This implementation respects the repository's modular design, maintains backward compatibility, and leverages existing LLM abstractions to minimize configuration overhead.

Frequently Asked Questions

What Python dependencies are required for blog enrichment?

The solution requires httpx for async HTTP requests and beautifulsoup4 for HTML parsing. These integrate with the existing LLM provider infrastructure in llm_utils.py. No additional API keys are needed beyond those already configured for the LLM back-end.

How does the system handle websites that block web scraping?

The WebEnricher class uses httpx.AsyncClient with a 10-second timeout and standard headers. If a website returns a non-200 status code or blocks the request, fetch_html raises an exception that should be caught in enrich_blog to skip enrichment gracefully. For production deployments, consider adding retry logic or rotating user agents in web_enricher.py.

Can I customize how many blog posts are analyzed?

Yes. In web_enricher.py, modify the slice urls[:5] in the summarize_posts method to adjust the number of posts processed. You can also refine the URL detection logic in the enrich method to target specific platforms like Medium (medium.com) or Dev.to by adjusting the href pattern matching.

Does this work with platforms like Medium or Dev.to?

The current implementation detects posts via common URL patterns (/post, /blog), which covers most personal blogs and Medium publications. For platform-specific extraction (such as Dev.to's API or Medium's RSS feeds), extend WebEnricher with platform-specific parsers while keeping the BlogProfile output format consistent across all sources.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →