How to Add Blog or Portfolio Enrichment to Hiring-Agent Beyond GitHub Profiles
Extend the Pydantic models in models.py, implement a WebEnricher class in a new web_enricher.py module, and modify score.py to fetch and summarize personal blogs or portfolio websites alongside existing GitHub data.
The interviewstreet/hiring-agent repository currently enriches candidate evaluations by pulling GitHub profiles and repository data via github.py. To capture additional technical signals from personal blogs, Medium articles, or portfolio websites, you can extend the architecture with a modular web enrichment pipeline. This approach keeps the existing workflow intact while adding support for arbitrary web content analysis.
Understanding the Current Enrichment Architecture
The hiring agent uses a pipeline pattern where score.py orchestrates data extraction. Currently, it processes PDF résumés through pdf.py, extracts structured data using Jinja templates in prompts/templates/, and enriches results via github.py. To add blog or portfolio enrichment, you will mirror the GitHub enrichment pattern: create data models, implement a fetcher/parser class, and wire it into the evaluation flow.
Step 1: Extend the Data Model in models.py
First, define Pydantic schema objects to represent blog content. Add these classes to models.py after the existing GitHubProfile model:
# models.py
from datetime import datetime
from typing import Optional, List
from pydantic import BaseModel
class BlogPost(BaseModel):
title: str
url: str
published: Optional[datetime] = None
summary: Optional[str] = None
class BlogProfile(BaseModel):
"""Information extracted from a personal blog or portfolio."""
url: str
author: Optional[str] = None
description: Optional[str] = None
recent_posts: List[BlogPost] = []
Update the Basics model to capture website URLs from résumés:
# models.py – extend the existing Basics class
class Basics(BaseModel):
name: str
email: Optional[str] = None
phone: Optional[str] = None
website: Optional[str] = None # New field for blog/portfolio URL
summary: Optional[str] = None
These models become part of the top-level EvaluationData container, allowing downstream code to safely ignore them when absent.
Step 2: Create the WebEnricher Module
Create a new file web_enricher.py that handles HTTP fetching, HTML parsing, and LLM-powered summarization. This module reuses the existing LLM provider abstraction from llm_utils.py:
# web_enricher.py
import httpx
from bs4 import BeautifulSoup
from typing import List
from .models import BlogProfile, BlogPost
from .llm_utils import LLMProvider, get_llm_response
class WebEnricher:
def __init__(self, llm: LLMProvider):
self.llm = llm
self.client = httpx.AsyncClient(timeout=10)
async def fetch_html(self, url: str) -> str:
resp = await self.client.get(url)
resp.raise_for_status()
return resp.text
def extract_metadata(self, html: str) -> dict:
soup = BeautifulSoup(html, "html.parser")
title = soup.title.string if soup.title else ""
description = (
soup.find("meta", attrs={"name": "description"}) or
soup.find("meta", attrs={"property": "og:description"})
)
return {
"title": title.strip(),
"description": (description["content"] if description else "").strip(),
}
async def summarize_posts(self, urls: List[str]) -> List[BlogPost]:
posts = []
for u in urls[:5]: # Limit to 5 recent items
html = await self.fetch_html(u)
meta = self.extract_metadata(html)
summary = await get_llm_response(
self.llm,
f"Summarize the article at {u} in 2-3 sentences."
)
posts.append(
BlogPost(
title=meta["title"],
url=u,
summary=summary,
)
)
return posts
async def enrich(self, blog_url: str) -> BlogProfile:
html = await self.fetch_html(blog_url)
meta = self.extract_metadata(html)
# Find recent post URLs using common path patterns
soup = BeautifulSoup(html, "html.parser")
post_links = [
a["href"] for a in soup.find_all("a", href=True)
if "/post" in a["href"] or "/blog" in a["href"]
][:5]
recent_posts = await self.summarize_posts(post_links)
return BlogProfile(
url=blog_url,
author=meta["title"], # Often contains author name
description=meta["description"],
recent_posts=recent_posts,
)
Why this pattern works:
- Async HTTP: Uses
httpx.AsyncClientfor non-blocking requests - Metadata extraction: Leverages BeautifulSoup to parse OpenGraph and standard meta tags
- LLM abstraction: Reuses the existing
get_llm_responseutility fromllm_utils.py, ensuring consistent token handling and provider-agnostic prompts
Step 3: Detect Blog URLs in Résumé Parsing
Extend the extraction template to capture website fields. Modify prompts/templates/basics.jinja to include a website field:
# prompts/templates/basics.jinja
{
"name": "{{ name }}",
"email": "{{ email }}",
"phone": "{{ phone }}",
"website": "{{ website }}",
"summary": "{{ summary }}"
}
When PDFHandler in pdf.py processes the résumé, the LLM will populate the website field if present in the candidate's contact information or header section.
Step 4: Wire the Enrichment into the Pipeline
Modify score.py to invoke WebEnricher after the existing GitHub enrichment. Add the blog enrichment logic to the orchestration flow:
# score.py
from web_enricher import WebEnricher
import asyncio
async def enrich_blog(resume: JSONResume, llm):
"""Enrich resume with blog/portfolio data if website URL is present."""
if resume.basics.website and "github.com" not in resume.basics.website:
enricher = WebEnricher(llm)
blog_profile = await enricher.enrich(resume.basics.website)
# Add to evaluation data (ensure EvaluationData model accepts this field)
resume.evaluation_data.blog = blog_profile
return resume
async def main(pdf_path: str):
# Existing PDF parsing flow
resume = parse_pdf_resume(pdf_path)
llm = init_llm() # From existing initialization
# Existing GitHub enrichment
resume = await enrich_github(resume, llm)
# New blog enrichment
resume = await enrich_blog(resume, llm)
# Continue with evaluation
evaluation = evaluator.evaluate(resume)
return evaluation
Key implementation details:
- Conditional execution: Only runs when
websiteexists and is not a GitHub URL - Backward compatibility: The
blogfield is optional; existing tests without website data continue to function - Unified LLM interface: All LLM calls flow through the existing provider abstraction
Step 5: Update Evaluation Logic (Optional)
To incorporate blog signals into scoring, extend evaluator.py to read the new blog field:
# evaluator.py
def calculate_bonus_points(resume):
bonus = 0
evidence = []
if resume.evaluation_data.blog:
# Heuristic: bonus point per substantial recent post
post_bonus = sum(
1 for p in resume.evaluation_data.blog.recent_posts
if len(p.summary.split()) > 30
)
bonus += post_bonus
evidence.append(f"Blog contributed {post_bonus} bonus point(s) from recent posts.")
return bonus, evidence
Because the fields are optional, the logic automatically degrades to original behavior when no blog data is present.
Practical Usage Example
When a candidate includes their portfolio in their résumé:
# Jane Smith
jane@example.com | https://janedev.io
## Summary
Full-stack engineer with 5 years experience...
The system automatically:
- Extracts
https://janedev.iofrom thewebsitefield - Fetches the HTML using
WebEnricher.fetch_html - Extracts metadata and recent post URLs via BeautifulSoup
- Generates summaries using the configured LLM provider
- Stores structured data in
BlogProfileandBlogPostmodels
Debug the enrichment manually with:
import asyncio
from llm_utils import get_provider
from web_enricher import WebEnricher
async def debug_enrichment():
llm = get_provider() # Uses LLM_PROVIDER env var
enricher = WebEnricher(llm)
profile = await enricher.enrich("https://janedev.io")
print(profile.json(indent=2))
asyncio.run(debug_enrichment())
Summary
Adding blog or portfolio enrichment to Hiring-Agent requires four architectural changes:
- Extend
models.pywithBlogPostandBlogProfilePydantic models to type the new data - Create
web_enricher.pyto handle async HTTP fetching, HTML parsing, and LLM summarization - Update
prompts/templates/basics.jinjaandscore.pyto detect and process website URLs from résumés - Modify
evaluator.py(optionally) to weight blog content in final scoring
This implementation respects the repository's modular design, maintains backward compatibility, and leverages existing LLM abstractions to minimize configuration overhead.
Frequently Asked Questions
What Python dependencies are required for blog enrichment?
The solution requires httpx for async HTTP requests and beautifulsoup4 for HTML parsing. These integrate with the existing LLM provider infrastructure in llm_utils.py. No additional API keys are needed beyond those already configured for the LLM back-end.
How does the system handle websites that block web scraping?
The WebEnricher class uses httpx.AsyncClient with a 10-second timeout and standard headers. If a website returns a non-200 status code or blocks the request, fetch_html raises an exception that should be caught in enrich_blog to skip enrichment gracefully. For production deployments, consider adding retry logic or rotating user agents in web_enricher.py.
Can I customize how many blog posts are analyzed?
Yes. In web_enricher.py, modify the slice urls[:5] in the summarize_posts method to adjust the number of posts processed. You can also refine the URL detection logic in the enrich method to target specific platforms like Medium (medium.com) or Dev.to by adjusting the href pattern matching.
Does this work with platforms like Medium or Dev.to?
The current implementation detects posts via common URL patterns (/post, /blog), which covers most personal blogs and Medium publications. For platform-specific extraction (such as Dev.to's API or Medium's RSS feeds), extend WebEnricher with platform-specific parsers while keeping the BlogProfile output format consistent across all sources.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →