How Yappuccino Calculates Relevance Scores for Search Results
Yappuccino calculates relevance scores using a weighted point system where title matches contribute 3 points and content matches contribute 1 point, implemented via Django's Case and When database annotations in the PostListView class.
When implementing full-text search in Django applications, calculating meaningful relevance scores requires careful weighting of different field matches. The Yappuccino blog platform implements a custom search relevance algorithm that prioritizes title matches over content matches using conditional database aggregations in blog/views.py.
How the Search Functionality Calculates Relevance Scoring
The entire relevance calculation logic resides in blog/views.py within the PostListView.get_queryset() method (approximately lines 166-190). This implementation follows a three-stage pipeline: filtering candidate posts, annotating weighted scores, and ordering by relevance.
Filtering Posts by Search Term
First, the view narrows the queryset to posts containing the search string in high-priority fields. The method uses Django's Q objects to perform case-insensitive searches across the title, content, and author username fields, ensuring no duplicate entries appear in results.
queryset = queryset.filter(
Q(title__icontains=search) |
Q(content__icontains=search) |
Q(author__username__icontains=search)
).distinct()
While the filter includes author usernames, note that only title and content matches contribute to the numerical relevance score.
Weighted Scoring Using Case/When Annotations
The core relevance calculation occurs through Django's Count, Case, and When constructs. The system assigns 3 points for each occurrence in the title and 1 point for each occurrence in the content, creating a weighted aggregate that reflects search importance.
queryset = queryset.annotate(
relevance_score=Count(
Case(
When(title__icontains=search, then=3), # title matches are worth 3 points
When(content__icontains=search, then=1), # content matches are worth 1 point
default=0,
output_field=IntegerField()
)
)
)
This approach generates a temporary column yielding weighted values per row, which Count aggregates into a single integer representing the post's total relevance value.
Sorting Results by Relevance Score
After annotation, the view maps the sort query parameter to database fields. When users request sort=relevance, the system orders results by the computed relevance_score, falling back to date_posted if no search term exists.
sort_mapping = {
'date': 'date_posted',
'popularity': 'popularity_score',
'comments': 'comment_count',
'votes': 'vote_score',
'relevance': 'relevance_score' if search else 'date_posted'
}
The final order_by() call respects the chosen direction (asc or desc), placing the highest-scored posts first when descending order is selected.
Practical Implementation Examples
HTTP Request Format
To search for posts containing "postgres" and sort by highest relevance:
GET /?search=postgres&sort=relevance&order=desc HTTP/1.1
Host: yappuccino.example.com
Generating Search URLs Programmatically
Use Django's URL resolver to construct search links with relevance sorting:
from django.urls import reverse
search_url = (
reverse('post-list')
+ f"?search={keyword}&sort=relevance&order=desc"
)
Reusable Search Function
Extract the scoring logic for use in custom views or management commands:
from django.db.models import Count, Case, When, IntegerField, Q
def annotated_search(queryset, term):
return (
queryset.filter(
Q(title__icontains=term) |
Q(content__icontains=term) |
Q(author__username__icontains=term)
)
.annotate(
relevance_score=Count(
Case(
When(title__icontains=term, then=3),
When(content__icontains=term, then=1),
default=0,
output_field=IntegerField(),
)
)
)
.order_by("-relevance_score")
)
Summary
- The
PostListView.get_queryset()method inblog/views.pyimplements a three-stage relevance calculation pipeline. - Title matches receive a weight of 3 points, while content matches receive 1 point, using Django's
CaseandWhenannotations. - Author username matches are included in filtering but do not contribute to the numerical score.
- Results sort by the annotated
relevance_scorefield whensort=relevanceis specified in the query parameters. - The implementation translates directly to efficient SQL
CASEstatements withCOUNTaggregations.
Frequently Asked Questions
How does Yappuccino handle author username matches in relevance scoring?
While the search filter includes author__username__icontains to broaden result discovery, username matches do not contribute points to the relevance_score calculation. Only title and content fields generate weighted points (3 and 1 respectively), meaning posts matching solely by author appear in results but receive a base score of zero unless they also match title or content criteria.
Can the relevance point weights be customized?
Yes. Developers can modify the then=3 and then=1 values in the Case construct within blog/views.py to adjust the relative importance of title versus content matches. For instance, changing title matches to 5 points and content to 2 points would further prioritize posts with keywords in headlines, though this requires editing the source code directly.
Does this relevance scoring approach impact database performance?
The implementation uses standard SQL CASE and COUNT operations supported by all major databases. Performance depends on whether the title, content, and author_id columns are properly indexed. While icontains queries utilize LIKE operations that may table-scan on large datasets, the scoring calculation itself adds minimal overhead beyond the initial filtering.
What happens when users request relevance sorting without providing a search term?
According to the sort_mapping dictionary in blog/views.py, the view automatically falls back to date_posted ordering when sort=relevance is requested but no search parameter exists. This prevents attempting to sort by a non-existent relevance_score annotation and ensures users always receive a meaningful, chronologically ordered list.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →