How Yappuccino Calculates Relevance Scores for Search Results

Yappuccino calculates relevance scores using a weighted point system where title matches contribute 3 points and content matches contribute 1 point, implemented via Django's Case and When database annotations in the PostListView class.

When implementing full-text search in Django applications, calculating meaningful relevance scores requires careful weighting of different field matches. The Yappuccino blog platform implements a custom search relevance algorithm that prioritizes title matches over content matches using conditional database aggregations in blog/views.py.

How the Search Functionality Calculates Relevance Scoring

The entire relevance calculation logic resides in blog/views.py within the PostListView.get_queryset() method (approximately lines 166-190). This implementation follows a three-stage pipeline: filtering candidate posts, annotating weighted scores, and ordering by relevance.

Filtering Posts by Search Term

First, the view narrows the queryset to posts containing the search string in high-priority fields. The method uses Django's Q objects to perform case-insensitive searches across the title, content, and author username fields, ensuring no duplicate entries appear in results.

queryset = queryset.filter(
    Q(title__icontains=search) |
    Q(content__icontains=search) |
    Q(author__username__icontains=search)
).distinct()

While the filter includes author usernames, note that only title and content matches contribute to the numerical relevance score.

Weighted Scoring Using Case/When Annotations

The core relevance calculation occurs through Django's Count, Case, and When constructs. The system assigns 3 points for each occurrence in the title and 1 point for each occurrence in the content, creating a weighted aggregate that reflects search importance.

queryset = queryset.annotate(
    relevance_score=Count(
        Case(
            When(title__icontains=search, then=3),   # title matches are worth 3 points

            When(content__icontains=search, then=1), # content matches are worth 1 point

            default=0,
            output_field=IntegerField()
        )
    )
)

This approach generates a temporary column yielding weighted values per row, which Count aggregates into a single integer representing the post's total relevance value.

Sorting Results by Relevance Score

After annotation, the view maps the sort query parameter to database fields. When users request sort=relevance, the system orders results by the computed relevance_score, falling back to date_posted if no search term exists.

sort_mapping = {
    'date': 'date_posted',
    'popularity': 'popularity_score',
    'comments': 'comment_count',
    'votes': 'vote_score',
    'relevance': 'relevance_score' if search else 'date_posted'
}

The final order_by() call respects the chosen direction (asc or desc), placing the highest-scored posts first when descending order is selected.

Practical Implementation Examples

HTTP Request Format

To search for posts containing "postgres" and sort by highest relevance:

GET /?search=postgres&sort=relevance&order=desc HTTP/1.1
Host: yappuccino.example.com

Generating Search URLs Programmatically

Use Django's URL resolver to construct search links with relevance sorting:

from django.urls import reverse

search_url = (
    reverse('post-list')
    + f"?search={keyword}&sort=relevance&order=desc"
)

Reusable Search Function

Extract the scoring logic for use in custom views or management commands:

from django.db.models import Count, Case, When, IntegerField, Q

def annotated_search(queryset, term):
    return (
        queryset.filter(
            Q(title__icontains=term) |
            Q(content__icontains=term) |
            Q(author__username__icontains=term)
        )
        .annotate(
            relevance_score=Count(
                Case(
                    When(title__icontains=term, then=3),
                    When(content__icontains=term, then=1),
                    default=0,
                    output_field=IntegerField(),
                )
            )
        )
        .order_by("-relevance_score")
    )

Summary

  • The PostListView.get_queryset() method in blog/views.py implements a three-stage relevance calculation pipeline.
  • Title matches receive a weight of 3 points, while content matches receive 1 point, using Django's Case and When annotations.
  • Author username matches are included in filtering but do not contribute to the numerical score.
  • Results sort by the annotated relevance_score field when sort=relevance is specified in the query parameters.
  • The implementation translates directly to efficient SQL CASE statements with COUNT aggregations.

Frequently Asked Questions

How does Yappuccino handle author username matches in relevance scoring?

While the search filter includes author__username__icontains to broaden result discovery, username matches do not contribute points to the relevance_score calculation. Only title and content fields generate weighted points (3 and 1 respectively), meaning posts matching solely by author appear in results but receive a base score of zero unless they also match title or content criteria.

Can the relevance point weights be customized?

Yes. Developers can modify the then=3 and then=1 values in the Case construct within blog/views.py to adjust the relative importance of title versus content matches. For instance, changing title matches to 5 points and content to 2 points would further prioritize posts with keywords in headlines, though this requires editing the source code directly.

Does this relevance scoring approach impact database performance?

The implementation uses standard SQL CASE and COUNT operations supported by all major databases. Performance depends on whether the title, content, and author_id columns are properly indexed. While icontains queries utilize LIKE operations that may table-scan on large datasets, the scoring calculation itself adds minimal overhead beyond the initial filtering.

What happens when users request relevance sorting without providing a search term?

According to the sort_mapping dictionary in blog/views.py, the view automatically falls back to date_posted ordering when sort=relevance is requested but no search parameter exists. This prevents attempting to sort by a non-existent relevance_score annotation and ensures users always receive a meaningful, chronologically ordered list.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →