How Yappuccino Sanitizes Content with the Bleach Library for XSS Prevention

Yappuccino prevents cross-site scripting (XSS) attacks by using the Bleach library to parse and clean all user-generated HTML against strict whitelists before storing or displaying content.

The Yappuccino blog platform implements a defense-in-depth strategy for XSS prevention by centralizing content sanitization with the Bleach library. This Django-based application ensures that both blog posts and comments are rendered safely without executing malicious scripts, leveraging Bleach's HTML5 parser to rebuild the DOM from allowed elements only.

Centralized Whitelist Configuration in settings.py

All sanitization rules are defined centrally in blogpost/settings.py to ensure consistency across the application. The configuration specifies exactly which HTML tags, attributes, and CSS properties are considered safe.


# blogpost/settings.py

BLEACH_ALLOWED_TAGS = [
    'a', 'abbr', 'acronym', 'b', 'blockquote', 'code', 'em', 'i',
    'li', 'ol', 'p', 'strong', 'ul', 'h1', 'h2', 'h3', 'h4',
    'h5', 'h6', 'pre', 'span', 'img', 'table', 'tbody', 'tr',
    'td', 'th', 'thead', 'tfoot', 'hr', 'br', 'div', 'caption',
    'sub', 'sup', 'figure', 'figcaption',
]

BLEACH_ALLOWED_ATTRIBUTES = {
    'a': ['href', 'title', 'target', 'rel'],
    'abbr': ['title'],
    'acronym': ['title'],
    'img': ['src', 'alt', 'title', 'style', 'width', 'height', 'class'],
    '*': ['class', 'style'],
}

BLEACH_ALLOWED_STYLES = [
    'font-family', 'font-size', 'font-weight', 'font-style',
    'text-decoration', 'text-align', 'color', 'background-color',
    'margin', 'padding', 'width', 'height', 'border',
    'border-color', 'border-width', 'border-style', 'float',
]

css_sanitizer = CSSSanitizer(allowed_css_properties=BLEACH_ALLOWED_STYLES)
BLEACH_CSS_SANITIZER = css_sanitizer
BLEACH_STRIP_COMMENTS = True

This whitelist approach explicitly denies all tags not listed, including <script>, <iframe>, and event handlers like onclick, effectively neutralizing common XSS vectors.

Safe Content Rendering in Django Models

Yappuccino implements sanitization at the model layer through get_safe_content() methods, ensuring that every retrieval of user-generated content is automatically cleaned.

Post.get_safe_content() Implementation

In blog/models.py, the Post model uses bleach.clean() with the centralized settings and includes a fallback for older Bleach versions:


# blog/models.py – Post.get_safe_content()

def get_safe_content(self):
    try:
        return bleach.clean(
            self.content,
            tags=settings.BLEACH_ALLOWED_TAGS,
            attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
            css=settings.BLEACH_ALLOWED_STYLES,
            strip=True,
        )
    except TypeError:
        # Fallback for older Bleach versions that don’t accept `css`

        return bleach.clean(
            self.content,
            tags=settings.BLEACH_ALLOWED_TAGS,
            attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
            strip=True,
        )

Comment.get_safe_content() Implementation

The Comment model applies similar sanitization but ensures consistent paragraph wrapping for display:


# blog/models.py – Comment.get_safe_content()

def get_safe_content(self):
    if not self.content:
        return mark_safe("<p></p>")
    cleaned = bleach.clean(
        self.content,
        tags=settings.BLEACH_ALLOWED_TAGS,
        attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
        strip=True,
    )
    # Ensure a paragraph wrapper for consistent rendering

    if not cleaned.startswith('<p>'):
        cleaned = f'<p>{cleaned}</p>'
    return mark_safe(cleaned)

Both methods guarantee that any HTML stored in the database is rendered only with permitted elements, attributes, and CSS, removing malicious script tags before they reach the template.

Input Sanitization Before Database Storage

Yappuccino applies defense-in-depth by sanitizing content both on output (rendering) and on input (storage). When users edit comments, the view sanitizes the posted raw HTML before saving to the database.

The edit_comment View

In blog/views.py, the edit_comment function processes POST data through bleach.clean() before persistence:


# blog/views.py – edit_comment()

def edit_comment(request, pk):
    # ... authentication and retrieval logic ...

    raw_content = request.POST.get('content', '').strip()
    cleaned_content = bleach.clean(
        raw_content,
        tags=BLEACH_ALLOWED_TAGS,
        attributes=BLEACH_ALLOWED_ATTRIBUTES,
        strip=True,
    )
    if not cleaned_content or cleaned_content.strip() == '':
        return JsonResponse({'success': False, 'message': 'Comment cannot be empty after formatting'}, status=400)
    comment.content = cleaned_content
    comment.is_edited = True
    comment.save()
    # ... return success response ...

By sanitizing before persisting, Yappuccino guarantees that no unsafe markup ever reaches the database, creating a fail-safe barrier against stored XSS attacks.

Why Bleach? Architecture and Security Benefits

Bleach serves as a thin wrapper around the HTML5 parsing library html5lib. Unlike simple regex-based filters, Bleach parses the input HTML, rebuilds a safe Document Object Model (DOM) based strictly on the whitelist, and optionally sanitizes inline CSS via the CSSSanitizer.

This architectural approach prevents classic XSS payloads including:

  • <script> tags and JavaScript event handlers (onclick, onerror)
  • javascript: and data: URLs in anchor tags
  • Malformed HTML that might bypass string-based filters
  • CSS-based attacks via expression() or dangerous properties

The library's strict whitelist approach ensures that any element not explicitly permitted is either stripped or escaped, providing deterministic security regardless of input complexity.

Fail-Safe Defaults and Backward Compatibility

Yappuccino implements several defensive defaults to maximize security:

  • strip=True: Removes disallowed tags completely rather than escaping them, preventing accidental execution of malicious code
  • BLEACH_STRIP_COMMENTS = True: Discards HTML comments that could potentially hide malicious code or break parsers
  • TypeError fallback: The Post.get_safe_content() method catches TypeError to maintain compatibility with older Bleach releases that don't support the css argument, ensuring the application remains secure across dependency versions

Summary

Yappuccino's XSS defense relies on a layered architecture that centralizes content sanitization with the Bleach library:

  • Centralized configuration in blogpost/settings.py defines strict whitelists for HTML tags, attributes, and CSS properties
  • Model-layer sanitization via get_safe_content() methods ensures all rendered content is cleaned before reaching templates
  • Input validation in views like edit_comment sanitizes data before database persistence
  • Fail-safe defaults including strip=True and comment stripping prevent common bypass techniques

This approach ensures that user-generated content in blog posts and comments is displayed safely without executing malicious scripts.

Frequently Asked Questions

What is the Bleach library used for in Yappuccino?

The Bleach library is used as an HTML sanitization engine to prevent XSS attacks. It parses user-generated HTML using html5lib, rebuilds the DOM using only whitelisted tags and attributes defined in blogpost/settings.py, and strips or escapes any potentially malicious content such as <script> tags or event handlers.

How does Yappuccino prevent XSS attacks in comments?

Yappuccino prevents XSS in comments through a two-layer sanitization strategy. First, the edit_comment view in blog/views.py sanitizes incoming POST data using bleach.clean() before saving to the database. Second, the Comment.get_safe_content() method in blog/models.py sanitizes content again when rendering, ensuring defense in depth against stored XSS vulnerabilities.

What HTML tags are allowed in Yappuccino's Bleach configuration?

The whitelist in blogpost/settings.py allows semantic and formatting tags including: a, abbr, acronym, b, blockquote, code, em, i, li, ol, p, strong, ul, h1 through h6, pre, span, img, table and related elements, hr, br, div, caption, sub, sup, figure, and figcaption. Notably, <script> and <iframe> are excluded.

Does Yappuccino sanitize content on input or output?

Yappuccino sanitizes content both on input and output for maximum security. Input sanitization occurs in views like edit_comment before database persistence, ensuring malicious code never reaches storage. Output sanitization happens in model methods like get_safe_content() when rendering templates, providing a fail-safe layer in case any unsanitized data bypasses input validation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →