# How Yappuccino Sanitizes Content with the Bleach Library for XSS Prevention

> Learn how Yappuccino sanitizes content using the Bleach library to effectively prevent XSS attacks by cleaning user generated HTML against strict whitelists.

- Repository: [Ja'farbek Yusupov/yappuccino](https://github.com/jafarbekyusupov/yappuccino)
- Tags: best-practices
- Published: 2026-03-04

---

**Yappuccino prevents cross-site scripting (XSS) attacks by using the Bleach library to parse and clean all user-generated HTML against strict whitelists before storing or displaying content.**

The Yappuccino blog platform implements a defense-in-depth strategy for XSS prevention by centralizing content sanitization with the Bleach library. This Django-based application ensures that both blog posts and comments are rendered safely without executing malicious scripts, leveraging Bleach's HTML5 parser to rebuild the DOM from allowed elements only.

## Centralized Whitelist Configuration in settings.py

All sanitization rules are defined centrally in [`blogpost/settings.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blogpost/settings.py) to ensure consistency across the application. The configuration specifies exactly which HTML tags, attributes, and CSS properties are considered safe.

```python

# blogpost/settings.py

BLEACH_ALLOWED_TAGS = [
    'a', 'abbr', 'acronym', 'b', 'blockquote', 'code', 'em', 'i',
    'li', 'ol', 'p', 'strong', 'ul', 'h1', 'h2', 'h3', 'h4',
    'h5', 'h6', 'pre', 'span', 'img', 'table', 'tbody', 'tr',
    'td', 'th', 'thead', 'tfoot', 'hr', 'br', 'div', 'caption',
    'sub', 'sup', 'figure', 'figcaption',
]

BLEACH_ALLOWED_ATTRIBUTES = {
    'a': ['href', 'title', 'target', 'rel'],
    'abbr': ['title'],
    'acronym': ['title'],
    'img': ['src', 'alt', 'title', 'style', 'width', 'height', 'class'],
    '*': ['class', 'style'],
}

BLEACH_ALLOWED_STYLES = [
    'font-family', 'font-size', 'font-weight', 'font-style',
    'text-decoration', 'text-align', 'color', 'background-color',
    'margin', 'padding', 'width', 'height', 'border',
    'border-color', 'border-width', 'border-style', 'float',
]

css_sanitizer = CSSSanitizer(allowed_css_properties=BLEACH_ALLOWED_STYLES)
BLEACH_CSS_SANITIZER = css_sanitizer
BLEACH_STRIP_COMMENTS = True

```

This whitelist approach explicitly denies all tags not listed, including `<script>`, `<iframe>`, and event handlers like `onclick`, effectively neutralizing common XSS vectors.

## Safe Content Rendering in Django Models

Yappuccino implements sanitization at the model layer through `get_safe_content()` methods, ensuring that every retrieval of user-generated content is automatically cleaned.

### Post.get_safe_content() Implementation

In [`blog/models.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blog/models.py), the `Post` model uses `bleach.clean()` with the centralized settings and includes a fallback for older Bleach versions:

```python

# blog/models.py – Post.get_safe_content()

def get_safe_content(self):
    try:
        return bleach.clean(
            self.content,
            tags=settings.BLEACH_ALLOWED_TAGS,
            attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
            css=settings.BLEACH_ALLOWED_STYLES,
            strip=True,
        )
    except TypeError:
        # Fallback for older Bleach versions that don’t accept `css`

        return bleach.clean(
            self.content,
            tags=settings.BLEACH_ALLOWED_TAGS,
            attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
            strip=True,
        )

```

### Comment.get_safe_content() Implementation

The `Comment` model applies similar sanitization but ensures consistent paragraph wrapping for display:

```python

# blog/models.py – Comment.get_safe_content()

def get_safe_content(self):
    if not self.content:
        return mark_safe("<p></p>")
    cleaned = bleach.clean(
        self.content,
        tags=settings.BLEACH_ALLOWED_TAGS,
        attributes=settings.BLEACH_ALLOWED_ATTRIBUTES,
        strip=True,
    )
    # Ensure a paragraph wrapper for consistent rendering

    if not cleaned.startswith('<p>'):
        cleaned = f'<p>{cleaned}</p>'
    return mark_safe(cleaned)

```

Both methods guarantee that any HTML stored in the database is rendered only with permitted elements, attributes, and CSS, removing malicious script tags before they reach the template.

## Input Sanitization Before Database Storage

Yappuccino applies defense-in-depth by sanitizing content both on output (rendering) and on input (storage). When users edit comments, the view sanitizes the posted raw HTML before saving to the database.

### The edit_comment View

In [`blog/views.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blog/views.py), the `edit_comment` function processes POST data through `bleach.clean()` before persistence:

```python

# blog/views.py – edit_comment()

def edit_comment(request, pk):
    # ... authentication and retrieval logic ...

    raw_content = request.POST.get('content', '').strip()
    cleaned_content = bleach.clean(
        raw_content,
        tags=BLEACH_ALLOWED_TAGS,
        attributes=BLEACH_ALLOWED_ATTRIBUTES,
        strip=True,
    )
    if not cleaned_content or cleaned_content.strip() == '':
        return JsonResponse({'success': False, 'message': 'Comment cannot be empty after formatting'}, status=400)
    comment.content = cleaned_content
    comment.is_edited = True
    comment.save()
    # ... return success response ...

```

By sanitizing before persisting, Yappuccino guarantees that no unsafe markup ever reaches the database, creating a fail-safe barrier against stored XSS attacks.

## Why Bleach? Architecture and Security Benefits

Bleach serves as a thin wrapper around the HTML5 parsing library `html5lib`. Unlike simple regex-based filters, Bleach parses the input HTML, rebuilds a safe Document Object Model (DOM) based strictly on the whitelist, and optionally sanitizes inline CSS via the `CSSSanitizer`.

This architectural approach prevents classic XSS payloads including:
- `<script>` tags and JavaScript event handlers (`onclick`, `onerror`)
- `javascript:` and `data:` URLs in anchor tags
- Malformed HTML that might bypass string-based filters
- CSS-based attacks via `expression()` or dangerous properties

The library's strict whitelist approach ensures that any element not explicitly permitted is either stripped or escaped, providing deterministic security regardless of input complexity.

## Fail-Safe Defaults and Backward Compatibility

Yappuccino implements several defensive defaults to maximize security:

- **`strip=True`**: Removes disallowed tags completely rather than escaping them, preventing accidental execution of malicious code
- **`BLEACH_STRIP_COMMENTS = True`**: Discards HTML comments that could potentially hide malicious code or break parsers
- **TypeError fallback**: The `Post.get_safe_content()` method catches `TypeError` to maintain compatibility with older Bleach releases that don't support the `css` argument, ensuring the application remains secure across dependency versions

## Summary

Yappuccino's XSS defense relies on a layered architecture that centralizes content sanitization with the Bleach library:

- **Centralized configuration** in [`blogpost/settings.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blogpost/settings.py) defines strict whitelists for HTML tags, attributes, and CSS properties
- **Model-layer sanitization** via `get_safe_content()` methods ensures all rendered content is cleaned before reaching templates
- **Input validation** in views like `edit_comment` sanitizes data before database persistence
- **Fail-safe defaults** including `strip=True` and comment stripping prevent common bypass techniques

This approach ensures that user-generated content in blog posts and comments is displayed safely without executing malicious scripts.

## Frequently Asked Questions

### What is the Bleach library used for in Yappuccino?

The Bleach library is used as an HTML sanitization engine to prevent XSS attacks. It parses user-generated HTML using `html5lib`, rebuilds the DOM using only whitelisted tags and attributes defined in [`blogpost/settings.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blogpost/settings.py), and strips or escapes any potentially malicious content such as `<script>` tags or event handlers.

### How does Yappuccino prevent XSS attacks in comments?

Yappuccino prevents XSS in comments through a two-layer sanitization strategy. First, the `edit_comment` view in [`blog/views.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blog/views.py) sanitizes incoming POST data using `bleach.clean()` before saving to the database. Second, the `Comment.get_safe_content()` method in [`blog/models.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blog/models.py) sanitizes content again when rendering, ensuring defense in depth against stored XSS vulnerabilities.

### What HTML tags are allowed in Yappuccino's Bleach configuration?

The whitelist in [`blogpost/settings.py`](https://github.com/jafarbekyusupov/yappuccino/blob/main/blogpost/settings.py) allows semantic and formatting tags including: `a`, `abbr`, `acronym`, `b`, `blockquote`, `code`, `em`, `i`, `li`, `ol`, `p`, `strong`, `ul`, `h1` through `h6`, `pre`, `span`, `img`, `table` and related elements, `hr`, `br`, `div`, `caption`, `sub`, `sup`, `figure`, and `figcaption`. Notably, `<script>` and `<iframe>` are excluded.

### Does Yappuccino sanitize content on input or output?

Yappuccino sanitizes content both on input and output for maximum security. Input sanitization occurs in views like `edit_comment` before database persistence, ensuring malicious code never reaches storage. Output sanitization happens in model methods like `get_safe_content()` when rendering templates, providing a fail-safe layer in case any unsanitized data bypasses input validation.