How user-scanner Prevents CSV Injection Attacks: Formula Neutralization Explained
user-scanner sanitizes every CSV cell by prepending a single quote to any value that starts with formula-triggering characters, forcing spreadsheets to treat the content as plain text.
CSV injection (also known as formula injection) occurs when malicious data starting with =, +, -, or @ causes spreadsheet programs to execute arbitrary commands. The open-source tool user-scanner eliminates this attack vector through a dedicated sanitization routine applied to all exported results. This article examines the implementation in the kaifcodec/user-scanner repository, tracing how user-controlled data is neutralized before reaching the CSV output.
The CSV Injection Threat
Spreadsheet applications like Microsoft Excel and LibreOffice Calc automatically interpret cell values beginning with certain characters as formulas. An attacker who controls a username field could inject =HYPERLINK("http://malicious") or +cmd|' /C calc'!A0—executing commands or exfiltrating data when the CSV is opened.
According to the user-scanner source code, the defense centers on detecting these triggers and rendering them harmless.
Core Sanitization: _neutralize_csv_cell in result.py
The protection mechanism lives in user_scanner/core/result.py within the Result class. The static method _neutralize_csv_cell examines each value before CSV serialization:
# user_scanner/core/result.py
def _neutralize_csv_cell(value):
FORMULA_TRIGGER_CHARS = ("=", "+", "-", "@", "\t", "\r", "\n")
if value is None:
return value
s = str(value)
if s.lstrip().startswith(FORMULA_TRIGGER_CHARS):
return "'" + s # prepend a quote to neutralise
return value
How It Works
- Trigger detection: The function checks if the string representation—after stripping leading whitespace—begins with any formula-initiating character
- Neutralization: A single leading quote (
') is prepended, which spreadsheet programs interpret as a text indicator rather than a formula - Preservation:
Nonevalues pass through unchanged; non-triggering values return unmodified for efficiency
This approach follows the OWASP-recommended defense for CSV injection: prefixing dangerous values with a single quote to force text treatment.
Application During CSV Export
Sanitization is applied uniformly in Result.to_csv. The method transforms the result into a dictionary, flattens nested structures, and runs every value through _neutralize_csv_cell:
# user_scanner/core/result.py (excerpt)
data = self.as_dict()
...
data = {k: _neutralize_csv_cell(v) for k, v in data.items()}
writer = csv.DictWriter(output, fieldnames=CSV_FIELDS, lineterminator="")
writer.writerow(data)
The dictionary comprehension ensures no field bypasses sanitization—whether username, site_name, url, or dynamically added extra metadata.
Formatter-Level Consistency
Higher-level CSV generation in user_scanner/core/formatter.py delegates to Result.to_csv, ensuring the protection applies globally:
# user_scanner/core/formatter.py
def into_csv(results: List[Result]) -> str:
return CSV_HEADER + "\n" + "\n".join(result.to_csv() for result in results)
This design guarantees that all code paths producing CSV output—whether single results or bulk exports—inherently include the injection defense.
Practical Examples
Neutralizing a Malicious Username
from user_scanner.core.result import Result
# A username that starts with "=" would normally be treated as a formula.
r = Result.taken(username="=HYPERLINK(\"http://malicious\")", site_name="ExampleSite")
print(r.to_csv())
# Output:
# username,category,site_name,status,url,extra,media,reason
# '=HYPERLINK("http://malicious")',,ExampleSite,Found,,,
The leading single quote prevents Excel from executing the hyperlink formula.
Bulk Export with Mixed Content
from user_scanner.core.result import Result
from user_scanner.core.formatter import into_csv
results = [
Result.taken(username="+cmd|' /C calc'!A0", site_name="CalcSite"),
Result.available(username="normal_user", site_name="SafeSite")
]
csv_output = into_csv(results)
print(csv_output)
# The first row's username is prefixed with a single quote, neutralising the formula.
Key Implementation Files
| File | Purpose |
|---|---|
user_scanner/core/result.py |
Defines Result._neutralize_csv_cell and applies sanitization in to_csv() |
user_scanner/core/formatter.py |
Orchestrates CSV output via into_csv(), delegating per-row sanitization to Result |
Summary
- user-scanner prevents CSV injection by detecting formula-triggering characters and prepending a single quote to force text interpretation
_neutralize_csv_cellinresult.pyimplements the core defense, checking for=,+,-,@, tab, carriage return, and newline- Universal application: Every value passes through sanitization via dictionary comprehension in
to_csv(), with no opt-out path - Defense in depth: The
formatter.pymodule'sinto_csv()ensures all bulk exports inherit the protection automatically
Frequently Asked Questions
What characters trigger CSV injection in user-scanner's defense?
The FORMULA_TRIGGER_CHARS tuple in _neutralize_csv_cell includes: equals sign (=), plus (+), minus (-), at symbol (@), tab (\t), carriage return (\r), and newline (\n). These are the characters that spreadsheet applications interpret as formula initiators.
Does the single quote prefix appear in the exported CSV file?
Yes—the leading single quote is written to the CSV file. Spreadsheet programs display the content without the visible quote but treat it as text rather than executing it as a formula. Raw text viewers and other parsers see the quote character.
Is user-scanner's CSV injection prevention configurable or optional?
No. The sanitization is mandatory and non-configurable in the current implementation. The _neutralize_csv_cell call in to_csv() applies unconditionally to all values, ensuring consistent protection regardless of how results are generated.
How does this compare to other CSV injection defenses?
The single-quote prefix method used in user-scanner is the most reliable portable defense. Alternatives like removing dangerous characters risk data loss, while tab-separated formats only shift the attack surface. The user-scanner implementation specifically follows OWASP guidance for formula injection prevention.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →