How to Validate Input and Handle Edge Cases with OpenMed's validate_input
The validate_input function in openmed/utils/validation.py provides a centralized utility that None-checks, type-coerces, strips whitespace, enforces length constraints, and detects suspicious content before returning a cleaned string or raising descriptive ValueError exceptions.
OpenMed is an open-source medical NLP framework that provides robust preprocessing utilities to ensure clinical text is safe for downstream models. The validate_input function serves as the primary gatekeeper for input sanitation, handling everything from null values to potential denial-of-service payloads. Understanding how to leverage this utility allows you to prevent runtime errors and protect your inference pipeline from malformed or malicious data.
How validate_input Works
The validation pipeline implemented in openmed/utils/validation.py follows a strict sequence of checks to guarantee that downstream components receive well-formed text.
None Checking and Type Coercion
The function first verifies that the input is not None, raising a ValueError with the message "Input text cannot be None" unless you explicitly set allow_empty=True. For non-string inputs, it automatically coerces the value using str(), ensuring that integers, floats, or other objects are safely converted to strings before processing.
Whitespace and Empty String Handling
After coercion, the function calls text.strip() to remove accidental leading or trailing whitespace. If the resulting string is empty and allow_empty=False (the default), it raises a ValueError stating "Input text cannot be empty". This prevents empty documents from wasting compute cycles in the NER pipeline.
Length Validation
The function enforces a min_length parameter (defaulting to 1) and an optional max_length to guarantee sufficient context for the model while protecting against memory exhaustion from extremely long inputs. Violations raise specific ValueError messages indicating whether the text is too short or too long.
Suspicious Content Detection
Before returning the cleaned string, validate_input calls the private helper _contains_suspicious_content. This detector identifies pathological inputs such as sequences with over 100 repeated characters, text containing more than 50% special symbols, or large non-ASCII blocks that could indicate binary blobs or DoS attacks. When suspicious content is detected, the function raises a ValueError with the message "Input text contains suspicious content".
Handling Common Edge Cases
The validate_input function explicitly handles several edge cases that commonly cause failures in production medical NLP pipelines.
Noneinput – RaisesValueErrorunlessallow_empty=True, preventingAttributeErrorwhen calling string methods onNone.- Empty strings – Returns
''only whenallow_empty=True; otherwise raisesValueError. - Non-string types – Silently coerces values like
123to"123"viastr(). - Length violations – Enforces
min_lengthandmax_lengthconstraints with descriptive error messages. - Malicious payloads – Blocks inputs with repetitive characters, high special-character density, or non-ASCII anomalies.
These behaviors are verified in the unit tests located in tests/unit/test_utils.py, specifically lines 57‑101 for normal cases, lines 115‑129 for error paths, and lines 132‑140 for suspicious content detection.
Practical Code Examples
Basic Usage
from openmed.utils.validation import validate_input
# Normal text – passes through unchanged
clean_text = validate_input("Patient presents with cough and fever.")
print(clean_text) # → Patient presents with cough and fever.
Enforcing Length Constraints
# Require at least 50 characters, reject anything over 500
text = "Short note."
try:
validate_input(text, min_length=50, max_length=500)
except ValueError as e:
print(e) # → Input text too short. Minimum length: 50
Allowing Empty Strings
empty = validate_input("", allow_empty=True)
print(repr(empty)) # → ''
Detecting Suspicious Content
from openmed.utils.validation import validate_input
# Very long repeated character sequence – will be blocked
bad = "a" * 200
try:
validate_input(bad)
except ValueError as e:
print(e) # → Input text contains suspicious content
Integrating into a Processing Pipeline
def process_document(raw_text: str):
# Validate first – guarantees safe downstream processing
text = validate_input(raw_text, min_length=10, max_length=2000)
# ... feed `text` to the OpenMed NER model ...
predictions = run_ner(text)
return predictions
Key Implementation Files
openmed/utils/validation.py– Core validation utilities includingvalidate_inputand the private_contains_suspicious_contenthelper.tests/unit/test_utils.py– Unit tests covering normal cases (lines 57‑101), error paths (lines 115‑129), and suspicious content detection (lines 132‑140).openmed/processing/text.py– Example consumer that usesvalidate_inputwhen preparing raw clinical notes for NER.openmed/service/app.py– FastAPI service entry point that validates input before model inference.
Summary
validate_inputinopenmed/utils/validation.pyprovides centralized input validation for the OpenMed medical NLP framework.- The function performs type coercion, whitespace stripping, length validation, and suspicious content detection in a strict pipeline.
- Edge cases like
None, empty strings, and non-string types are handled gracefully with configurableallow_emptyparameters. - Security features include detection of repeated characters, high special-character density, and non-ASCII blocks to prevent DoS attacks.
- Integration is straightforward: call
validate_inputat the entry point of any function that processes raw clinical text before passing data to the NER model.
Frequently Asked Questions
What happens if I pass None to validate_input?
Unless you set allow_empty=True, the function raises a ValueError with the message "Input text cannot be None". This prevents downstream AttributeError exceptions when string methods are called on None objects.
How does validate_input handle non-string data types?
The function automatically converts non-string values to strings using str(), so passing an integer like 123 returns the string "123". This coercion ensures consistent return types while allowing flexible input from various sources.
What triggers suspicious content detection in OpenMed's validation?
The private helper _contains_suspicious_content identifies pathological inputs such as sequences with over 100 repeated characters, text where more than 50% of characters are special symbols, or large non-ASCII blocks. When detected, validate_input raises a ValueError stating "Input text contains suspicious content".
Can I allow empty strings while still validating other constraints?
Yes. By setting allow_empty=True, empty strings pass through the empty-string guard and return '' without raising an error. However, if you also specify a min_length greater than 0, the empty string will still fail the minimum length check.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →