How to Validate Input and Handle Edge Cases with OpenMed's validate_input

The validate_input function in openmed/utils/validation.py provides a centralized utility that None-checks, type-coerces, strips whitespace, enforces length constraints, and detects suspicious content before returning a cleaned string or raising descriptive ValueError exceptions.

OpenMed is an open-source medical NLP framework that provides robust preprocessing utilities to ensure clinical text is safe for downstream models. The validate_input function serves as the primary gatekeeper for input sanitation, handling everything from null values to potential denial-of-service payloads. Understanding how to leverage this utility allows you to prevent runtime errors and protect your inference pipeline from malformed or malicious data.

How validate_input Works

The validation pipeline implemented in openmed/utils/validation.py follows a strict sequence of checks to guarantee that downstream components receive well-formed text.

None Checking and Type Coercion

The function first verifies that the input is not None, raising a ValueError with the message "Input text cannot be None" unless you explicitly set allow_empty=True. For non-string inputs, it automatically coerces the value using str(), ensuring that integers, floats, or other objects are safely converted to strings before processing.

Whitespace and Empty String Handling

After coercion, the function calls text.strip() to remove accidental leading or trailing whitespace. If the resulting string is empty and allow_empty=False (the default), it raises a ValueError stating "Input text cannot be empty". This prevents empty documents from wasting compute cycles in the NER pipeline.

Length Validation

The function enforces a min_length parameter (defaulting to 1) and an optional max_length to guarantee sufficient context for the model while protecting against memory exhaustion from extremely long inputs. Violations raise specific ValueError messages indicating whether the text is too short or too long.

Suspicious Content Detection

Before returning the cleaned string, validate_input calls the private helper _contains_suspicious_content. This detector identifies pathological inputs such as sequences with over 100 repeated characters, text containing more than 50% special symbols, or large non-ASCII blocks that could indicate binary blobs or DoS attacks. When suspicious content is detected, the function raises a ValueError with the message "Input text contains suspicious content".

Handling Common Edge Cases

The validate_input function explicitly handles several edge cases that commonly cause failures in production medical NLP pipelines.

  • None input – Raises ValueError unless allow_empty=True, preventing AttributeError when calling string methods on None.
  • Empty strings – Returns '' only when allow_empty=True; otherwise raises ValueError.
  • Non-string types – Silently coerces values like 123 to "123" via str().
  • Length violations – Enforces min_length and max_length constraints with descriptive error messages.
  • Malicious payloads – Blocks inputs with repetitive characters, high special-character density, or non-ASCII anomalies.

These behaviors are verified in the unit tests located in tests/unit/test_utils.py, specifically lines 57‑101 for normal cases, lines 115‑129 for error paths, and lines 132‑140 for suspicious content detection.

Practical Code Examples

Basic Usage

from openmed.utils.validation import validate_input

# Normal text – passes through unchanged

clean_text = validate_input("Patient presents with cough and fever.")
print(clean_text)   # → Patient presents with cough and fever.

Enforcing Length Constraints


# Require at least 50 characters, reject anything over 500

text = "Short note."
try:
    validate_input(text, min_length=50, max_length=500)
except ValueError as e:
    print(e)   # → Input text too short. Minimum length: 50

Allowing Empty Strings

empty = validate_input("", allow_empty=True)
print(repr(empty))   # → ''

Detecting Suspicious Content

from openmed.utils.validation import validate_input

# Very long repeated character sequence – will be blocked

bad = "a" * 200
try:
    validate_input(bad)
except ValueError as e:
    print(e)   # → Input text contains suspicious content

Integrating into a Processing Pipeline

def process_document(raw_text: str):
    # Validate first – guarantees safe downstream processing

    text = validate_input(raw_text, min_length=10, max_length=2000)
    
    # ... feed `text` to the OpenMed NER model ...

    predictions = run_ner(text)
    return predictions

Key Implementation Files

  • openmed/utils/validation.py – Core validation utilities including validate_input and the private _contains_suspicious_content helper.
  • tests/unit/test_utils.py – Unit tests covering normal cases (lines 57‑101), error paths (lines 115‑129), and suspicious content detection (lines 132‑140).
  • openmed/processing/text.py – Example consumer that uses validate_input when preparing raw clinical notes for NER.
  • openmed/service/app.py – FastAPI service entry point that validates input before model inference.

Summary

  • validate_input in openmed/utils/validation.py provides centralized input validation for the OpenMed medical NLP framework.
  • The function performs type coercion, whitespace stripping, length validation, and suspicious content detection in a strict pipeline.
  • Edge cases like None, empty strings, and non-string types are handled gracefully with configurable allow_empty parameters.
  • Security features include detection of repeated characters, high special-character density, and non-ASCII blocks to prevent DoS attacks.
  • Integration is straightforward: call validate_input at the entry point of any function that processes raw clinical text before passing data to the NER model.

Frequently Asked Questions

What happens if I pass None to validate_input?

Unless you set allow_empty=True, the function raises a ValueError with the message "Input text cannot be None". This prevents downstream AttributeError exceptions when string methods are called on None objects.

How does validate_input handle non-string data types?

The function automatically converts non-string values to strings using str(), so passing an integer like 123 returns the string "123". This coercion ensures consistent return types while allowing flexible input from various sources.

What triggers suspicious content detection in OpenMed's validation?

The private helper _contains_suspicious_content identifies pathological inputs such as sequences with over 100 repeated characters, text where more than 50% of characters are special symbols, or large non-ASCII blocks. When detected, validate_input raises a ValueError stating "Input text contains suspicious content".

Can I allow empty strings while still validating other constraints?

Yes. By setting allow_empty=True, empty strings pass through the empty-string guard and return '' without raising an error. However, if you also specify a min_length greater than 0, the empty string will still fail the minimum length check.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →