How to Use Regex and Text-Based Search in Scrapling Selectors
Scrapling provides find_by_text() for exact or partial string matching and find_by_regex() for pattern-based extraction, both accessible through the Selector class in scrapling/parser.py.
Scrapling is a high-performance web scraping library that simplifies HTML parsing with an intuitive selector API. When you need to locate elements based on their visible content rather than structural attributes, regex and text-based search in Scrapling selectors offer precise targeting capabilities. This guide examines the implementation details and practical usage of these methods as defined in the source code.
Text-Based Element Discovery with find_by_text
The find_by_text method, implemented in scrapling/parser.py (lines 1057-1084), enables exact or partial string matching against element text content.
This method traverses all descendant elements via _find_all_elements_with_spaces, normalizes case when case_sensitive=False, and optionally collapses whitespace with clean_match. It performs either a direct equality check (text == node_text) or a containment check (text in node_text) when partial=True.
from scrapling import Page
page = Page(url="https://example.com")
add_to_cart = page.find_by_text("Add to cart")
print(add_to_cart.href)
Pattern Matching with find_by_regex
For complex extraction scenarios, find_by_regex (lines 1122-1150 in scrapling/parser.py) leverages regular expressions to match dynamic text patterns.
This method retrieves each element's TextHandler and calls TextHandler.re with the supplied pattern. It uses the check_match=True flag to short-circuit evaluation, adding only matching elements to the results.
stock = page.find_by_regex(r"In stock: \d+", clean_match=True)
print(stock.text) # → "In stock: 42"
Core Regex Engine in TextHandler
The low-level regex implementation resides in scrapling/core/custom_types.py (lines 146-182) within the TextHandler.re method.
This engine:
- Compiles patterns with
re.IGNORECASEwhencase_sensitive=False - Optionally sanitizes whitespace via
clean_match - Returns a list of
TextHandlerobjects or a boolean whencheck_match=True
Customizing Search Behavior
Both methods accept parameters that fine-tune matching behavior:
| Parameter | Description | Effect |
|---|---|---|
first_match |
True returns only the first match; False returns all matches |
Controls result type (Selector vs Selectors) |
partial (find_by_text only) |
When True, searches for substring containment rather than exact equality |
Enables "contains" searches |
case_sensitive |
When False, converts both source and pattern to lowercase before comparison |
Simplifies case-insensitive matching |
clean_match |
Strips extra whitespace and normalizes spaces before matching | Handles HTML-induced line breaks or multiple spaces |
Practical Scrapling Examples
Partial Match Collection
Locate all elements containing a specific keyword:
offers = page.find_by_text("discount", first_match=False, partial=True)
for offer in offers:
print(offer.text)
Multiple Regex Extractions
Extract all price patterns from a page:
prices = page.find_by_regex(r"\$\d+\.\d{2}", first_match=False, case_sensitive=True)
for price in prices:
print(price.text) # → "$19.99", "$24.99", etc.
Case-Insensitive Date Matching
Find dates with normalized whitespace:
dates = page.find_by_regex(r"\bJan\s+\d{1,2},\s+202[0-5]\b",
case_sensitive=False,
clean_match=True,
first_match=False)
Summary
find_by_textinscrapling/parser.pyprovides exact or partial string matching with optional case and whitespace normalization.find_by_regexleverages theTextHandler.reengine inscrapling/core/custom_types.pyfor pattern-based extraction.- Both methods return either a single
Selectoror aSelectorscollection based on thefirst_matchparameter. - Use
clean_match=Trueto handle HTML formatting artifacts, andcase_sensitive=Falsefor flexible matching.
Frequently Asked Questions
How does find_by_text differ from standard CSS selectors?
While CSS selectors target elements by tag name, class, or ID, find_by_text searches the actual text content within elements. According to the implementation in scrapling/parser.py, this method traverses descendant elements and performs string comparison against normalized text content, making it ideal for locating buttons, headings, or labels with known text.
Can I use compiled regex patterns with find_by_regex?
Yes, the TextHandler.re method in scrapling/core/custom_types.py accepts both string patterns and compiled regex objects. When you pass a string, the method compiles it internally with re.IGNORECASE if case_sensitive=False. Using pre-compiled patterns can improve performance when running the same regex against multiple selectors.
What happens when first_match=False returns no results?
When first_match=False is set, both find_by_text and find_by_regex return a Selectors object (essentially a list-like collection). If no elements match the criteria, this collection will be empty rather than raising an exception. You can check the length or iterate safely without error handling for null results.
How does clean_match handle whitespace normalization?
The clean_match parameter triggers whitespace sanitization in the TextHandler class. When enabled, the implementation strips leading and trailing whitespace and collapses multiple consecutive spaces or line breaks into single spaces. This is particularly useful for HTML where formatting introduces \n or \t characters between inline elements, ensuring that patterns like "Price: $20" match even when the source HTML contains "Price:\n$20".
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →