How FckSignups Handles Hyphenated Search Terms: Tokenization and Exact Match Scoring

FckSignups processes hyphenated search queries by splitting them into individual tokens using a regex-based tokenizer in src/hooks/useTools.ts, requiring every resulting keyword to appear in a tool's name, description, or tags for a result to qualify.

When users enter hyphenated phrases like "video-editor" into the BraveOPotato/FckSignups search interface, the application treats these as separate keywords rather than literal strings. This behavior stems from a custom normalization pipeline implemented in the React hooks layer, ensuring that compound terms match tools containing either word independently or together.

Tokenization Logic in src/hooks/useTools.ts

The search pipeline begins in src/hooks/useTools.ts, which exports a tokenize function designed to normalize user input. This function applies the regular expression /[^a-z0-9+]+/ to split query strings on any character sequence that isn't a lowercase letter, number, or plus sign.

When processing a term like "video-editor", the regex matches the hyphen as a delimiter, producing the array ["video", "editor"]. The tokenizer strips case sensitivity and treats multiple hyphens, spaces, or other special characters identically, collapsing them into single split points.

React Optimization with useMemo

To prevent redundant processing during component re-renders, the hook wraps tokenization in a useMemo hook:

searchKeywords = useMemo(() => tokenize(searchQuery), [searchQuery])

This ensures the keyword array only recalculates when the actual search string changes, providing stable references for downstream filtering operations.

Search Scoring and Filtering Architecture

After tokenization, the matchScore function evaluates potential matches by constructing a searchable haystack from three tool properties: the name, description, and tags array. The function concatenates these fields into a single lowercase string, then counts how many query tokens appear within it:

keywords.filter((kw) => haystack.includes(kw)).length

Exact Match Enforcement

FckSignups implements strict AND-logic filtering. The system retains only tools where the match score equals the total number of generated tokens. For a hyphenated query producing two tokens, both must appear in the tool's metadata—partial matches scoring 1 out of 2 are filtered out. This guarantees that multi-term searches return only highly relevant results.

Practical Implementation Example

The following TypeScript code illustrates how the tokenization and scoring pipeline handles hyphenated input against tool data:

import { tokenize, matchScore } from "./hooks/useTools";

// User query containing a hyphen
const query = "video-editor";

// Tokenize splits on the hyphen → ["video", "editor"]
const tokens = tokenize(query);

// Example tool object matching the structure in src/types/index.ts
const tool = {
  name: "Video Editor Pro",
  description: "A powerful video-editing suite",
  tags: ["media", "video"],
  stars: 120,
};

// Calculate relevance score (returns 2 since both tokens match)
const score = matchScore(tool, tokens);

// The tool qualifies for display only if score === tokens.length (2)

In the live application, this logic integrates with default tool data from src/constants/fallbackData.ts when remote fetching fails, ensuring consistent search functionality across network conditions.

Summary

  • Normalization: The tokenize function in src/hooks/useTools.ts splits hyphenated terms using the regex /[^a-z0-9+]+/, converting "video-editor" into separate keywords.
  • Scoring: The matchScore function checks token presence against tool name, description, and tags fields defined in src/types/index.ts.
  • Filtering: Only tools matching all generated tokens (AND logic) appear in results, with partial matches excluded.
  • Performance: useMemo caches token arrays to optimize React rendering cycles during user input.

Frequently Asked Questions

What regex pattern does FckSignups use to split hyphenated search terms?

According to the source code in src/hooks/useTools.ts, the tokenizer uses the regular expression /[^a-z0-9+]+/ to split strings. This pattern targets any sequence of characters that isn't a lowercase letter, digit, or plus sign, effectively treating hyphens, spaces, and other punctuation as delimiters.

Which tool properties does the search algorithm evaluate?

The matchScore function constructs its searchable haystack from three specific fields: the tool's name, description, and tags array. The algorithm performs case-insensitive matching against the concatenated content of these properties.

Does FckSignups use AND or OR logic for multi-token searches?

The implementation uses strict AND logic. For a query like "video-editor" that tokenizes into two separate keywords, a tool must contain both "video" and "editor" in its metadata to qualify. The filtering condition requires score === tokens.length, ensuring partial matches are excluded from results.

How does the application prevent performance issues during rapid typing?

The useTools hook memoizes token generation using React's useMemo(() => tokenize(searchQuery), [searchQuery]). This optimization ensures that the expensive regex splitting and array generation only occur when the actual search string changes, not on every component re-render.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →