How Plausible Analytics Resolves Segments and Filters in the Query Pipeline
Plausible Analytics resolves segments by replacing ["is", "segment", [id]] placeholders with their stored filter definitions during the query building phase, transforming them into concrete filter trees before ClickHouse query generation.
The Plausible analytics engine allows users to reference saved segments within standard filter syntax, enabling complex, reusable audience definitions. When building a query, the pipeline extracts these segment references, loads their definitions from the database, and expands them into native filter AST nodes. This resolution process ensures that downstream components work with a unified filter structure regardless of whether filters were written inline or pulled from saved segments.
Query Pipeline Entry Point
The resolution process begins in Plausible.Stats.QueryBuilder.build/3, which receives a %ParsedQueryParams{} struct containing the raw filter list from the client. The builder immediately delegates to the segments module to handle any segment references before proceeding with further validation.
In lib/plausible/stats/query_builder.ex (lines 79-84), the code wraps segment preloading and resolution in a with statement:
with {:ok, preloaded_segments} <-
Segments.Filters.preload_needed_segments(site, parsed_query_params.filters),
{:ok, filters} <-
Segments.Filters.resolve_segments(parsed_query_params.filters, preloaded_segments) do
{:ok, struct!(parsed_query_params, filters: filters)}
end
If segment resolution succeeds, the builder continues with remaining validation steps (ordering, dimensions, metrics) and ultimately produces a %Plausible.Stats.Query{} for the ClickHouse query generator.
Extracting Segment IDs from Filter AST
Before loading any data from the database, the system must identify which segments are referenced. The function Plausible.Segments.Filters.get_segment_ids/1 traverses the filter AST using Filters.traverse/1 and collects every numeric ID appearing in a "segment" clause.
From lib/plausible/segments/filters.ex (lines 20-38):
ids =
filters
|> Filters.traverse()
|> Enum.flat_map(fn
{[_operation, "segment", clauses], _} -> clauses
_ -> []
end)
This function also enforces a hard limit of ten segment filters per query, returning a %QueryError{} if the client exceeds this threshold. This prevents excessive database load from deeply nested segment definitions.
Preloading Segment Definitions from the Database
Once the system knows which segments are needed, Plausible.Segments.Filters.preload_needed_segments/2 fetches the corresponding %Segment{} records using Segments.get_many/3. Each segment stores its filter definition in a JSON column segment_data["filters"].
In lib/plausible/segments/filters.ex (lines 41-73), the raw JSON is parsed back into a filter AST using ApiQueryParser.parse_filters/1. The result is a lookup map %{segment_id => filter_tree} that will be substituted into the original query.
# Simplified concept of the preloading result
%{
42 => [[:is, "visit:country", ["PL"]], [:contains, "event:goal", ["Signup"]]],
43 => [[:is, "visit:device", ["Desktop"]]]
}
The retrieval helpers (get_many/3, get_one/4) are defined in lib/plausible/segments/segments.ex, which implements the Ecto schema for the segments table.
Resolving Segments into Concrete Filter Trees
The core transformation happens in Plausible.Segments.Filters.resolve_segments/2, defined in lib/plausible/segments/filters.ex (lines 123-134). This function performs two passes over the filter list:
replace_segment_with_filter_tree/2swaps a["segment", [id]]clause for the concrete filter tree. If a single ID is present, the result is wrapped in[:and, …]; if multiple IDs are present, it creates an[:or, …]of[:and, …]groups.expand_first_level_and_filters/1flattens top-level[:and, clauses]nodes into a plain list for uniform downstream processing.
original_filters
|> Filters.transform_filters(fn f -> replace_segment_with_filter_tree(f, preloaded_segments) end)
|> Filters.transform_filters(&expand_first_level_and_filters/1)
Code Example: Single Segment Resolution
Consider a client query referencing segment 42:
filters = [
[:is, "visit:entry_page", ["/blog"]],
[:is, "segment", [42]]
]
After resolution, assuming segment 42 contains filters for country and goal:
[
[:is, "visit:entry_page", ["/blog"]],
[:is, "visit:country", ["PL"]],
[:contains, "event:goal", ["Signup"]]
]
Code Example: Multiple Segment Combination
When multiple segment IDs are provided, they are combined with OR logic:
filters = [
[:is, "segment", [1, 2]]
]
# Resolves to:
[
[:or, [
[:and, [[:contains, "event:goal", ["Signup"]], [:is, "visit:country", ["PL"]]]],
[:and, [[:contains, "event:goal", ["Purchase"]], [:is, "visit:country", ["EE"]]]]
]]
]
Segment Storage and Data Structure
Segments are stored as regular Ecto schemas defined in lib/plausible/segments/segments.ex. The segment_data column holds a JSON map that must contain a "filters" key with the filter AST. The schema enforces no specific structure beyond valid JSON, but the query pipeline expects this key to exist when resolving segments.
The separation between segment storage (persisting reusable definitions) and filter resolution (runtime AST transformation) allows the rest of the analytics engine to remain agnostic about whether a filter originated from a saved segment or inline user input.
Summary
- Resolution happens early:
Plausible.Stats.QueryBuilder.build/3triggers segment resolution before any ClickHouse query generation occurs. - ID extraction limits: The system enforces a maximum of ten segment references per query via
get_segment_ids/1. - JSON storage: Segments store their filter definitions in
segment_data["filters"], parsed byApiQueryParser.parse_filters/1during preloading. - AST transformation:
resolve_segments/2replaces placeholders with concrete trees, wrapping single segments in[:and, …]and multiple segments in[:or, …]of[:and, …]groups. - Flattening: The final pass expands top-level AND clauses to ensure downstream components receive a uniform filter list.
Frequently Asked Questions
What is the maximum number of segments allowed per query in Plausible Analytics?
Plausible enforces a hard limit of ten segment filters per query in the get_segment_ids/1 function. If a client submits a filter list containing more than ten segment references, the system returns a %QueryError{} and halts query processing. This limit prevents excessive database load from complex nested segment definitions.
How does Plausible combine multiple segments in a single filter clause?
When a filter contains multiple segment IDs such as [:is, "segment", [1, 2]], the resolution logic creates an OR-of-AND structure. Each segment’s filter tree is wrapped in its own [:and, …] node, and these are combined under a single [:or, …] node. This ensures that matching any of the specified segments satisfies the filter condition.
Where are segment filter definitions stored in the Plausible codebase?
Segment definitions are stored in the segment_data JSON column of the segments table, accessible via the Plausible.Segments.Segment Ecto schema in lib/plausible/segments/segments.ex. The JSON must contain a "filters" key holding the filter AST. When resolving queries, the system fetches these records using Segments.get_many/3 and parses the JSON through ApiQueryParser.parse_filters/1.
What happens if a query references a segment that does not exist?
If preload_needed_segments/2 cannot find a segment ID referenced in the filters, the function returns an error tuple, causing the with chain in query_builder.ex to fail. This prevents the query from proceeding to the ClickHouse generation phase and ensures the user receives a validation error rather than incomplete or incorrect analytics data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →