Elasticsearch Term Query on Text Fields: Best Practices and Analyzer Interaction
Use term queries on text fields only when matching exact analyzed tokens; for user searches, prefer match queries that apply the analyzer, or use a keyword sub-field for exact string matching.
When working with Elasticsearch, understanding how the term query interacts with text fields is critical for accurate search results. Unlike full-text queries, a term query seeks exact matches in the inverted index without applying the field's analyzer to the query string. This article examines the source code implementation in elastic/elasticsearch to explain when and how to use term queries on analyzed text fields.
How Term Queries Work with Text Fields in Elasticsearch
A term query operates on the exact terms stored in the inverted index. When you index a document with a text field, Elasticsearch runs the configured analyzer during indexing, producing tokens that are stored in the index. The term query does not run the analyzer on your query input—it looks for the exact byte sequence you provide.
This creates a common pitfall: searching for "The Great Gatsby" with a term query on a standard analyzed text field returns no results, because the analyzer lowercased the text and removed punctuation, storing tokens like "the", "great", and "gatsby".
The Inverted Index vs. Doc-Values Fallback
According to the source code in TextFieldMapper.java, the termQuery method first checks whether the field has indexed terms:
@Override
public Query termQuery(Object value, SearchExecutionContext context) {
if (indexType().hasTerms()) {
return super.termQuery(value, context); // Fast inverted index path
}
failIfNotIndexedNorDocValuesFallback(context);
if (usesBinaryDocValues) {
return new SlowCustomBinaryDocValuesTermQuery(name(),
indexedValueForSearch(value));
} else {
return SortedSetDocValuesField.newSlowExactQuery(name(),
indexedValueForSearch(value));
}
}
(see server/src/main/java/org/elasticsearch/index/mapper/TextFieldMapper.java, lines 558‑571)
If indexType().hasTerms() returns true, Elasticsearch uses the fast inverted index lookup via super.termQuery(). If the field has no indexed terms (for example, if index: false is set but doc_values: true), the query falls back to a slow doc-values scan using SortedSetDocValuesField.newSlowExactQuery or SlowCustomBinaryDocValuesTermQuery.
Source Code Analysis: Term Query Execution Path
The TermQueryBuilder class serves as the DSL entry point for term queries. As shown in TermQueryBuilder.java, it extends BaseTermQueryBuilder and passes the value directly to the low-level Lucene query without analysis:
public class TermQueryBuilder extends BaseTermQueryBuilder<TermQueryBuilder> {
public static final String NAME = "term";
private boolean caseInsensitive = DEFAULT_CASE_INSENSITIVITY;
// Value is passed verbatim to the underlying TermQuery
}
(see server/src/main/java/org/elasticsearch/index/query/TermQueryBuilder.java, lines 30‑38)
For the newer match_only_text field type, implemented in MatchOnlyTextFieldMapper.java, the behavior is similar: if the field contains terms, it wraps the query in a ConstantScoreQuery; otherwise, it falls back to the same doc-values logic as standard text fields.
Best Practices for Elasticsearch Term Queries on Text Fields
-
Use
matchormatch_phrasefor user-facing search
These queries apply the field's analyzer to the input, ensuring that tokenization matches the indexed terms. This is the correct approach for full-text search. -
Create a
keywordsub-field for exact matches
Map your text field with a multi-field that includes akeywordvariant:"title": { "type": "text", "fields": { "keyword": { "type": "keyword" } } }Run
termqueries againsttitle.keywordfor exact string matching. -
Avoid
termqueries on analyzedtextfields unless you know the exact token
If you must query thetextfield directly, ensure your query value matches the analyzer output exactly (e.g., lowercased for thestandardanalyzer). -
Do not enable
fielddataontextfields to supporttermqueries
Loading fielddata into the JVM heap is expensive and can cause out-of-memory errors. Use thekeywordsub-field pattern instead. -
Be aware of doc-values fallback performance
If atextfield hasindex: falsebutdoc_values: true, atermquery will execute a slow scan. Ensuredoc_valuesis enabled (the default fortextfields in recent versions) if you must query unindexed fields, but prefer indexed fields for performance.
Practical Examples: Mapping and Query Strategies
Recommended Mapping Structure
PUT /books
{
"mappings": {
"properties": {
"title": {
"type": "text",
"analyzer": "standard",
"fields": {
"keyword": {
"type": "keyword",
"ignore_above": 256
}
}
}
}
}
}
Correct: Term Query on Keyword Sub-field
GET /books/_search
{
"query": {
"term": {
"title.keyword": {
"value": "The Great Gatsby"
}
}
}
}
Incorrect: Term Query on Analyzed Text Field
This query returns no results because the standard analyzer lowercases tokens during indexing:
GET /books/_search
{
"query": {
"term": {
"title": {
"value": "The Great Gatsby"
}
}
}
}
Corrected: Match Query for Full-Text Search
GET /books/_search
{
"query": {
"match": {
"title": "The Great Gatsby"
}
}
}
Term Query with Doc-Values Fallback
When querying a text field with index: false but doc_values: true:
GET /books/_search
{
"query": {
"term": {
"title": {
"value": "gatsby",
"case_insensitive": false
}
}
}
}
Summary
- Term queries perform exact matching against the inverted index without applying the field's analyzer, making them unsuitable for full-text search on analyzed
textfields. - Use
matchormatch_phrasequeries for user search input to ensure the query string undergoes the same analysis as the indexed text. - Implement
keywordsub-fields for exact string matching requirements, querying them withtermqueries instead of the parenttextfield. - Avoid fielddata on
textfields due to memory overhead; rely on doc-values or keyword fields instead. - Understand the doc-values fallback in
TextFieldMapper.java(lines 558-571) for unindexed fields, recognizing that this path usesSlowCustomBinaryDocValuesTermQueryorSortedSetDocValuesField.newSlowExactQuerywith significant performance penalties.
Frequently Asked Questions
Why does my term query return no results on a text field?
A term query searches for the exact byte sequence you provide without running the field's analyzer. If your text field uses the standard analyzer, it lowercases tokens and removes punctuation during indexing. Searching for "The Great Gatsby" (mixed case) fails because the index contains "the", "great", and "gatsby". Use a match query instead, or query a keyword sub-field for exact matches.
What is the difference between term and match queries in Elasticsearch?
A term query performs exact value matching against the inverted index and does not analyze the query string. A match query is a full-text search that applies the field's analyzer to the query input, tokenizing and normalizing it to match the indexed terms. Use term for exact values (IDs, enums, keywords) and match for user-provided search text.
When should I use the keyword sub-field instead of querying the text field directly?
Always use the keyword sub-field when you need exact string matching, sorting, or aggregations on textual data. The keyword type stores the original string value unchanged, making it compatible with term queries. Querying the parent text field with a term query only works if you know the exact analyzed token, which is fragile and analyzer-dependent. The keyword sub-field provides a reliable, performant path for exact matches.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →