Lemon AI Thinking Utility Parameters for Reasoning Processing
The thinking utility in Lemon AI utilizes nine core parameters—prompt, model, temperature, top_p, max_tokens, stream, enable_thinking, thinking_budget, and max_reasoning_tokens—along with a structured context object to generate and parse LLM reasoning blocks.
The thinking utility serves as the central reasoning engine in the Lemon AI framework (hexdocom/lemonai), responsible for constructing specialized prompts that instruct the LLM to wrap its internal reasoning process within <thinking> tags. According to the source code in src/utils/thinking.js, this utility manages both the parameter configuration for the LLM call and the subsequent extraction of structured reasoning from the model's response.
Core Parameters for Reasoning Processing
The thinking utility accepts a comprehensive parameter set defined in src/utils/thinking.js that controls both the LLM generation behavior and the reasoning extraction logic.
Required Inference Parameters
When invoking the thinking pipeline, the following parameters configure the base LLM request:
prompt– The complete instruction text sent to the model, often constructed viaresolveThinkingPromptto include system context and task descriptions.model– The specific model identifier (e.g., GPT-4, Claude) retrieved viagetDefaultModel()or passed explicitly to override defaults.temperature– Controls randomness in the reasoning generation, typically set between 0 and 1 for deterministic logical processing.top_p– Nucleus sampling parameter that filters token selection to the top probability mass.max_tokens– Hard limit on the total response length, encompassing both the reasoning block and final answer.
Reasoning Control Flags
Additional boolean and numeric parameters specifically govern the reasoning extraction behavior:
enable_thinking– Boolean flag that activates the reasoning parser; when true, the utility expects and extracts content wrapped in<thinking></thinking>tags.thinking_budget– Numeric allocation that may limit the token count reserved for the internal reasoning process versus the final output.max_reasoning_tokens– Upper bound specifically for the reasoning block content, preventing excessive token consumption during the thought process.
Streaming Configuration
stream– Boolean indicating whether to process the LLM response as a stream; when enabled alongsideenable_thinking, the streaming handler wraps reasoning segments in real-time<thinking>tags before final parsing.
Context Object Structure
Beyond direct parameters, the thinking utility accepts a context object that supplies runtime state and memory access. According to the implementation in src/agent/code-act/thinking, this object must contain:
memory– A memory instance providinggetMessages()andaddMessage()methods for conversation history retrieval and storage.runtime– Object containing active port mappings (e.g.,app_port_1: 3000,app_port_2: 3001) and execution environment details.goal– String describing the high-level objective (e.g., "Generate a project scaffold").reflection– String containing previous iteration feedback or self-correction notes.files– Array of uploaded file objects; empty array[]when no attachments exist.depth– Integer indicating recursion depth for multi-step reasoning (typically starts at 1).task_manager– Optional reference to the current plan state provider.
Parsing the Reasoning Block
After the LLM returns a response, src/utils/thinking.js executes a tag-based extraction algorithm to separate internal reasoning from the actionable output.
The parser searches for the <thinking> opening tag and </thinking> closing tag within the raw content string. When detected, it slices the string at these boundaries:
- The
thinkingfield receives the complete reasoning block including the XML tags (everything from the opening<thinking>through the closing</thinking>). - The
contentfield receives the sanitized actionable reply—the portion of the response occurring after the closing</thinking>tag.
If no thinking tags are present, the entire response populates the content field while thinking remains empty.
Implementation Flow: Standard vs. Local Processing
The thinking() function serves as the primary entry point, which branches based on model subscription status:
- Model Selection – The utility first calls
getDefaultModel()to determine the target LLM. - Routing Decision – If the model is not a subscribed cloud service, execution routes to
thinking_local(). - Prompt Resolution –
thinking_local()invokesresolveThinkingPrompt, passing all nine core parameters plus the context object to construct the final prompt string. - Execution – The LLM receives the configured request; streaming handlers wrap reasoning segments in
<thinking>tags whenenable_thinkingandstreamare both active. - Extraction – The response passes through the tag parser, returning an object with separate
thinkingandcontentfields.
Code Example: Invoking the Thinking Utility
The following example demonstrates proper parameter initialization and context preparation when calling the reasoning processor:
// Import the thinking utility from the agent code-act module
const thinking = require('@src/agent/code-act/thinking');
// Configure the context object with required runtime state
const context = {
memory, // Memory instance with getMessages()/addMessage()
runtime: {
app_port_1: 3000,
app_port_2: 3001
},
goal: 'Generate a project scaffold',
reflection: '',
files: [], // Empty array when no files uploaded
depth: 1,
task_manager: taskMgr // Optional: provides current plan state
};
// Execute reasoning with user requirement
(async () => {
const result = await thinking(
'Create a Node.js project with ESLint',
context
);
console.log('Reasoning block:', result.thinking);
console.log('Final answer:', result.content);
})();
This invocation triggers the full parameter resolution pipeline: thinking() selects the default model, routes to thinking_local() for prompt construction using all listed parameters, executes the LLM call, and returns the parsed result containing both the raw reasoning XML and the cleaned actionable response.
Summary
- The thinking utility in
src/utils/thinking.jsprocesses nine core parameters:prompt,model,temperature,top_p,max_tokens,stream,enable_thinking,thinking_budget, andmax_reasoning_tokens. - A structured context object containing
memory,runtime,goal,reflection,files,depth, and optionaltask_managersupplies runtime state for prompt construction. - The parser extracts content wrapped in
<thinking>XML tags, separating internal reasoning from final output into distinctthinkingandcontentfields. - Execution branches between standard and local processing via
thinking_local()based on model subscription status. - When
enable_thinkingis true, the streaming handler wraps reasoning segments in real-time tags before final extraction.
Frequently Asked Questions
What parameters control the reasoning token limits in Lemon AI?
The thinking utility accepts thinking_budget and max_reasoning_tokens to constrain how many tokens the LLM allocates to its internal reasoning process versus the final answer. The standard max_tokens parameter sets the overall response ceiling, while these reasoning-specific parameters prevent the model from consuming the entire allocation during the thought phase.
How does the thinking utility separate reasoning from the final answer?
According to src/utils/thinking.js, the utility parses the raw LLM response for <thinking> and </thinking> XML tags. Content between these tags populates the thinking field (including the tags themselves), while everything after the closing tag populates the content field. This separation allows agent logic to inspect reasoning metadata while acting only on the sanitized output.
When does the thinking utility use local processing instead of cloud models?
The thinking() function checks model subscription status immediately after calling getDefaultModel(). If the selected model is not a subscribed service, execution routes to thinking_local(), which handles prompt resolution and inference using local parameter configurations. This branching ensures the utility adapts to both hosted APIs and local LLM deployments.
What must the context object contain for the thinking utility to function?
The context object requires a memory instance with getMessages() and addMessage() methods, a runtime object with environment details like port mappings, a string goal, a reflection field (which may be empty), a files array, and an integer depth. An optional task_manager may be included to provide the current plan state during multi-step reasoning tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →