AI Chat Response System Data Flow in Maybe: A Complete Technical Guide
The AI chat response system in Maybe processes user messages through a tightly-coupled Rails pipeline that moves from browser controllers to background jobs, then streams LLM responses via Turbo Streams while persisting conversation state using models like Chat and AssistantMessage.
The maybe-finance/maybe repository implements its AI chat feature as an asynchronous, event-driven pipeline spanning from initial prompt submission to real-time response rendering. This system leverages OpenAI's API with a clear separation between web request handling and LLM processing, utilizing specific controller actions, background jobs, and model callbacks to manage data flow. Understanding this architecture is essential for developers extending chat functionality or debugging conversation state.
How the AI Chat Response System Works
The data flow follows a sequential 10-step pipeline that begins with a user POST request and ends with a Turbo Stream broadcast. Each step is implemented in specific files within the Rails application.
Step 1: Initiating a Chat via Controllers
When a user starts a new conversation, the UI posts to ChatsController#create. According to the maybe-finance/maybe source code, this controller delegates to Current.user.chats.start!, which instantiates a Chat record and creates the initial UserMessage.
# File: app/controllers/chats_controller.rb
# Entry point for new conversations
chat = Current.user.chats.start!("How much did I spend?", model: "gpt-4.1")
Step 2: Persisting Messages and Scheduling Jobs
For follow-up messages, MessagesController#create handles the POST request. The implementation in app/models/user_message.rb uses an after_create_commit callback named request_response_later to enqueue background processing asynchronously without blocking the HTTP response.
# File: app/models/user_message.rb
# Creates message and schedules async processing
message = UserMessage.create!(chat: chat, content: "Show breakdown", ai_model: "gpt-4.1")
# Automatically triggers: AssistantResponseJob.perform_later(message)
Step 3: Background Processing with AssistantResponseJob
The AssistantResponseJob class in app/jobs/assistant_response_job.rb serves as the asynchronous bridge between the web layer and the LLM. The job simply calls message.request_response, which delegates to the chat orchestration layer to begin the actual LLM interaction.
# File: app/jobs/assistant_response_job.rb
AssistantResponseJob.perform_later(message) # Queued via Active Job
AssistantResponseJob.perform_now(message) # Synchronous execution for testing
Step 4: Orchestrating the Assistant and Provider Resolution
The UserMessage#request_response method forwards the request to Chat#ask_assistant in app/models/chat.rb. This method builds an Assistant instance via Assistant.for_chat, which mixes in Provided to resolve the LLM provider through Provider::Registry.for_concept(:llm).
The provider resolution in Assistant::Provided#get_model_provider supplies the concrete OpenAI client configuration. Subsequently, Assistant#respond_to creates an AssistantMessage record to store the pending response and instantiates Assistant::Responder with the message, system instructions, and a FunctionToolCaller.
Step 5: Streaming LLM Responses
Inside Assistant::Responder#respond (defined in app/models/assistant/responder.rb), the system calls llm.chat_response with a streaming proc. This proc receives "output_text" and "response" chunks, emitting events that drive the real-time update mechanism.
The first "output_text" chunk creates the AssistantMessage record via AssistantMessage#append_text!, while subsequent chunks append to the existing record. The implementation also updates Chat#latest_assistant_response_id through Chat#update_latest_response! to maintain conversation state.
Step 6: Parsing and Handling Function Calls
The OpenAI client returns raw JSON that Provider::Openai::ChatParser (in app/models/provider/openai/chat_parser.rb) translates into ChatResponse, ChatMessage, and ChatFunctionRequest objects.
If the response contains function_requests, the responder invokes FunctionToolCaller.fulfill_requests. The resulting tool calls return as function_results, triggering a follow-up LLM request via Responder#handle_follow_up_response before finalizing the assistant's text output.
Step 7: Broadcasting to the UI
The Message model in app/models/message.rb implements after_create_commit and after_update_commit callbacks that broadcast Turbo Stream updates to the messages target. This pushes real-time updates to the browser without polling, completing the data flow cycle.
Errors during streaming are rescued, logged on the Chat record via Chat#add_error, and broadcast to the UI to maintain user feedback during failures.
Code Examples for the AI Chat Response System
These practical examples demonstrate how to interact with the AI chat response system programmatically.
Start a New AI Chat
# Create a chat with an initial user prompt
chat = Current.user.chats.start!("How much did I spend on groceries last month?", model: "gpt-4.1")
# The first assistant reply streams automatically via background job
Send a Follow-Up Message
# Simulate the POST /messages endpoint behavior
message = UserMessage.create!(
chat: chat,
content: "Show me the breakdown by category.",
ai_model: "gpt-4.1"
)
# The after_create_commit callback queues AssistantResponseJob automatically
Manually Trigger Background Processing
# Useful for testing or synchronous scripts
AssistantResponseJob.perform_now(message)
Inspect the Latest Assistant Reply
# Access the persisted assistant message
assistant_msg = chat.assistant_message
puts assistant_msg.content
Key Files in the Maybe AI Chat Architecture
Understanding these source files is critical for modifying the AI chat response system behavior:
app/models/chat.rb- Orchestrates assistant calls viaask_assistant, manages conversation state, and updateslatest_assistant_response_idapp/models/user_message.rb- Persists user input and schedules async processing viarequest_response_laterapp/models/message.rb- Base class with Turbo Stream broadcast hooks and status enumapp/models/assistant.rb- Entry point that wraps provider configuration and builds the responder viarespond_toapp/models/assistant/responder.rb- Handles streaming LLM output and function call loopsapp/models/assistant/provided.rb- Resolves LLM provider viaget_model_providerand registry lookupapp/models/provider/openai/chat_parser.rb- Translates OpenAI JSON into domain objectsapp/models/provider/openai/chat_config.rb- Configures OpenAI client parametersapp/jobs/assistant_response_job.rb- Active Job wrapper for asynchronous processingapp/controllers/chats_controller.rb- Routes initial chat creationapp/controllers/messages_controller.rb- Routes follow-up message creation
Summary
- The AI chat response system uses a controller-to-job pipeline where
MessagesController#createinitiates async processing viaAssistantResponseJob UserMessageandChatmodels inapp/models/chat.rbcoordinate to persist state and delegate to theAssistantclass- The
Assistant::Responderstreams LLM output in real-time, parsing responses withProvider::Openai::ChatParser - Function calls trigger a follow-up request loop via
FunctionToolCallerbefore finalizing the assistant message - Turbo Stream broadcasts in
app/models/message.rbpush updates to the browser without requiring page reloads
Frequently Asked Questions
How does the AI chat response system handle real-time updates?
The system uses Turbo Stream broadcasts defined in app/models/message.rb via after_create_commit and after_update_commit callbacks. When AssistantMessage content is appended during streaming, these callbacks push updates to the messages target in the browser, creating a real-time chat experience without polling.
What triggers the background job for AI responses?
The after_create_commit callback in app/models/user_message.rb named request_response_later automatically enqueues AssistantResponseJob.perform_later(message) immediately after the database commits the user message record. This ensures the LLM request happens asynchronously without blocking the HTTP response.
How does Maybe support different LLM providers?
The Assistant class mixes in Provided (from app/models/assistant/provided.rb), which looks up the provider via Provider::Registry.for_concept(:llm). This registry pattern allows the system to resolve the concrete OpenAI client or other LLM implementations dynamically based on configuration.
Where is the streaming LLM response parsed?
The Provider::Openai::ChatParser class in app/models/provider/openai/chat_parser.rb handles translation of raw OpenAI JSON into structured domain objects like ChatResponse and ChatFunctionRequest. The Assistant::Responder uses this parser to process streaming chunks and detect function call requests.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →