# AI Chat Response System Data Flow in Maybe: A Complete Technical Guide

> Explore the AI chat response system data flow in Maybe finance through a Rails pipeline. Understand browser controllers, background jobs, LLM streaming, and state persistence.

- Repository: [Maybe/maybe](https://github.com/maybe-finance/maybe)
- Tags: architecture
- Published: 2026-03-07

---

**The AI chat response system in Maybe processes user messages through a tightly-coupled Rails pipeline that moves from browser controllers to background jobs, then streams LLM responses via Turbo Streams while persisting conversation state using models like `Chat` and `AssistantMessage`.**

The maybe-finance/maybe repository implements its AI chat feature as an asynchronous, event-driven pipeline spanning from initial prompt submission to real-time response rendering. This system leverages OpenAI's API with a clear separation between web request handling and LLM processing, utilizing specific controller actions, background jobs, and model callbacks to manage data flow. Understanding this architecture is essential for developers extending chat functionality or debugging conversation state.

## How the AI Chat Response System Works

The data flow follows a sequential 10-step pipeline that begins with a user POST request and ends with a Turbo Stream broadcast. Each step is implemented in specific files within the Rails application.

### Step 1: Initiating a Chat via Controllers

When a user starts a new conversation, the UI posts to `ChatsController#create`. According to the maybe-finance/maybe source code, this controller delegates to `Current.user.chats.start!`, which instantiates a `Chat` record and creates the initial `UserMessage`.

```ruby

# File: app/controllers/chats_controller.rb

# Entry point for new conversations

chat = Current.user.chats.start!("How much did I spend?", model: "gpt-4.1")

```

### Step 2: Persisting Messages and Scheduling Jobs

For follow-up messages, `MessagesController#create` handles the POST request. The implementation in [`app/models/user_message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/user_message.rb) uses an `after_create_commit` callback named `request_response_later` to enqueue background processing asynchronously without blocking the HTTP response.

```ruby

# File: app/models/user_message.rb

# Creates message and schedules async processing

message = UserMessage.create!(chat: chat, content: "Show breakdown", ai_model: "gpt-4.1")

# Automatically triggers: AssistantResponseJob.perform_later(message)

```

### Step 3: Background Processing with AssistantResponseJob

The `AssistantResponseJob` class in [`app/jobs/assistant_response_job.rb`](https://github.com/maybe-finance/maybe/blob/main/app/jobs/assistant_response_job.rb) serves as the asynchronous bridge between the web layer and the LLM. The job simply calls `message.request_response`, which delegates to the chat orchestration layer to begin the actual LLM interaction.

```ruby

# File: app/jobs/assistant_response_job.rb

AssistantResponseJob.perform_later(message)  # Queued via Active Job

AssistantResponseJob.perform_now(message)    # Synchronous execution for testing

```

### Step 4: Orchestrating the Assistant and Provider Resolution

The `UserMessage#request_response` method forwards the request to `Chat#ask_assistant` in [`app/models/chat.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/chat.rb). This method builds an `Assistant` instance via `Assistant.for_chat`, which mixes in `Provided` to resolve the LLM provider through `Provider::Registry.for_concept(:llm)`.

The provider resolution in `Assistant::Provided#get_model_provider` supplies the concrete OpenAI client configuration. Subsequently, `Assistant#respond_to` creates an `AssistantMessage` record to store the pending response and instantiates `Assistant::Responder` with the message, system instructions, and a `FunctionToolCaller`.

### Step 5: Streaming LLM Responses

Inside `Assistant::Responder#respond` (defined in [`app/models/assistant/responder.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/assistant/responder.rb)), the system calls `llm.chat_response` with a streaming proc. This proc receives "output_text" and "response" chunks, emitting events that drive the real-time update mechanism.

The first "output_text" chunk creates the `AssistantMessage` record via `AssistantMessage#append_text!`, while subsequent chunks append to the existing record. The implementation also updates `Chat#latest_assistant_response_id` through `Chat#update_latest_response!` to maintain conversation state.

### Step 6: Parsing and Handling Function Calls

The OpenAI client returns raw JSON that `Provider::Openai::ChatParser` (in [`app/models/provider/openai/chat_parser.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/provider/openai/chat_parser.rb)) translates into `ChatResponse`, `ChatMessage`, and `ChatFunctionRequest` objects.

If the response contains `function_requests`, the responder invokes `FunctionToolCaller.fulfill_requests`. The resulting tool calls return as `function_results`, triggering a follow-up LLM request via `Responder#handle_follow_up_response` before finalizing the assistant's text output.

### Step 7: Broadcasting to the UI

The `Message` model in [`app/models/message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/message.rb) implements `after_create_commit` and `after_update_commit` callbacks that broadcast Turbo Stream updates to the `messages` target. This pushes real-time updates to the browser without polling, completing the data flow cycle.

Errors during streaming are rescued, logged on the `Chat` record via `Chat#add_error`, and broadcast to the UI to maintain user feedback during failures.

## Code Examples for the AI Chat Response System

These practical examples demonstrate how to interact with the AI chat response system programmatically.

### Start a New AI Chat

```ruby

# Create a chat with an initial user prompt

chat = Current.user.chats.start!("How much did I spend on groceries last month?", model: "gpt-4.1")

# The first assistant reply streams automatically via background job

```

### Send a Follow-Up Message

```ruby

# Simulate the POST /messages endpoint behavior

message = UserMessage.create!(
  chat: chat,
  content: "Show me the breakdown by category.",
  ai_model: "gpt-4.1"
)

# The after_create_commit callback queues AssistantResponseJob automatically

```

### Manually Trigger Background Processing

```ruby

# Useful for testing or synchronous scripts

AssistantResponseJob.perform_now(message)

```

### Inspect the Latest Assistant Reply

```ruby

# Access the persisted assistant message

assistant_msg = chat.assistant_message
puts assistant_msg.content

```

## Key Files in the Maybe AI Chat Architecture

Understanding these source files is critical for modifying the AI chat response system behavior:

- **[`app/models/chat.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/chat.rb)** - Orchestrates assistant calls via `ask_assistant`, manages conversation state, and updates `latest_assistant_response_id`
- **[`app/models/user_message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/user_message.rb)** - Persists user input and schedules async processing via `request_response_later`
- **[`app/models/message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/message.rb)** - Base class with Turbo Stream broadcast hooks and status enum
- **[`app/models/assistant.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/assistant.rb)** - Entry point that wraps provider configuration and builds the responder via `respond_to`
- **[`app/models/assistant/responder.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/assistant/responder.rb)** - Handles streaming LLM output and function call loops
- **[`app/models/assistant/provided.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/assistant/provided.rb)** - Resolves LLM provider via `get_model_provider` and registry lookup
- **[`app/models/provider/openai/chat_parser.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/provider/openai/chat_parser.rb)** - Translates OpenAI JSON into domain objects
- **[`app/models/provider/openai/chat_config.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/provider/openai/chat_config.rb)** - Configures OpenAI client parameters
- **[`app/jobs/assistant_response_job.rb`](https://github.com/maybe-finance/maybe/blob/main/app/jobs/assistant_response_job.rb)** - Active Job wrapper for asynchronous processing
- **[`app/controllers/chats_controller.rb`](https://github.com/maybe-finance/maybe/blob/main/app/controllers/chats_controller.rb)** - Routes initial chat creation
- **[`app/controllers/messages_controller.rb`](https://github.com/maybe-finance/maybe/blob/main/app/controllers/messages_controller.rb)** - Routes follow-up message creation

## Summary

- The AI chat response system uses a controller-to-job pipeline where `MessagesController#create` initiates async processing via `AssistantResponseJob`
- `UserMessage` and `Chat` models in [`app/models/chat.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/chat.rb) coordinate to persist state and delegate to the `Assistant` class
- The `Assistant::Responder` streams LLM output in real-time, parsing responses with `Provider::Openai::ChatParser`
- Function calls trigger a follow-up request loop via `FunctionToolCaller` before finalizing the assistant message
- Turbo Stream broadcasts in [`app/models/message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/message.rb) push updates to the browser without requiring page reloads

## Frequently Asked Questions

### How does the AI chat response system handle real-time updates?

The system uses Turbo Stream broadcasts defined in [`app/models/message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/message.rb) via `after_create_commit` and `after_update_commit` callbacks. When `AssistantMessage` content is appended during streaming, these callbacks push updates to the `messages` target in the browser, creating a real-time chat experience without polling.

### What triggers the background job for AI responses?

The `after_create_commit` callback in [`app/models/user_message.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/user_message.rb) named `request_response_later` automatically enqueues `AssistantResponseJob.perform_later(message)` immediately after the database commits the user message record. This ensures the LLM request happens asynchronously without blocking the HTTP response.

### How does Maybe support different LLM providers?

The `Assistant` class mixes in `Provided` (from [`app/models/assistant/provided.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/assistant/provided.rb)), which looks up the provider via `Provider::Registry.for_concept(:llm)`. This registry pattern allows the system to resolve the concrete OpenAI client or other LLM implementations dynamically based on configuration.

### Where is the streaming LLM response parsed?

The `Provider::Openai::ChatParser` class in [`app/models/provider/openai/chat_parser.rb`](https://github.com/maybe-finance/maybe/blob/main/app/models/provider/openai/chat_parser.rb) handles translation of raw OpenAI JSON into structured domain objects like `ChatResponse` and `ChatFunctionRequest`. The `Assistant::Responder` uses this parser to process streaming chunks and detect function call requests.