How Folia's AI Theme Generation Uses Google Gemini and OpenAI

Folia generates synchronized light and dark UI themes by sending song lyrics to either Google Gemini or OpenAI, then parses the structured JSON response through a shared schema validator.

Folia-major is an open-source music player that dynamically themes its interface based on the emotional content of the currently playing track. The project implements a provider-agnostic AI theme generation system that can switch between Google Gemini and OpenAI (or compatible APIs) using a single environment variable.

Provider Selection Architecture

The system determines which AI provider to use at runtime by checking the VITE_AI_PROVIDER environment variable. In src/services/gemini.ts, the generateThemeFromLyrics function selects the appropriate backend endpoint:

const provider = import.meta.env.VITE_AI_PROVIDER;
const endpoint = provider === 'openai'
  ? '/api/generate-theme_openai'
  : '/api/generate-theme';

When set to "openai", requests route to the OpenAI-compatible handler. Any other value (commonly "gemini") triggers the Google Gemini code path.

OpenAI Implementation

The OpenAI integration lives in worker/generate-theme_openai.ts. This handler constructs a strict schema-constrained request to ensure valid JSON output.

Key implementation details:

  • Model: Defaults to gpt-4o but supports configuration via environment variables
  • Endpoint: OpenAI Chat Completions API (/v1/chat/completions)
  • Response format: Uses response_format: { type: "json_schema", ... } to enforce structured output matching THEME_JSON_SCHEMA
  • Prompt construction:
    • buildThemeSystemPrompt() creates the system instructions describing the JSON schema and artistic requirements
    • buildThemeSourcePrompt() packages the lyrics snippet, isPureMusic boolean flag, and optional song title

The handler normalizes the API URL, detects the provider type, and sends the payload with the schema constraint. Upon receiving the response, it strips markdown fences and passes the result through sanitizeDualTheme.

Google Gemini Implementation

The Gemini path is implemented in worker/generate-theme.ts (exposed via the /api/generate-theme endpoint). This mirrors the OpenAI logic but adapts to Google's API structure.

Key differences from OpenAI:

  • Endpoint: https://generativelanguage.googleapis.com/v1beta/models/gemini-pro:generateContent
  • Payload structure: Uses contents array with role and parts instead of Chat Completions format
  • Temperature: Fixed at 0.7 for creative consistency
  • Response extraction: Parses candidates[0].content.parts[0].text from the Gemini response object

The backend extracts the JSON content, cleans markdown fences, and passes it through the same sanitizeDualTheme function used by the OpenAI path.

Shared Schema and Sanitization

Both providers target an identical JSON structure defined in THEME_JSON_SCHEMA. The schema requires a dual-theme object containing separate configurations for light and dark modes:

{
  light: {
    name: string,
    description: string,
    backgroundColor: string,
    primaryColor: string,
    accentColor: string,
    secondaryColor: string,
    wordColors: [{ word: string, color: string }],
    lyricsIcons: string[]
  },
  dark: { /* same shape */ }
}

After parsing the AI response, both code paths call sanitizeDualTheme from ../shared/themeSanitizer.mjs to validate and normalize the data. Finally, applyStoredAnimationIntensityToDualTheme merges user-specific animation preferences before returning the result to the UI.

Frontend Integration

The frontend service in src/services/gemini.ts orchestrates the request flow. The generateThemeFromLyrics function accepts lyrics text and optional metadata, then handles the network request:

export const generateThemeFromLyrics = async (
  lyricsText: string,
  options?: { isPureMusic?: boolean; songTitle?: string }
): Promise<DualTheme> => {
  const provider = import.meta.env.VITE_AI_PROVIDER;
  const endpoint = provider === 'openai'
    ? '/api/generate-theme_openai'
    : '/api/generate-theme';

  const response = await fetch(endpoint, {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ lyricsText, ...options })
  });
  // ... response handling
};

For pure instrumental tracks, the system sends the song title instead of lyrics, allowing the AI to generate themes based on textual metadata alone.

Summary

  • Provider switching happens via the VITE_AI_PROVIDER environment variable, requiring no code changes to swap between Google Gemini and OpenAI
  • OpenAI path uses schema-constrained chat completions with gpt-4o in worker/generate-theme_openai.ts
  • Gemini path calls the Google Generative Language API in worker/generate-theme.ts with a temperature of 0.7
  • Shared validation ensures both providers return compatible DualTheme objects through sanitizeDualTheme
  • Frontend abstraction in src/services/gemini.ts unifies the interface regardless of backend provider

Frequently Asked Questions

How does Folia handle instrumental tracks without lyrics?

When the isPureMusic flag is set to true, Folia sends the song title instead of lyrics to the AI provider. The prompt in buildThemeSourcePrompt includes this flag, allowing the model to generate appropriate color schemes based solely on the track name and any available metadata.

Can I use a custom OpenAI-compatible API instead of the official OpenAI service?

Yes. The worker/generate-theme_openai.ts handler normalizes the API URL and auto-detects the provider type, allowing integration with compatible services such as DeepSeek or self-hosted models. Simply set the appropriate API endpoint and key in your environment configuration.

What happens if the AI returns malformed JSON?

Both backend handlers strip markdown code fences from the response before parsing. The sanitizeDualTheme function then validates the structure against THEME_JSON_SCHEMA, ensuring that missing fields or invalid color values do not crash the UI. Invalid themes fall back to default values or trigger an error boundary in the frontend.

Why does the Gemini implementation use "gemini-pro" specifically?

The endpoint https://generativelanguage.googleapis.com/v1beta/models/gemini-pro:generateContent targets Google's Gemini Pro model, which offers a balance of creative generation speed and cost efficiency suitable for real-time theme generation. The architecture allows for easy model upgrades by changing the model identifier in the URL while maintaining the same request/response parsing logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →