# What Kind of Data Is Used to Train Orca? Clarifying the Orchestration Platform

> Discover what kind of data trains Orca, a powerful orchestration platform, not a language model. Understand its integration with LLM agents like Claude Code, Codex, and Gemini for unified output routing.

- Repository: [Stably/orca](https://github.com/stablyai/orca)
- Tags: faq
- Published: 2026-05-25

---

**Orca does not use any training data because it is not a language model; it is an orchestration platform that integrates existing LLM agents such as Claude Code, Codex, and Gemini, routing their outputs through a unified work-tree interface.**

The question "what kind of data is used to train Orca?" reflects a common misconception about the `stablyai/orca` repository. Orca is not an AI model trained on static datasets, code repositories, or human feedback. Instead, it is a runtime environment and CLI tool that coordinates multiple third-party large language models (LLMs) without containing any model-training pipeline internally. Understanding this architecture is critical for developers evaluating Orca for agent orchestration.

## Orca Is an Orchestrator, Not a Language Model

Orca contains **no neural network weights** and **no training corpus**. According to the repository's [`README.md`](https://github.com/stablyai/orca/blob/main/README.md), Orca functions as a work-tree management system that launches and coordinates external AI agents. The codebase provides terminal management, browser automation, and a unified CLI (`orca`), but all heavy inference is delegated to the underlying models hosted by their respective providers.

The repository explicitly lists supported agents including Claude Code, Codex, Grok, Gemini, Antigravity, Pi, Hermes, OpenCode, Goose, Amp, and others. These agents are pre-trained models owned by companies like Anthropic, OpenAI, and Google—not components of Orca itself.

## Where the "Training Data" Actually Lives

Since Orca routes requests to external LLMs, the training data question applies to those underlying providers rather than to Orca. The models accessible through Orca are typically trained on:

- Massive collections of public source code from GitHub and GitLab.
- Technical documentation and software manuals.
- Stack Overflow posts and developer Q&A forums.
- General-purpose text corpora (for foundational models).

These datasets are proprietary to each provider (Anthropic, OpenAI, Google, etc.) and are not present in the `stablyai/orca` repository. Orca’s role is limited to **integrating** these already-trained models via their respective APIs and CLIs.

## Technical Implementation of Agent Integration

Orca communicates with these external models through a skill-based architecture defined in [`skills/orca-cli/SKILL.md`](https://github.com/stablyai/orca/blob/main/skills/orca-cli/SKILL.md) and [`skills/orchestration/SKILL.md`](https://github.com/stablyai/orca/blob/main/skills/orchestration/SKILL.md). The CLI acts as a router, forwarding commands to the appropriate agent runtime while managing the local work-tree state.

### Launching Claude Code via Orca

```bash

# Create a new work-tree for the session

orca worktree create --repo my-repo --comment "Claude Code session"

# Launch the Claude agent in a managed terminal

orca terminal create --command "claude" --title "Claude Code"

```

### Running OpenAI Codex Through Orca

```bash

# Initialize workspace for Codex

orca worktree create --repo my-repo --comment "Codex session"

# Spawn Codex in the Orca-managed environment

orca terminal create --command "codex" --title "Codex"

```

### Coordinating Multiple Agents

The orchestration skill ([`skills/orchestration/SKILL.md`](https://github.com/stablyai/orca/blob/main/skills/orchestration/SKILL.md)) enables multi-agent workflows where different LLMs collaborate on the same codebase:

```bash

# Send a request to Claude

orca orchestration send \
  --from $ORCA_TERMINAL_HANDLE \
  --to claude \
  --message "Generate a skeleton for a React component."

# Forward results to Codex for refactoring

orca orchestration send \
  --from $ORCA_TERMINAL_HANDLE \
  --to codex \
  --message "Refactor the component to use TypeScript."

```

Each command interacts with the running Orca runtime, which delegates the actual language modeling to the external services. No model training occurs within these operations.

## Key Files That Document the Architecture

Several files in the repository clarify that Orca is purely an integration layer:

- **[`README.md`](https://github.com/stablyai/orca/blob/main/README.md)**: Lists supported AI agents and defines Orca as an orchestration platform rather than a model.
- **[`skills/orca-cli/SKILL.md`](https://github.com/stablyai/orca/blob/main/skills/orca-cli/SKILL.md)**: Specifies that the CLI "talks to a running Orca editor" and routes commands to agents rather than processing natural language internally.
- **[`skills/orchestration/SKILL.md`](https://github.com/stablyai/orca/blob/main/skills/orchestration/SKILL.md)**: Describes coordination of multiple agents, explicitly stating reliance on the "running Orca runtime" with no mention of model training.
- **[`AGENTS.md`](https://github.com/stablyai/orca/blob/main/AGENTS.md)**: Provides design-system guidelines for how agent outputs are rendered in the UI.
- **[`docs/STYLEGUIDE.md`](https://github.com/stablyai/orca/blob/main/docs/STYLEGUIDE.md)**: Governs UI conventions for displaying results from the various integrated LLMs.

## Summary

- **Orca is not an LLM**: It is an orchestration platform that manages work-trees and routes commands to external AI agents.
- **No training data exists in the repository**: The codebase contains integration logic only, with no datasets, weights, or training pipelines.
- **Training data belongs to providers**: The underlying models (Claude, Codex, Gemini, etc.) are trained by their respective companies on proprietary code and text corpora.
- **Integration via CLI and skills**: Orca uses a skill-based architecture (`skills/orca-cli/`, `skills/orchestration/`) to manage interactions with these third-party models.

## Frequently Asked Questions

### Does Orca have its own AI model?

No. Orca does not contain a language model. It is a runtime environment and CLI tool that launches and manages external LLM agents such as Claude Code and Codex. The repository contains only orchestration logic, not neural network weights or training infrastructure.

### What AI models can I use with Orca?

Orca supports a broad ecosystem of agents including Claude Code, Codex (OpenAI), Grok, Gemini (Google), Antigravity, Pi, Hermes, OpenCode, Goose, Amp, Auggie, Autohand Code, Charm, Cline, Codebuff, Continue, Cursor, Droid, GitHub Copilot, Kilocode, Kimi, Kiro, Mistral Vibe, Qwen Code, and Rovo Dev. Each agent must be installed separately and is subject to its own provider's terms and training data policies.

### Where is Orca's training data stored?

There is no training data stored in the Orca repository because Orca is not trained. The platform does not ingest codebases to learn patterns or fine-tune models. Any data processing occurs transiently during agent execution and is handled according to the policies of the underlying LLM providers (Anthropic, OpenAI, etc.).

### How does Orca process natural language commands?

Orca passes natural language commands to the underlying agent CLIs through its orchestration layer. When you run `orca terminal create --command "claude"`, Orca spawns the Claude Code CLI in a managed terminal. The language processing happens inside Anthropic's Claude model, not within Orca's codebase.