How to Use FreeLLMAPI for Anthropic Claude Messages API: Complete Integration Guide
FreeLLMAPI implements a fully Anthropic-compatible Messages endpoint at POST /v1/messages that accepts standard Claude API payloads, translates them internally, and routes through free-model providers.
FreeLLMAPI's Anthropic Claude Messages API integration lets you point any Claude client—whether the official Anthropic SDK, Claude Code, or custom scripts—at your self-hosted instance. The service transparently converts Anthropic-shaped requests into its internal OpenAI-compatible format, executes its free-model fallback engine, and returns Anthropic-compatible responses including Server-Sent Events (SSE) for streaming.
Architecture Overview
The Anthropic compatibility layer lives in server/src/routes/anthropic.ts. Here's how a request flows through the system:
| Step | Action | Source Location |
|---|---|---|
| 1 | HTTP POST /v1/messages enters via anthropicRouter |
anthropic.ts:45 |
| 2 | authenticate helper validates x-api-key against unified API key |
anthropic.ts:95-103 |
| 3 | Zod messagesSchema validates payload structure |
anthropic.ts:81-98 |
| 4 | convertRequest translates Anthropic blocks to internal ChatMessage[] |
anthropic.ts:65-66 |
| 5 | applyTokenBudget enforces token limits with optional compressRequest |
anthropic.ts:106-108 |
| 6 | resolveAnthropicModel maps Claude family names to concrete models |
anthropic.ts:30 |
| 7 | Sticky session handling via x-claude-code-session-id header |
anthropic.ts:36-38 |
| 8 | runFallbackLoop executes the shared fallback engine |
anthropic.ts:110-118 |
| 9 | toAnthropicContent converts provider response back to Anthropic shape |
anthropic.ts:94-114 |
| 10 | streamCompletion handles SSE streaming with Anthropic event types |
anthropic.ts:70-73 |
Authentication and API Keys
FreeLLMAPI uses a unified API key system. The same key works for both OpenAI-compatible and Anthropic-compatible routes.
Set the x-api-key header to your key from the dashboard:
curl -X POST https://your-freellmapi-host/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-d '{ "model": "auto", "messages": [...] }'
The authenticate function in server/src/routes/anthropic.ts validates this against getUnifiedApiKey from the database.
Core API Requests
Basic Chat Completion
Send a simple non-streaming request:
curl -X POST https://api.myfreellmapi.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-d '{
"model": "auto",
"max_tokens": 512,
"messages": [
{ "role": "user", "content": "Explain quantum entanglement in plain English." }
]
}'
Key parameters:
model: Any Claude family name (claude-3-opus,claude-3.5-sonnet,claude-3-haiku) or"auto"for automatic free-model selectionmax_tokens: Required output token limit (defaults to 1024 if omitted)messages: Array of{role, content}objects following Anthropic's format
Streaming Responses
Enable SSE streaming with stream: true:
curl -N -X POST https://api.myfreellmapi.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-d '{
"model": "claude-3.5-sonnet",
"stream": true,
"messages": [
{ "role": "user", "content": "Tell me a joke." }
]
}'
Use curl -N (no buffering) to receive events in real-time. The server emits standard Anthropic SSE events:
event: message_start
event: content_block_start
event: content_block_delta
event: content_block_stop
event: message_delta
event: message_stop
The streamCompletion function handles translation of OpenAI provider streams into this Anthropic event format.
Tool Use and Function Calling
FreeLLMAPI forwards tool definitions to capable providers and repairs returned tool calls:
curl -X POST https://api.myfreellmapi.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-d '{
"model": "auto",
"tools": [
{
"name": "search",
"description": "Web search",
"input_schema": {
"type": "object",
"properties": { "query": { "type": "string" } },
"required": ["query"]
}
}
],
"tool_choice": { "type": "auto" },
"messages": [
{ "role": "user", "content": "Find the tallest mountain in the world." }
]
}'
Tool arguments are validated via server/src/lib/tool-validate.ts and repaired via server/src/lib/tool-args.ts.
Multimodal Inputs (Images)
Send images using Anthropic's block format:
curl -X POST https://api.myfreellmapi.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-d '{
"model": "auto",
"messages": [
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": "<BASE64_DATA>"
}
},
{ "type": "text", "text": "Describe this picture." }
]
}
]
}'
Images are normalized via server/src/lib/image-normalize.ts, converted to image_url blocks, and token-cost estimated before budget checks.
Model Resolution and Mapping
The resolveAnthropicModel function in server/src/services/anthropic-map.ts handles model selection:
| Input | Behavior |
|---|---|
claude-3-opus |
Maps to highest-capability free provider |
claude-3.5-sonnet |
Maps to Sonnet-class free model |
claude-3-haiku |
Maps to fastest/lightest free model |
auto (default) |
FreeLLMAPI selects cheapest available provider |
You can customize mappings in the dashboard or directly edit server/src/services/anthropic-map.ts.
Session Persistence and Sticky Routing
For clients like Claude Code that require consistent model behavior across turns, supply a session ID:
curl -X POST https://api.myfreellmapi.com/v1/messages \
-H "Content-Type: application/json" \
-H "x-api-key: $FREE_LLMAPI_KEY" \
-H "x-claude-code-session-id: my-session-123" \
-d '{ "model": "auto", "messages": [...] }'
The router extracts this header (also accepts x-session-id) and pins the session to the same provider/model, preventing model flapping. See anthropic.ts:36-38.
Token Budgeting and Compression
FreeLLMAPI enforces per-request token limits through server/src/lib/guardrails.ts:
estimateTokenscalculates prompt token countapplyTokenBudgetchecks against limits- If over budget,
compressRequestmay truncate or summarize - Compression status returned in
X-FreeLLM-Compressheader
Set explicit budgets via max_tokens or rely on DEFAULT_MAX_TOKENS (1024).
Error Handling
Provider errors map to Anthropic-compatible error types:
| Error Type | Trigger |
|---|---|
api_error |
Generic provider failure |
overloaded_error |
Rate limit or quota exceeded |
request_too_large |
Token budget exceeded |
not_found_error |
Invalid model reference |
The anthropicErrorType helper at anthropic.ts:165-170 performs this mapping.
Key Source Files
| File | Purpose |
|---|---|
server/src/routes/anthropic.ts |
Main Anthropic Messages endpoint implementation |
server/src/services/anthropic-map.ts |
Claude family → catalog model resolution |
server/src/lib/fallback-loop.ts |
Shared provider fallback and retry engine |
server/src/lib/content.ts |
Content block utilities |
server/src/lib/image-normalize.ts |
Image preprocessing and token estimation |
server/src/lib/guardrails.ts |
Token budget enforcement |
server/src/lib/tool-args.ts |
Tool argument repair |
server/src/lib/tool-validate.ts |
Tool schema validation |
server/src/lib/error-redaction.ts |
Error message sanitization |
Summary
- FreeLLMAPI's Anthropic Claude Messages API provides drop-in compatibility at
POST /v1/messages - Unified API key authentication works across OpenAI and Anthropic routes
- Automatic model routing via
"auto"or explicit Claude family names - Full feature parity including streaming, tools, images, and session persistence
- Token budgeting and compression protect against oversized requests
- Sticky sessions ensure consistent behavior for multi-turn conversations
Frequently Asked Questions
What model names can I use with the FreeLLMAPI Anthropic Claude Messages API?
Use any Claude family identifier (claude-3-opus, claude-3-sonnet, claude-3.5-sonnet, claude-3-haiku) or "auto" to let FreeLLMAPI select the best available free provider. The mapping logic in server/src/services/anthropic-map.ts resolves these to concrete catalog models at runtime.
Do I need a separate Anthropic API key to use this integration?
No. FreeLLMAPI uses a unified API key system. The same key from your dashboard works for both OpenAI-compatible (/v1/chat/completions) and Anthropic-compatible (/v1/messages) endpoints. The service handles all provider authentication internally through its free-model pool.
How does streaming work with Claude Code or other Anthropic clients?
Set stream: true in your request. FreeLLMAPI's streamCompletion function translates OpenAI provider SSE streams into Anthropic-native events (message_start, content_block_delta, message_stop). Use the x-claude-code-session-id header to pin sessions and prevent model switching mid-conversation.
Can I use tools and image inputs together?
Yes. The convertRequest function in server/src/routes/anthropic.ts handles mixed content blocks including text, images, and tool definitions. Images are normalized via server/src/lib/image-normalize.ts, tools are validated via server/src/lib/tool-validate.ts, and the fallback engine routes to providers capable of handling both modalities.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →