Optimizing Token Usage and Reducing LLM Operational Costs with MetaGPT

MetaGPT reduces LLM operational costs through built-in token counting, configurable message compression, and granular cost tracking across multiple providers.

MetaGPT is a multi-agent framework for orchestrating Large Language Model workflows. By implementing token counters, message compression strategies, and cost management subsystems, the framework helps developers minimize API spend while preserving the quality of complex agent interactions.

Core Architecture for Token Optimization

MetaGPT's cost-control system operates through four integrated components that work transparently during every LLM call.

Token Counter

The count_message_tokens and count_output_tokens functions in metagpt/utils/token_counter.py calculate exact token consumption for any message list or string. This provides ground-truth measurements before sending requests to the API, enabling data-driven decisions about when to prune content.

Message Compression

Implemented in metagpt/provider/base_llm.py, the compress_messages method supports four strategies defined in metagpt/configs/compress_msg_config.py:

  • POST_CUT_BY_MSG – Retains the newest messages up to the token threshold
  • POST_CUT_BY_TOKEN – Truncates newest messages to exact token budgets
  • PRE_CUT_BY_MSG – Retains the oldest messages, discarding recent ones
  • PRE_CUT_BY_TOKEN – Truncates from the beginning of the conversation

By trimming or truncating conversations to fit within TOKEN_MAX limits (defined per model in metagpt/utils/token_counter.py), the system prevents costly context window overflows and API retries.

Cost Management

The CostManager class in metagpt/utils/cost_manager.py tracks prompt and completion tokens using per-model pricing tables for OpenAI, Anthropic, Gemini, and Fireworks. For self-hosted models, TokenCostManager provides zero-cost tracking that monitors usage without monetary calculation.

Context Integration

The Context class in metagpt/context.py wires LLM instances to their appropriate cost managers through _select_costmanager, ensuring every API call is automatically accounted for. The create_llm_instance factory in metagpt/provider/llm_provider_registry.py injects these dependencies consistently across all provider implementations.

Practical Implementation

Counting Tokens Before API Calls

Use count_message_tokens to preview costs before sending requests:

from metagpt.utils.token_counter import count_message_tokens

msgs = [
    {"role": "user", "content": "Explain quantum entanglement in simple terms."},
    {"role": "assistant", "content": "Quantum entanglement is ..."}
]

num = count_message_tokens(msgs, model="gpt-4-0314")
print(f"Tokens required: {num}")      # → Tokens required: 15

Source: metagpt/utils/token_counter.py

Compressing Long Conversations

Apply compress_messages to enforce context window limits dynamically:

from metagpt.provider.base_llm import BaseLLM
from metagpt.configs.compress_msg_config import CompressType

# Assume `llm` is a concrete subclass (e.g., OpenAIChatLLM)

messages = [...]
compressed = llm.compress_messages(
    messages,
    compress_type=CompressType.POST_CUT_BY_TOKEN,
    max_token=128000,
    threshold=0.8,
)
print(f"Compressed to {len(compressed)} messages")

This implementation retains system messages, then applies the selected cutting strategy to user/assistant messages to maintain keep_token = max_token * threshold.

Source: metagpt/provider/base_llm.py

Tracking Real-Time Costs

Enable automatic cost tracking through the Context object:

from metagpt.context import Context

ctx = Context()
llm = ctx.llm()                     # creates LLM and attaches appropriate CostManager

response = await llm.achat("Write a brief summary of the latest AI research.") 

# After the call, inspect accumulated cost:

print(f"Prompt tokens: {ctx.cost_manager.get_total_prompt_tokens()}")
print(f"Completion tokens: {ctx.cost_manager.get_total_completion_tokens()}")
print(f"Total cost (USD): {ctx.cost_manager.get_total_cost():.4f}")

Source: metagpt/context.py and metagpt/utils/cost_manager.py

Configuring Budget Controls for Self-Hosted Models

Switch to non-monetary tracking for local LLMs:

from metagpt.context import Context
from metagpt.utils.cost_manager import TokenCostManager

ctx = Context()
ctx.cost_manager = TokenCostManager()  # Override default manager

llm = ctx.llm()

# Tracks token counts only; no monetary cost recorded

Source: metagpt/utils/cost_manager.py

Configuration and Integration

YAML-Based Compression Settings

Configure compression globally via configuration files. Create or edit your LLM config (e.g., config/examples/openai-gpt-4-turbo.yaml):

llm:
  model: gpt-4-turbo
  compress_type: post_cut_by_token   # or pre_cut_by_msg, etc.

  token_threshold: 0.8                # Keep 80% of context window

When Config.from_yaml loads this file, the LLM instance automatically applies the specified compression strategy before each API call.

Model-Specific Token Limits

The TOKEN_MAX dictionary in metagpt/utils/token_counter.py defines default context windows for each supported model (e.g., 128,000 tokens for GPT-4-turbo). These limits inform the compression algorithms, ensuring automatic boundary enforcement without manual calculation.

Summary

MetaGPT provides a complete toolchain for optimizing token usage and reducing LLM operational costs:

  • Accurate token accounting via count_message_tokens eliminates guesswork about payload sizes
  • Four compression strategies in compress_messages automatically fit conversations within context windows
  • Flexible cost policies support commercial APIs (OpenAI, Anthropic) and self-hosted models through swappable cost managers
  • Transparent integration through the Context class ensures all LLM calls are tracked without boilerplate code

Frequently Asked Questions

How does MetaGPT count tokens for different models?

MetaGPT uses model-specific encodings in metagpt/utils/token_counter.py. The count_message_tokens function accepts a model parameter (e.g., "gpt-4-0314") and applies the appropriate tokenizer to calculate exact consumption, including special tokens for message formatting.

What compression strategies does MetaGPT support?

The framework supports four strategies defined in CompressType: POST_CUT_BY_MSG and POST_CUT_BY_TOKEN retain newest messages (truncating oldest), while PRE_CUT_BY_MSG and PRE_CUT_BY_TOKEN retain oldest messages (truncating newest). The *_BY_TOKEN variants truncate individual messages to exact byte boundaries, while *_BY_MSG operates on whole message boundaries.

Can I use MetaGPT's cost tracking with self-hosted LLMs?

Yes. Instantiate TokenCostManager instead of the default CostManager. This class tracks prompt and completion token counts without applying monetary pricing, making it ideal for local models or custom endpoints where per-token costs are zero or unknown.

Where are token limits configured in MetaGPT?

Token limits are defined in the TOKEN_MAX dictionary within metagpt/utils/token_counter.py. These defaults can be overridden via YAML configuration files or passed directly to compress_messages as the max_token parameter. The system automatically references these limits when calculating compression thresholds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →