Optimizing Token Usage and Reducing LLM Operational Costs with MetaGPT
MetaGPT reduces LLM operational costs through built-in token counting, configurable message compression, and granular cost tracking across multiple providers.
MetaGPT is a multi-agent framework for orchestrating Large Language Model workflows. By implementing token counters, message compression strategies, and cost management subsystems, the framework helps developers minimize API spend while preserving the quality of complex agent interactions.
Core Architecture for Token Optimization
MetaGPT's cost-control system operates through four integrated components that work transparently during every LLM call.
Token Counter
The count_message_tokens and count_output_tokens functions in metagpt/utils/token_counter.py calculate exact token consumption for any message list or string. This provides ground-truth measurements before sending requests to the API, enabling data-driven decisions about when to prune content.
Message Compression
Implemented in metagpt/provider/base_llm.py, the compress_messages method supports four strategies defined in metagpt/configs/compress_msg_config.py:
POST_CUT_BY_MSG– Retains the newest messages up to the token thresholdPOST_CUT_BY_TOKEN– Truncates newest messages to exact token budgetsPRE_CUT_BY_MSG– Retains the oldest messages, discarding recent onesPRE_CUT_BY_TOKEN– Truncates from the beginning of the conversation
By trimming or truncating conversations to fit within TOKEN_MAX limits (defined per model in metagpt/utils/token_counter.py), the system prevents costly context window overflows and API retries.
Cost Management
The CostManager class in metagpt/utils/cost_manager.py tracks prompt and completion tokens using per-model pricing tables for OpenAI, Anthropic, Gemini, and Fireworks. For self-hosted models, TokenCostManager provides zero-cost tracking that monitors usage without monetary calculation.
Context Integration
The Context class in metagpt/context.py wires LLM instances to their appropriate cost managers through _select_costmanager, ensuring every API call is automatically accounted for. The create_llm_instance factory in metagpt/provider/llm_provider_registry.py injects these dependencies consistently across all provider implementations.
Practical Implementation
Counting Tokens Before API Calls
Use count_message_tokens to preview costs before sending requests:
from metagpt.utils.token_counter import count_message_tokens
msgs = [
{"role": "user", "content": "Explain quantum entanglement in simple terms."},
{"role": "assistant", "content": "Quantum entanglement is ..."}
]
num = count_message_tokens(msgs, model="gpt-4-0314")
print(f"Tokens required: {num}") # → Tokens required: 15
Source: metagpt/utils/token_counter.py
Compressing Long Conversations
Apply compress_messages to enforce context window limits dynamically:
from metagpt.provider.base_llm import BaseLLM
from metagpt.configs.compress_msg_config import CompressType
# Assume `llm` is a concrete subclass (e.g., OpenAIChatLLM)
messages = [...]
compressed = llm.compress_messages(
messages,
compress_type=CompressType.POST_CUT_BY_TOKEN,
max_token=128000,
threshold=0.8,
)
print(f"Compressed to {len(compressed)} messages")
This implementation retains system messages, then applies the selected cutting strategy to user/assistant messages to maintain keep_token = max_token * threshold.
Source: metagpt/provider/base_llm.py
Tracking Real-Time Costs
Enable automatic cost tracking through the Context object:
from metagpt.context import Context
ctx = Context()
llm = ctx.llm() # creates LLM and attaches appropriate CostManager
response = await llm.achat("Write a brief summary of the latest AI research.")
# After the call, inspect accumulated cost:
print(f"Prompt tokens: {ctx.cost_manager.get_total_prompt_tokens()}")
print(f"Completion tokens: {ctx.cost_manager.get_total_completion_tokens()}")
print(f"Total cost (USD): {ctx.cost_manager.get_total_cost():.4f}")
Source: metagpt/context.py and metagpt/utils/cost_manager.py
Configuring Budget Controls for Self-Hosted Models
Switch to non-monetary tracking for local LLMs:
from metagpt.context import Context
from metagpt.utils.cost_manager import TokenCostManager
ctx = Context()
ctx.cost_manager = TokenCostManager() # Override default manager
llm = ctx.llm()
# Tracks token counts only; no monetary cost recorded
Source: metagpt/utils/cost_manager.py
Configuration and Integration
YAML-Based Compression Settings
Configure compression globally via configuration files. Create or edit your LLM config (e.g., config/examples/openai-gpt-4-turbo.yaml):
llm:
model: gpt-4-turbo
compress_type: post_cut_by_token # or pre_cut_by_msg, etc.
token_threshold: 0.8 # Keep 80% of context window
When Config.from_yaml loads this file, the LLM instance automatically applies the specified compression strategy before each API call.
Model-Specific Token Limits
The TOKEN_MAX dictionary in metagpt/utils/token_counter.py defines default context windows for each supported model (e.g., 128,000 tokens for GPT-4-turbo). These limits inform the compression algorithms, ensuring automatic boundary enforcement without manual calculation.
Summary
MetaGPT provides a complete toolchain for optimizing token usage and reducing LLM operational costs:
- Accurate token accounting via
count_message_tokenseliminates guesswork about payload sizes - Four compression strategies in
compress_messagesautomatically fit conversations within context windows - Flexible cost policies support commercial APIs (OpenAI, Anthropic) and self-hosted models through swappable cost managers
- Transparent integration through the
Contextclass ensures all LLM calls are tracked without boilerplate code
Frequently Asked Questions
How does MetaGPT count tokens for different models?
MetaGPT uses model-specific encodings in metagpt/utils/token_counter.py. The count_message_tokens function accepts a model parameter (e.g., "gpt-4-0314") and applies the appropriate tokenizer to calculate exact consumption, including special tokens for message formatting.
What compression strategies does MetaGPT support?
The framework supports four strategies defined in CompressType: POST_CUT_BY_MSG and POST_CUT_BY_TOKEN retain newest messages (truncating oldest), while PRE_CUT_BY_MSG and PRE_CUT_BY_TOKEN retain oldest messages (truncating newest). The *_BY_TOKEN variants truncate individual messages to exact byte boundaries, while *_BY_MSG operates on whole message boundaries.
Can I use MetaGPT's cost tracking with self-hosted LLMs?
Yes. Instantiate TokenCostManager instead of the default CostManager. This class tracks prompt and completion token counts without applying monetary pricing, making it ideal for local models or custom endpoints where per-token costs are zero or unknown.
Where are token limits configured in MetaGPT?
Token limits are defined in the TOKEN_MAX dictionary within metagpt/utils/token_counter.py. These defaults can be overridden via YAML configuration files or passed directly to compress_messages as the max_token parameter. The system automatically references these limits when calculating compression thresholds.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →