How ML Intern Tracks Token Usage Across the Agent Loop
ML Intern tracks token usage by instrumenting LLM call sites with a telemetry helper that extracts usage fields from Litellm responses, emits structured llm_call events to the session store, and aggregates them via post-processing scripts for cost analysis.
The huggingface/ml-intern repository implements a granular telemetry system to monitor LLM consumption in real time. By intercepting model responses at the agent loop level, the system captures prompt tokens, completion tokens, and cache-related metrics, enabling precise cost attribution and performance analysis across autonomous agent executions.
Normalizing Token Usage from Litellm Responses
Token tracking begins with normalization. In agent/core/telemetry.py, the extract_usage function (lines 29-69) sanitizes the raw usage field from Litellm responses or plain dictionaries. It returns a stable dictionary containing prompt_tokens, completion_tokens, total_tokens, cache_read_tokens, and cache_creation_tokens, ensuring consistent field names regardless of the underlying model provider.
Emitting Telemetry from the Agent Loop
The record_llm_call async function in agent/core/telemetry.py serves as the central instrumentation hook. It is invoked from two specific locations in agent/core/agent_loop.py: lines 390-401 for streaming responses and lines 474-482 for non-streaming responses. This dual instrumentation ensures that every LLM invocation, whether delivered via chunked streaming or complete response objects, triggers telemetry capture.
When called, record_llm_call executes the following steps:
- Invokes
extract_usageon the raw response to normalize token counts. - Calculates USD cost using
litellm.completion_cost, defaulting to0.0if the calculation fails. - Constructs an
llm_callevent containing the model name, latency, finish reason, cost, and all token fields. - Sends the event to the current
Sessionviasession.send_event.
# agent/core/telemetry.py
async def record_llm_call(
session: Any,
*,
model: str,
response: Any = None,
latency_ms: int,
finish_reason: str | None,
) -> dict:
"""Emit an ``llm_call`` event and return the extracted usage dict."""
usage = extract_usage(response) if response is not None else {}
cost_usd = 0.0
if response is not None:
try:
from litellm import completion_cost
cost_usd = float(completion_cost(completion_response=response) or 0.0)
except Exception:
cost_usd = 0.0
from agent.core.session import Event
try:
await session.send_event(Event(
event_type="llm_call",
data={
"model": model,
"latency_ms": latency_ms,
"finish_reason": finish_reason,
"cost_usd": cost_usd,
**usage,
},
))
except Exception as e:
logger.debug("record_llm_call failed (non-fatal): %s", e)
return usage
Persisting Events to Session Storage
Once emitted, events are captured by Session.send_event in agent/core/session.py. The method appends each llm_call event to an in-memory logged_events list and forwards it to connected clients. When the session terminates, Session.save_trajectory_local persists the entire event stream—including all token usage records—to a JSON trajectory file, creating an immutable audit trail suitable for downstream analysis.
Aggregating Tokens for KPI Reporting
Post-execution analysis is handled by scripts/build_kpis.py. This script reads persisted session logs and iterates over every llm_call event (lines 42-48), extracting token fields from event["data"] to compute aggregate statistics. The aggregation loop sums prompt_tokens, completion_tokens, cache_read_tokens, and cache_creation_tokens, producing per-session and per-bucket KPIs for cost control and performance monitoring.
# scripts/build_kpis.py (inside the per‑event loop)
if et == "llm_call":
out["llm_calls"] += 1
out["tokens_prompt"] += int(data.get("prompt_tokens") or 0)
out["tokens_completion"] += int(data.get("completion_tokens") or 0)
out["tokens_cache_read"] += int(data.get("cache_read_tokens") or 0)
out["tokens_cache_creation"] += int(data.get("cache_creation_tokens") or 0)
out["cost_usd"] += float(data.get("cost_usd") or 0.0)
Accessing Token Totals from Saved Sessions
You can programmatically audit token consumption from any saved session file:
import json, pathlib
session_file = pathlib.Path("session_logs/session_abc123_20240424_101530.json")
with session_file.open() as f:
trajectory = json.load(f)
# Sum tokens across all LLM calls in this session
total_prompt = sum(ev["data"].get("prompt_tokens", 0)
for ev in trajectory["events"]
if ev["event_type"] == "llm_call")
print(f"Prompt tokens used: {total_prompt}")
Summary
- Normalization: The
extract_usagehelper inagent/core/telemetry.pystandardizes Litellm responses into a consistent schema that includes cache-related token fields. - Instrumentation: Call sites in
agent/core/agent_loop.py(streaming at lines 390-401 and non-streaming at lines 474-482) invokerecord_llm_callto capture usage at the moment of response receipt. - Cost Calculation: Each
llm_callevent includes USD cost computed vialitellm.completion_cost, with a safe fallback to0.0on failure. - Persistence: Events are stored in-memory via
Session.send_eventand flushed to JSON viaSession.save_trajectory_localinagent/core/session.py. - Aggregation: The
scripts/build_kpis.pyutility sums token fields across sessions, enabling precise cost attribution and performance reporting.
Frequently Asked Questions
How does ML Intern handle token tracking for streaming responses versus non-streaming responses?
Both paths use identical instrumentation. The agent/core/agent_loop.py file contains separate call sites for streaming (lines 390-401) and non-streaming (lines 474-482), but both invoke the same telemetry.record_llm_call method with the final response chunk or complete response object. This ensures token extraction and event emission remain consistent regardless of how the model delivers content.
What specific token fields does ML Intern track beyond standard prompt and completion counts?
According to the implementation in agent/core/telemetry.py, the system captures five distinct fields: prompt_tokens, completion_tokens, total_tokens, cache_read_tokens, and cache_creation_tokens. The latter two are particularly important for providers like Anthropic that support prompt caching, allowing the system to distinguish between novel and cached prompt content.
Where is token usage physically stored in ML Intern?
Token usage is stored as structured events within the logged_events list of a Session object defined in agent/core/session.py. When the session ends, save_trajectory_local persists these events to a JSON file on disk, typically located in a session_logs/ directory, creating a permanent record that scripts/build_kpis.py can later ingest for analysis.
How does the telemetry system handle cost calculation failures?
The record_llm_call function wraps the litellm.completion_cost invocation in a try-except block. If the cost calculation raises an exception or returns None, the function defaults the cost_usd field to 0.0 and continues emitting the event. This defensive design ensures that token tracking remains robust and non-blocking even when cost metadata is unavailable or the Litellm integration encounters errors.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →