How to Implement Chain-of-Thought Reasoning with the Cosmos Reasoner Using the Format Instruction
Append the specific format instruction Answer the question using the following format: \n here the final answer. to your prompt and send it to a Cosmos 3 Reasoner endpoint using the Qwen-3-VL message format to elicit structured chain-of-thought reasoning.
The NVIDIA Cosmos 3 Reasoner enables explicit chain-of-thought (CoT) reasoning for video and image understanding tasks. By implementing chain-of-thought reasoning with the Reasoner using the format instruction, you can force the model to emit intermediate reasoning steps wrapped in `here the final answer.
When processed by the `Cosmos3ReasonerForConditionalGeneration` architecture, the model generates tokens autoregressively, starting with the ``, and then outputting the final answer after the newline.
## Message Format Requirements
The Reasoner expects messages in **Qwen-3-VL-compatible** format—a list of `role`/`content` dictionaries. This structure supports multimodal inputs including text, images, and videos.
According to the README (lines 869-885), each message object contains:
- `role`: Either `"system"`, `"user"`, or `"assistant"`
- `content`: A list of content items, each with a `type` field (`"text"`, `"image_url"`, or `"video_url"`)
Example structure:
```python
messages = [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "video_url", "video_url": "https://example.com/scene.mp4"},
{"type": "text", "text": "What happens next?\nAnswer the question using the following format:\n\n here the final answer."}
]
}
]
Implementation Methods
Method 1: OpenAI-Compatible API with vLLM
For production deployments, send requests to a vLLM server hosting the Reasoner endpoint. The endpoint loads only the Reasoner branch of the model, which processes the prompt and produces the `\n" " here the final answer." )
user_query = "What are the main objects in this video and what will happen next?" prompt = f"{user_query}\n{cot_instruction}"
messages = [ { "role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}] }, { "role": "user", "content": [ {"type": "video_url", "video_url": "https://example.com/scene.mp4"}, {"type": "text", "text": prompt} ] }, ]
3️⃣ Call the Reasoner endpoint
client = openai.OpenAI(base_url="http://localhost:8000/v1") resp = client.chat.completions.create( model="cosmos3-reasoner-nano", # any model registered with the Reasoner architecture
messages=messages,
temperature=0.6,
top_p=0.95,
)
print(resp.choices[0].message.content)
The response will contain text structured as:
```text
The next immediate action is for the robot to grasp the cube.
Method 2: Cosmos Framework CLI
For quick experiments without writing Python code, use the Cosmos Framework CLI entrypoint. Create a JSON input file following the same message structure, then invoke the Reasoner task.
# 1️⃣ Create a JSON input file
cat > reasoner_input.json <<EOF
{
"messages": [
{
"role": "system",
"content": [{"type": "text", "text": "You are a helpful assistant."}]
},
{
"role": "user",
"content": [
{"type": "image_url", "image_url": "https://example.com/photo.jpg"},
{"type": "text", "text": "Describe the scene using chain-of-thought.\nAnswer the question using the following format:\n\n here the final answer."}
]
}
]
}
EOF
# 2️⃣ Run the Reasoner through the framework entrypoint
cosmos_framework/scripts/inference \
--model-path models/cosmos3-nano \
--task reasoner \
--input-file reasoner_input.json \
--output-file answer.json
The generated answer.json contains the model's response with the \n to separate the reasoning chain from the final answer. The content before the closing tag contains the model's internal reasoning process, while the content after represents the distilled output.
Summary
- Format instruction: Append the exact string
Answer the question using the following format:\n\n here the final answer.to trigger CoT generation. - Message format: Use Qwen-3-VL-style role/content dictionaries with support for
video_urlandimage_urlcontent types. - Endpoints: Deploy via vLLM, NVIDIA NIM, or the Cosmos Framework CLI using the
Cosmos3ReasonerForConditionalGenerationarchitecture. - Response parsing: Look for ``.
Frequently Asked Questions
What is the exact format instruction string required?
You must use the precise wording: Answer the question using the following format:\n\n here the final answer. As documented in the README (lines 887-898), this template instructs the Cosmos 3 Reasoner to structure its output with explicit reasoning tags. Variations in spacing or tag formatting may cause the model to omit the chain-of-thought segment.
Can I use chain-of-thought reasoning with video inputs?
Yes. The Reasoner supports multimodal chain-of-thought reasoning across video, image, and text inputs. Include a video_url content item in the user message list alongside your text prompt containing the format instruction. The model will analyze the video frames using its vision encoder before generating the textual reasoning steps.
How does the Reasoner architecture differ from the standard Cosmos 3 model?
The Reasoner uses Cosmos3ReasonerForConditionalGeneration, which activates only the autoregressive transformer path with causal self-attention. Standard Cosmos 3 models may use diffusion-based generation, whereas the Reasoner generates text token-by-token through mRoPE-fused visual and textual embeddings, making it suitable for step-by-step reasoning tasks.
Why are my <think> tags not being generated?
Ensure you are sending requests to a Reasoner-specific endpoint (vLLM server loaded with --task reasoner or the Cosmos Framework inference script). Standard Cosmos 3 endpoints without the Reasoner architecture ignore the format instruction. Also verify that your message format matches the Qwen-3-VL specification with content as a list of dictionaries, not a plain string.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →