Where to Find the Main Entry Point for the Marin Application
The main entry point for the Marin application is the marin-serve console script, which resolves to the main() function defined in lib/marin/src/marin/inference/iris_cli.py.
The Marin platform provides a unified interface for serving large language models on TPU or GPU slices. Understanding the main entry point for the Marin application is essential for developers who need to debug, extend, or programmatically interact with the serving infrastructure. This entry point handles both vLLM and Levanter backends through a single command-line interface.
Locating the Entry Point in the Source Code
The Marin application follows standard Python packaging conventions, with its executable logic clearly separated from its declaration.
The iris_cli.py Module
The primary implementation resides in lib/marin/src/marin/inference/iris_cli.py. This module contains the main() function that serves as the canonical entry point when users execute the marin-serve command:
# lib/marin/src/marin/inference/iris_cli.py
def main() -> None:
... # parses CLI arguments, builds a ServingPlan, and launches the Iris service
When invoked, this function orchestrates the entire serving pipeline, from argument parsing to job submission.
Console Script Declaration in pyproject.toml
The marin-serve command is declared in the project's pyproject.toml under the [project.scripts] section. According to the marin-community/marin source code, this configuration maps the console script directly to the entry point function:
[project.scripts]
marin-serve = "marin.inference.iris_cli:main"
This declaration ensures that when you run marin-serve after installing the package, Python executes the main() function from the marin.inference.iris_cli module.
How the marin-serve Entry Point Works
The main() function in iris_cli.py implements a four-stage execution flow:
- Parses CLI arguments via Click, validating backend-specific options for both vLLM and Levanter.
- Validates hardware configurations, ensuring compatibility between selected slices (TPU v6e-8 or GPU H100x8) and the chosen backend.
- Constructs a
ServingPlanthat selects the appropriate inference engine and computes required environment variables. - Submits an Iris job that starts the serving backend on the selected slice and registers an OpenAI-compatible endpoint.
Practical Usage Examples
The marin-serve entry point supports multiple hardware and backend combinations through a unified interface.
Running on TPU with vLLM
To serve a model on a TPU slice using the default vLLM backend:
marin-serve iris Qwen/Qwen3-0.6B \
--cluster my-cluster \
--tpu v6e-8 \
--backend vllm
Running on GPU with Levanter
To use the Levanter backend on a GPU slice instead:
marin-serve iris Qwen/Qwen3-0.6B \
--cluster my-cluster \
--gpu H100x8 \
--backend levanter
Both commands are processed by the same main() entry point, which dynamically adjusts the serving plan based on the hardware and backend flags.
Supporting Entry Point Files
While iris_cli.py serves as the primary entry point, the Marin application includes specialized entry points for specific verification tasks:
lib/marin/src/marin/inference/vllm_wheel_entrypoint.py: Helper entry point used when verifying the vLLM wheel before launching the actual vLLM CLI.lib/marin/src/marin/inference/vllm_smoke_test.py: Provides a lightweight smoke-test entry point for validating vLLM inference capabilities without a full model load.
These files supplement the main entry point by handling edge cases in the vLLM deployment pipeline.
Summary
- The main entry point for the Marin application is the
main()function inlib/marin/src/marin/inference/iris_cli.py. - The
marin-serveconsole script is declared inpyproject.tomland maps tomarin.inference.iris_cli:main. - This entry point supports both vLLM and Levanter backends across TPU and GPU hardware slices.
- The entry point constructs a
ServingPlanand submits an Iris job to register OpenAI-compatible endpoints. - Supplementary entry points exist for vLLM wheel verification and smoke testing.
Frequently Asked Questions
What file contains the main() function for marin-serve?
The main() function for the marin-serve command is located in lib/marin/src/marin/inference/iris_cli.py. This function handles CLI parsing and orchestrates the model serving pipeline for both vLLM and Levanter backends.
How is the marin-serve command registered in the Marin package?
The command is registered in pyproject.toml under the [project.scripts] section, where marin-serve is mapped to marin.inference.iris_cli:main. Python packaging tools use this entry point declaration to generate the executable script during package installation.
Can I use the Marin entry point to serve models on both TPU and GPU?
Yes. The marin-serve entry point accepts either --tpu (e.g., v6e-8) or --gpu (e.g., H100x8) flags, along with --backend options (vllm or levanter). The main() function in iris_cli.py validates these combinations and constructs an appropriate ServingPlan for the selected hardware.
What happens when I execute the marin-serve command?
When executed, the entry point parses command-line arguments using Click, validates backend-specific configurations, builds a ServingPlan specifying the inference engine and environment, and submits an Iris job. This job starts the serving backend on the designated slice and registers an OpenAI-compatible API endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →