marin
Open-source framework for the research and development of foundation models.
Learn to perform text deduplication with Marin-dupekit. Discover three distributed modes Exact Paragraph Exact Document and Fuzzy Document for petabyte-scale corpora using the Zephyr execution engine.
What Is the Marin‑Ducky Service? A DuckDB‑Powered SQL Dashboard for Cloud StorageExplore the Marin-Ducky service a DuckDB-powered SQL dashboard for querying data in Google Cloud Storage and other object stores. Run ad-hoc SQL queries easily.
How to Deploy Models with Marin-Deploy: A Complete Guide to ML Model ServingDeploy ML models effortlessly with Marin-Deploy. This guide shows how to use its command-line interface for seamless deployment to Kubernetes and cloud platforms.
What Is the Purpose of Marin-core? A Deep Dive into the Foundation of the Marin EcosystemDiscover the purpose of Marin-core, the foundational data layer for the Marin ecosystem. It standardizes data for question-answer examples and conversation logs, ensuring smooth data flow across pipelines.
How Marin Ensures Training Mixture Reproducibility: A Two‑Phase PipelineMarin ensures training mixture reproducibility with a two phase pipeline. Discover how it guarantees identical logical mixtures for consistent results.
Where to Find Delphi Checkpoints on Hugging Face: A Complete GuideEasily find Delphi checkpoints on Hugging Face. Discover specific model identifiers in the Marin repository for your NLP projects.
Delphi: Marin’s Open Scaling Suite for Systematic LLM TrainingDiscover Delphi, Marin's open scaling suite for systematic LLM training. Automate reproducible pipelines from 3×10¹⁸ to 10²³ FLOPs with a learned scaling law.
How Marin Handles Text Deduplication: Exact, Paragraph, and Fuzzy Matching ExplainedDiscover how Marin handles text deduplication with exact paragraph, exact document, and fuzzy matching. Learn about its powerful Zephyr engine and MinHash algorithms.
How Marin Curates Raw Source DataLearn how Marin curates raw source data through a deterministic download transform normalize pipeline. Ensure reproducibility and traceability with Datakit and StepSpec.
Haliax for Named Array Programming in JAX: A Complete GuideDiscover Haliax for named array programming in JAX. Replace positional dimensions with named axes to eliminate shape mismatch bugs and write self-documenting code.
How Marin-Fray Provides a Distributed Execution SubstrateMarin-Fray provides a distributed execution substrate by unifying Iris job orchestration. Learn how it translates specifications, supports actor hosting, coscheduling, and dynamic discovery.
Marin-Zephyr Pipeline Stages: A Complete Guide to the 8 Core Stage TypesExplore the 8 core Marin-Zephyr pipeline stages: MAP, FILTER, FLATMAP, REDUCE, GROUP_BY, SORTED_MERGE_JOIN, WRITE, and CUSTOM. Understand how ZephyrCoordinator orchestrates these steps.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →