How the Open Notebook Skill Integrates with Self‑Hosted AI Providers: A LangChain & Esperanto Guide
The Open Notebook skill integrates with self‑hosted AI providers through a dual‑layer architecture using LangChain for workflow orchestration and Esperanto as a multi‑provider abstraction layer, enabling seamless connection to local models like Ollama, vLLM, or LM Studio via HTTP endpoints.
The Open Notebook skill within the K‑Dense‑AI/scientific‑agent‑skills repository provides a self‑hosted, open‑source alternative to Google NotebookLM that plugs into any locally‑run AI model. By leveraging LangChain with the Esperanto multi‑provider library, the system abstracts provider‑specific differences, allowing the same notebook creation, summarization, and podcast generation APIs to work regardless of where the model is hosted. According to the source code in SKILL.md (lines 55‑61) and the architecture diagram in architecture.md (line 153), Esperanto registers model metadata while LangChain consumes it uniformly, ensuring consistent functionality across local and cloud deployments.
Understanding the Dual‑Layer Architecture
The integration relies on two distinct abstraction layers that separate workflow logic from provider specifics:
LangChain provides uniform workflow handling for prompt templates, chaining, and output parsing. As shown in the architecture diagram at line 153 of architecture.md, LangChain operates at the "AI integration" node alongside Esperanto, remaining entirely unaware of the underlying provider implementation.
Esperanto functions as a multi‑provider wrapper that normalizes differences between 16+ AI backends including Ollama, vLLM, OpenAI, Anthropic, and Groq. As documented in SKILL.md (lines 55‑61), Esperanto detects the specific provider, translates the Open Notebook /credentials schema into provider‑specific API calls (such as Ollama’s /api/generate), and forwards requests accordingly.
This architecture means LangChain treats a local Llama 2 instance identically to a cloud GPT‑4 endpoint, while Esperanto handles the protocol translation beneath.
Configuring Self‑Hosted AI Providers
Integrating a local model server requires exposing the model via HTTP and registering it through Open Notebook’s credential API.
Exposing the Model via HTTP
Most self‑hosted stacks expose models through HTTP endpoints by default. Compatible servers include:
- Ollama (serving models on port 11434 by default)
- vLLM (providing an OpenAI‑compatible API)
- LM Studio (running a local inference server)
These services run as containers within the Docker Compose environment defined in the configuration files, operating alongside the Open Notebook server and SurrealDB.
Registering Credentials via the REST API
Open Notebook exposes REST endpoints to register self‑hosted providers dynamically. The process involves creating a credential, discovering available models, and registering specific model IDs for use.
import requests
BASE_URL = "http://localhost:5055/api"
# 1️⃣ Add a new credential (e.g., Ollama)
resp = requests.post(
f"{BASE_URL}/credentials",
json={"provider": "ollama", "name": "My Ollama", "api_key": ""} # Ollama uses no key
)
cred = resp.json()
# 2️⃣ Discover the models offered by the provider
discover = requests.post(f"{BASE_URL}/credentials/{cred['id']}/discover")
models = discover.json()["models"]
# 3️⃣ Register the models you want to use
requests.post(
f"{BASE_URL}/credentials/{cred['id']}/register-models",
json={"model_ids": [m["id"] for m in models]},
)
As shown in SKILL.md (lines 64‑88), once registered, you can call any Open Notebook endpoint (such as /notebooks/{id}/summarize) and LangChain routes the request to the Ollama model via Esperanto.
Practical Example: Podcast Generation with Local Models
Once configured, the integration enables complex workflows like podcast generation entirely within your self‑hosted environment. The following example assumes Ollama is running locally with the Llama 2 model:
# Assume Ollama is running locally with Llama 2
BASE_URL = "http://localhost:5055/api"
# Register Ollama credential (as above) and pick the model "llama2:13b"
payload = {
"notebook_id": "nb-123",
"model_id": "ollama-llama2-13b",
"speaker_count": 2,
"prompt": "Summarize the key findings and read them as a two‑speaker podcast."
}
resp = requests.post(f"{BASE_URL}/podcast", json=payload)
print(resp.json()["audio_url"])
The request flow follows this path: Open Notebook API → LangChain → Esperanto → Ollama. Esperanto handles the translation to Ollama’s native /api/generate endpoint, while LangChain manages the prompt templating. The resulting audio file URL returns without data ever leaving your Docker network.
Key Architectural Components
The self‑hosted integration relies on specific components working within the Docker Compose stack described in architecture.md (lines 5‑9):
| Component | Responsibility |
|---|---|
| Docker Compose | Orchestrates the Open Notebook server, SurrealDB, and optional AI containers (Ollama, vLLM, etc.). |
| SurrealDB | Stores notebooks, source documents, chat histories, and model metadata persistently. |
| AI Service | Any HTTP‑exposed model server running self‑hosted. |
| Esperanto | Detects provider type, normalizes request payloads, and forwards to the specific model API. |
| LangChain | Handles prompt templates, chaining logic, and output parsing uniformly. |
Benefits of Self‑Hosted Provider Integration
Unified API: Esperanto translates Open Notebook’s credential schema into provider‑specific implementations. Whether connecting to Groq, Mistral, or a local Ollama instance, the integration code remains identical.
No API Key Requirement: For providers running locally (Ollama, LM Studio), the api_key field can remain empty during credential creation, simplifying security management.
Dynamic Model Discovery: The /credentials/{id}/discover endpoint queries the provider at runtime to retrieve model names, capabilities, and token limits, automatically exposing them to the UI without manual configuration.
Full Privacy: All notebook content, source files, and generated outputs remain inside your Docker network. When using self‑hosted providers, AI inference occurs locally without transmitting data to external APIs.
Summary
- Open Notebook integrates self‑hosted AI through LangChain and Esperanto, creating a provider‑agnostic abstraction layer documented in
architecture.mdandSKILL.md. - Registration occurs via the
/credentialsREST API, supporting dynamic model discovery through the/discoverendpoint and accepting empty API keys for local servers. - The architecture stores all data in SurrealDB and processes AI requests through containerized services, ensuring complete data privacy with no external telemetry.
- Once registered, endpoints like
/notebooks/{id}/summarizeand/podcastwork identically across cloud and self‑hosted models, with Esperanto handling provider‑specific translations.
Frequently Asked Questions
What is Esperanto in the Open Notebook architecture?
Esperanto is a multi‑provider wrapper library that abstracts differences between 16+ AI backends. It detects the provider type (such as Ollama or vLLM), normalizes the request payload, and translates Open Notebook API calls into provider‑specific endpoints like Ollama’s /api/generate. This allows LangChain to treat all models uniformly without provider‑specific code.
Can I use Ollama without an API key in Open Notebook?
Yes. When registering a self‑hosted provider like Ollama that runs locally within your Docker network, you can pass an empty string for the api_key parameter in the credential creation request to /credentials. Open Notebook recognizes that local providers do not require authentication keys, simplifying the setup process while maintaining secure local access.
How does LangChain handle different self‑hosted models?
LangChain remains provider‑agnostic by consuming models through Esperanto’s normalized interface. Whether processing a request for a local Llama 2 instance or a cloud GPT‑4 endpoint, LangChain handles prompt templates, chaining, and output parsing identically. Esperanto manages the translation to provider‑specific APIs underneath, as implemented in the AI integration node shown in architecture.md.
Is my data kept private when using self‑hosted AI providers?
Yes. All notebook content, source documents, chat histories, and generated outputs store in SurrealDB within your Docker network. When using self‑hosted providers like Ollama or vLLM, AI inference occurs locally without transmitting data to external APIs, ensuring complete privacy and no external telemetry according to the architecture described in the K‑Dense‑AI repository.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →