How the Run Method in lib/instance.go Orchestrates the Input-to-Output Conversion Pipeline
The Run method validates configuration, initializes a shared Container, sequentially executes all registered input converters to populate the container, then streams the aggregated data through all output converters to final destinations.
The Run method serves as the central conductor of the loyalsoldier/geoip library, transforming raw IP data from diverse sources into standardized output formats. Located in lib/instance.go, this method implements a strict pipeline architecture that ensures data flows deterministically from inputs to outputs through a shared in-memory store. Understanding this orchestration is essential for extending the library with custom converters or debugging complex transformation workflows.
Pipeline Architecture Overview
The Run method implements a linear three-phase pipeline that guarantees data integrity across transformation stages:
- Validation Phase: Verifies that at least one
InputConverterand oneOutputConverterhave been registered via configuration or programmatic registration - Ingestion Phase: Executes all input converters sequentially to populate a fresh
Containerinstance with parsed IP entries - Emission Phase: Streams the populated container through all output converters to generate final artifacts (JSON, MMDB, plaintext, etc.)
This architecture ensures that all input data aggregates in a single shared state before any output generation begins, preventing race conditions and enabling cross-referencing between disparate input sources.
Phase 1: Configuration Validation and Container Initialization
Upon invocation, Run immediately validates the instance state to prevent null-pointer exceptions during processing. According to the source code in lib/instance.go at lines 111-115, the method checks that both input and output slices contain registered converters:
if len(i.input) == 0 || len(i.output) == 0 {
return errors.New("missing input or output")
}
Once validation passes, Run initializes the shared state by calling NewContainer() at lines 116-117. This creates a fresh Container struct (defined in lib/container.go) containing a map-based storage layer (container.entries) that holds *Entry objects indexed by canonical names. The container serves as the single source of truth throughout the pipeline execution, passed by reference to all subsequent processing stages.
Phase 2: Input Processing via RunInput
After container initialization, Run delegates to the unexported helper method RunInput (lines 89-98 in lib/instance.go). This method iterates over the i.input slice and invokes each converter's Input(container Container) (Container, error) interface method:
for _, ic := range i.input {
var err error
container, err = ic.Input(container)
if err != nil {
return err
}
}
Each input converter—whether parsing CSV files, MaxMind DBs, or custom formats—reads its source data and populates the container using methods like container.Add(entry, ...), container.Lookup(name), or container.GetEntry(name). The container accumulates contributions from all inputs, with later inputs potentially enriching or overriding entries added by earlier ones based on the converter implementation.
Phase 3: Output Processing via RunOutput
Once all inputs have populated the container, Run invokes RunOutput (lines 101-108 in lib/instance.go) to execute the emission phase. This method iterates over the i.output slice, calling each converter's Output(container Container) error method:
for _, oc := range i.output {
if err := oc.Output(container); err != nil {
return err
}
}
Output converters read from the finalized container using iteration methods like container.Loop(callback) or direct lookups, then serialize the data to their respective targets (files, stdout, network sockets). Because outputs execute sequentially, they operate on a complete, immutable view of the aggregated dataset, ensuring consistency across all generated artifacts. Upon successful completion of all outputs, Run returns nil at lines 124-127, signaling successful pipeline completion.
Complete Implementation Example
The following example demonstrates the full pipeline orchestration, from configuration loading through execution:
package main
import (
"log"
"github.com/loyalsoldier/geoip/lib"
)
func main() {
// Create a fresh instance
inst, err := lib.NewInstance()
if err != nil {
log.Fatalf("instance creation failed: %v", err)
}
// Load configuration that registers input (CSV) and output (JSON) converters
if err := inst.InitConfig("examples/config.json"); err != nil {
log.Fatalf("config load failed: %v", err)
}
// Execute the complete pipeline: CSV → Container → JSON
if err := inst.Run(); err != nil {
log.Fatalf("run failed: %v", err)
}
}
For programmatic configuration without JSON files, converters can be added directly before invoking the orchestration:
inst.AddInput(lib.NewCSVInput("input.csv")) // Implements InputConverter
inst.AddOutput(lib.NewJSONOutput("out.json")) // Implements OutputConverter
if err := inst.Run(); err != nil { // Orchestrates validation → input → output
log.Fatal(err)
}
Summary
- Validation First: The
Runmethod inlib/instance.goenforces preconditions at lines 111-115, ensuring the pipeline has both input sources and output destinations before allocating resources. - Shared Container State: A single
Containerinstance created viaNewContainer()acts as the intermediate data structure, passed sequentially through all inputs and outputs to maintain data consistency. - Sequential Execution: Both input processing (
RunInput, lines 89-98) and output processing (RunOutput, lines 101-108) execute converters in registration order, enabling deterministic data layering and transformation. - Interface-Driven Design: The pipeline relies on the
InputConverterandOutputConverterinterfaces defined inlib/converter.go, allowing custom implementations without modifying core orchestration logic.
Frequently Asked Questions
What happens if no inputs or outputs are registered when calling Run?
The Run method returns an immediate error at lines 111-115 in lib/instance.go with the message "missing input or output". This validation prevents the initialization of empty containers or unnecessary processing cycles when the pipeline configuration is incomplete.
How does data flow between input and output converters?
Input converters populate a shared Container struct (defined in lib/container.go) using methods like Add() and GetEntry(). After all inputs complete, the same container instance passes to output converters, which read the aggregated data using Loop() or Lookup() methods. The container serves as the sole communication channel between phases.
Can multiple input converters process data in a single Run execution?
Yes. The RunInput method iterates over the entire i.input slice sequentially, allowing you to chain multiple input sources (e.g., CSV + MMDB + JSON) in a single pipeline. Each converter receives the container state left by the previous one, enabling incremental data enrichment and merging.
Where are the InputConverter and OutputConverter interfaces defined?
Both interfaces are defined in lib/converter.go. The InputConverter requires an Input(Container) (Container, error) method, while OutputConverter requires Output(Container) error. These interfaces enable the Run method to treat diverse data sources uniformly without knowing specific implementation details.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →