# How to Optimize the GeoIP Generation Process for Extremely Large Datasets

> Optimize GeoIP generation for millions of CIDR blocks using batch AddSet, parallel parsing, and IP version filtering. Boost performance and reduce overhead for massive datasets.

- Repository: [Loyalsoldier/geoip](https://github.com/loyalsoldier/geoip)
- Tags: performance
- Published: 2026-03-06

---

**Optimize GeoIP generation for massive datasets by leveraging batch `AddSet` operations, parallel input parsing, and selective IP version filtering to reduce memory pressure and CPU overhead when processing millions of CIDR blocks.**

The `loyalsoldier/geoip` repository provides a robust pipeline for converting and merging IP geolocation data, but default sequential processing struggles when handling datasets containing millions of CIDR entries. This guide examines the architectural bottlenecks in the source code and provides concrete optimization strategies based on the actual implementation in [`lib/instance.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/instance.go), [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go), and [`lib/entry.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/entry.go).

## Architectural Bottlenecks in Large-Scale Processing

The generation pipeline follows a linear data flow: **Input Converters** parse raw data and push CIDR blocks into a **Container**, which stores entries in `netipx.IPSetBuilder` instances, and finally **Output Converters** serialize the results. In [`lib/entry.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/entry.go), each `Entry` maintains separate `ipv4Builder` and `ipv6Builder` fields of type `*netipx.IPSetBuilder` that accumulate IP ranges until the final `IPSet` is materialized.

The critical performance hotspot is the **per-CIDR insertion overhead**. When processing extremely large datasets, calling `IPSetBuilder.Add` individually for millions of CIDR blocks creates severe memory pressure because each entry is retained in the builder's internal tree until `IPSet()` is called. This pattern, visible in the default `Container.Add` implementation in [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go), results in excessive CPU cycles spent on tree balancing and high heap utilization that triggers frequent garbage collection pauses.

## Performance Tuning Strategies

### Batch CIDR Processing with AddSet

Instead of inserting CIDR blocks one at a time, use the **`AddSet`** method to merge entire pre-built `IPSet` instances in a single operation. In [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go), when merging entries, calling `existing.ipv4Builder.AddSet(set4)` performs a batch union of radix trees rather than iterative insertion, dramatically reducing the computational complexity from O(N×logM) individual insertions to O(N+M) tree merging.

### Parallelize Input Parsing in Instance.RunInput

The default `RunInput` method in [`lib/instance.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/instance.go) processes input converters sequentially. For large datasets, refactor this to spawn **goroutines** for each `InputConverter`, having them write to a buffered channel of `*Entry` objects. After all parsers complete, merge the entries into the final container using `Container.Add` with the appropriate ignore options. This utilizes multi-core CPUs for I/O and parsing operations while maintaining thread safety through the channel-based collector.

### Early Deduplication and Streaming Merge

For datasets requiring only deduplication without format conversion, use the **`geoip merge`** command implemented in [`merge.go`](https://github.com/loyalsoldier/geoip/blob/main/merge.go). This command reads from `stdin` and writes to `stdout`, processing CIDRs in a streaming fashion rather than loading everything into memory. It internally utilizes `Container.Add` with `IgnoreIPv6()` or `IgnoreIPv4()` options to minimize builder overhead, emitting unique ranges as they are confirmed.

### Selective IP Version Handling

When sources contain only IPv4 or IPv6 data, explicitly pass **`IgnoreIPv4()`** or **`IgnoreIPv6()`** options to `Container.Add` as defined in [`lib/lib.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/lib.go). This prevents the instantiation of unnecessary `IPSetBuilder` instances for the unused IP version, cutting memory allocation by nearly half for single-stack datasets.

## Code Implementation Examples

### Concurrent Input Processing

Replace the sequential loop in [`lib/instance.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/instance.go) with this parallel implementation to saturate CPU cores during input parsing:

```go
func (i *instance) RunInput(container Container) error {
	var wg sync.WaitGroup
	entriesCh := make(chan *Entry, len(i.input))
	errCh := make(chan error, len(i.input))

	for _, ic := range i.input {
		wg.Add(1)
		go func(conv InputConverter) {
			defer wg.Done()
			tmp := NewContainer()
			if err := conv.Input(tmp); err != nil {
				errCh <- err
				return
			}
			for e := range tmp.Loop() {
				entriesCh <- e
			}
		}(ic)
	}

	wg.Wait()
	close(entriesCh)
	close(errCh)

	if len(errCh) > 0 {
		return <-errCh
	}

	for e := range entriesCh {
		if err := container.Add(e, IgnoreIPv4(), IgnoreIPv6()); err != nil {
			return err
		}
	}
	return nil
}

```

### Optimized Container Merging

Modify [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go) to leverage `AddSet` when merging entries with overlapping names:

```go
func (c *container) Add(entry *Entry, opts ...IgnoreIPOption) error {
	name := entry.GetName()
	if existing, found := c.GetEntry(name); found {
		if set4, _ := entry.ipv4Builder.IPSet(); set4 != nil {
			existing.ipv4Builder.AddSet(set4) // Batch union
		}
		if set6, _ := entry.ipv6Builder.IPSet(); set6 != nil {
			existing.ipv6Builder.AddSet(set6)
		}
		return nil
	}
	// Store new entry if not found
	c.entries[name] = entry
	return nil
}

```

### Streaming Deduplication Command

Process billion-line datasets with minimal memory footprint using the streaming merge command:

```bash

# Deduplicate a 10GB IPv4 CIDR list

cat massive-list.txt | ./geoip merge -t ipv4 > deduplicated.txt

```

## Summary

- **Batch operations** via `IPSetBuilder.AddSet` in [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go) eliminate per-CIDR insertion overhead when merging large datasets.
- **Parallel input parsing** using goroutines in [`lib/instance.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/instance.go) maximizes CPU utilization during I/O-bound input processing.
- **Streaming deduplication** through the [`merge.go`](https://github.com/loyalsoldier/geoip/blob/main/merge.go) command-line tool processes unlimited dataset sizes with constant memory usage.
- **IP version filtering** using `IgnoreIPv4()` and `IgnoreIPv6()` options reduces heap allocation by skipping unnecessary builder instantiation.

## Frequently Asked Questions

### What causes high memory usage when processing large GeoIP datasets?

High memory usage stems from the `netipx.IPSetBuilder` implementation in [`lib/entry.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/entry.go), which retains every CIDR block in memory until the final `IPSet` is materialized. When processing millions of entries, the builder's internal radix tree consumes significant heap space, triggering garbage collection pauses and potential out-of-memory errors in resource-constrained environments.

### How does the AddSet method improve performance over Add?

The `AddSet` method performs a **batch union** of two pre-built radix trees in O(N+M) time complexity, whereas iterative `Add` calls require O(N×logM) operations for individual tree insertions. When merging large entries in [`lib/container.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/container.go), using `existing.ipv4Builder.AddSet(set4)` reduces CPU overhead by eliminating repeated tree rebalancing and allocation for each CIDR prefix.

### Can I process multiple input sources in parallel?

Yes, by refactoring `RunInput` in [`lib/instance.go`](https://github.com/loyalsoldier/geoip/blob/main/lib/instance.go) to spawn goroutines for each `InputConverter` and collecting results through a channel, you can parse multiple large datasets concurrently. The container's `Add` method is not inherently thread-safe for simultaneous writes to the same entry, but parallel parsing into separate temporary containers followed by sequential merging into the final container achieves safe, scalable processing.

### What is the best approach for deduplicating billion-entry datasets?

For datasets exceeding available RAM, use the **`geoip merge`** command implemented in [`merge.go`](https://github.com/loyalsoldier/geoip/blob/main/merge.go), which streams data through `stdin` and `stdout` with constant memory usage. This approach avoids loading the entire dataset into the `Container` by checking uniqueness against the `IPSetBuilder` and immediately emitting results, making it suitable for billion-line CIDR lists on standard hardware.