# Differences Between GPU Acceleration Libraries in the Optimize for GPU Skill

> Explore GPU acceleration libraries in K-Dense-AI's Optimize for GPU skill. Learn about RAPIDS, CuPy, cuDF, Warp, and KvikIO for efficient CUDA array data sharing and zero-copy performance.

- Repository: [K-Dense/scientific-agent-skills](https://github.com/K-Dense-AI/scientific-agent-skills)
- Tags: deep-dive
- Published: 2026-05-14

---

**The "Optimize for GPU" skill in the K‑Dense‑AI/scientific‑agent‑skills repository documents 12 NVIDIA RAPIDS libraries that expose a common CUDA Array Interface, enabling zero‑copy data sharing between drop‑in replacements like CuPy and cuDF and specialized engines like Warp and KvikIO.**

The **K‑Dense‑AI/scientific‑agent‑skills** repository provides a comprehensive decision framework for GPU‑accelerated scientific computing. According to [`docs/scientific-skills.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/docs/scientific-skills.md), every library in the ecosystem implements the **CUDA Array Interface**, allowing arrays to be passed between CuPy, cuDF, cuML, and other components without host‑to‑device copies. This architectural consistency—confirmed by the runtime detection logic in [`scientific-skills/get-available-resources/scripts/detect_resources.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/get-available-resources/scripts/detect_resources.py)—enables you to build heterogeneous pipelines that remain entirely on the GPU from data ingestion to visualization.

## Drop‑In Replacement Libraries

These libraries prioritize API compatibility with popular CPU libraries, minimizing code changes during migration.

### CuPy: NumPy and SciPy Acceleration

**CuPy** serves as a GPU‑accelerated replacement for **NumPy** and **SciPy**, handling dense linear algebra, FFT, random number generation, and statistics. It is the optimal choice when you have existing NumPy code that requires a minimal‑change speed boost. Limitations include sparse matrix support (which requires the separate CuPy‑sparse module) and incomplete coverage of SciPy’s special functions.

```python
import cupy as cp

x = cp.arange(10)
y = cp.sin(x)
result = cp.linalg.norm(y)

```

### cuDF: GPU‑Accelerated DataFrames

**cuDF** replaces **pandas** for tabular ETL, groupby operations, and joins on GPU. It excels at large‑scale aggregations that exceed CPU memory but does not yet support certain pandas extensions such as complex custom `apply` functions.

```python
import cudf

pdf = cudf.read_csv('data.csv')
result = pdf.groupby('category').agg({'value':'mean'})

```

### cuML: Scikit‑Learn on GPU

**cuML** accelerates classical machine learning algorithms from **scikit‑learn**, including logistic regression, K‑means, and dimensionality reduction. It delivers 10×–100× speedups for supported algorithms, though newer scikit‑learn models may not yet be implemented.

```python
from cuml.linear_model import LogisticRegression

model = LogisticRegression()
model.fit(X_gpu, y_gpu)

```

### cuCIM: Image Processing

**cuCIM** provides GPU‑accelerated image filtering, morphology, and whole‑slide reading as a replacement for **scikit‑image**. It is optimized for digital pathology and microscopy workflows requiring fast convolution operations.

```python
import cucim

img = cucim.io.imread('slide.svs')
filtered = cucim.filters.gaussian(img, sigma=2)

```

### cuSpatial: Geospatial Analytics

**cuSpatial** accelerates spatial joins and geometry operations compatible with **GeoPandas**. It targets polygon‑heavy analytics that become bottlenecked on CPU, though feature parity with GeoPandas is still expanding.

```python
import cuspatial

joined = cuspatial.sjoin(df1, df2, how='inner')

```

### cuGraph: Graph Analytics

**cuGraph** replaces **NetworkX** for large‑scale graph analytics such as PageRank, centrality calculations, and shortest paths. It is ideal for graphs where CPU implementations are too slow or memory‑bound, but it does not support all NetworkX custom graph classes.

```python
import cugraph

g = cugraph.from_cudf_edgelist(edges, source='src', destination='dst')
pr = cugraph.pagerank(g)

```

## Specialized Compute and Custom Kernels

When drop‑in compatibility is insufficient, these libraries provide fine‑grained control or domain‑specific abstractions.

### Numba CUDA: JIT‑Compiled Kernels

**Numba CUDA** allows you to write custom GPU kernels in pure Python using just‑in‑time compilation. Use it when you need explicit control over thread hierarchies, shared memory, or rapid kernel prototyping. The trade‑off is verbose boilerplate and manual performance tuning for block sizes.

```python
from numba import cuda

@cuda.jit
def vec_add(a, b, out):
    i = cuda.grid(1)
    if i < a.size:
        out[i] = a[i] + b[i]

```

### Warp: Physics Simulation

**Warp** focuses on JIT‑compiled simulation kernels with built‑in spatial types (vectors, meshes, rays). It targets physics‑oriented workloads like particle simulations, robotics, and differentiable rendering. While it offers powerful geometry primitives, community examples are fewer than those for CuPy or Numba.

```python
import warp as wp

@wp.kernel
def gravity(p):
    wp.atomic_add(p.force, wp.vec3(0, -9.81, 0))

```

### RAFT: Low‑Level Primitives

**RAFT** (and **pylibraft**) provides foundational algorithms such as sparse eigensolvers, device memory management, and multi‑GPU communication primitives. It is intended for developers building custom solvers or multi‑GPU collectives who require NCCL and UCX integration without high‑level abstractions.

```python
from raft.linalg import eig

w, v = eig(A_gpu)

```

## Utility and Infrastructure Libraries

### KvikIO: GPUDirect Storage

**KvikIO** enables high‑throughput file I/O that streams directly into GPU memory, supporting S3, HTTP, and Zarr backends. It eliminates CPU‑mediated transfers in I/O‑bound pipelines, such as reading terabytes of Parquet files, but requires hardware and driver support for GPUDirect.

```python
import kvikio

buf = kvikio.read('s3://bucket/file.parquet')

# buf lives in GPU memory, ready for cuDF processing

```

### cuVS: Vector Similarity Search

**cuVS** implements GPU‑accelerated approximate nearest‑neighbor search algorithms (CAGRA, IVF‑Flat, IVF‑PQ). It is specialized for recommendation systems and RAG retrieval workflows involving large‑scale embedding queries.

```python
from cucim import cuvs

index = cuvs.ivf_flat.build(x_gpu)
neighbors = index.search(query_gpu, k=10)

```

### cuxfilter: Interactive Visualization

**cuxfilter** builds GPU‑accelerated dashboards using Bokeh, Datashader, and Deck.gl for smooth brushing and zooming across millions of rows. It is designed for rapid exploratory data analysis rather than heavy computation or model training.

```python
import cuxfilter

cuxfilter.plot(df, x='x', y='y')

```

## Zero‑Copy Interoperability and Pipeline Composition

All libraries adhere to the **CUDA Array Interface**, allowing a CuPy array to be passed directly to cuML or a cuDF DataFrame to be consumed by cuGraph without memory duplication. According to the architectural documentation in [`docs/scientific-skills.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/docs/scientific-skills.md), this enables heterogeneous pipelines that remain entirely on the device:

```python
import cudf, cuml, cugraph, cupy as cp

# 1️⃣ Load CSV directly into GPU memory

df = cudf.read_csv('large_table.csv')

# 2️⃣ Train a logistic regression model on GPU

X = df[['feat1','feat2']].values   # CuPy array

y = df['label'].values
model = cuml.linear_model.LogisticRegression()
model.fit(X, y)

# 3️⃣ Build a graph from output probabilities

probs = model.predict_proba(X)[:,1]
edges = cp.column_stack([df['src'], df['dst']])
g = cugraph.from_cudf_edgelist(cudf.DataFrame({'src':edges[:,0], 'dst':edges[:,1]}))
pr = cugraph.pagerank(g)

# 4️⃣ Visualise top‑ranked nodes with cuxfilter

import cuxfilter
cuxfilter.plot(df, x='feat1', y='feat2', color=pr)

```

Multi‑GPU scaling is supported internally by RAFT and cuVS when `CUDA_VISIBLE_DEVICES` lists multiple GPUs, with higher‑level libraries like cuML and cuGraph automatically utilizing these primitives.

## How to Select a GPU Acceleration Library

- **NumPy/SciPy compatibility**: Choose **CuPy** for dense linear algebra.
- **pandas DataFrame operations**: Choose **cuDF** for ETL and aggregations.
- **Scikit‑learn algorithms**: Choose **cuML** for classical ML model training.
- **Custom kernel logic**: Choose **Numba CUDA** for fine‑grained thread control.
- **Physics or robotics simulation**: Choose **Warp** for built‑in spatial primitives.
- **Massive graph analytics**: Choose **cuGraph** for PageRank and centrality.
- **Direct‑to‑GPU I/O**: Choose **KvikIO** for GPUDirect Storage reads.
- **Image processing**: Choose **cuCIM** for digital pathology workflows.
- **Vector search**: Choose **cuVS** for RAG and recommendation retrieval.
- **Geospatial joins**: Choose **cuSpatial** for polygon/point operations.
- **Low‑level solvers**: Choose **RAFT** for multi‑GPU primitives and eigenvalue decomposition.

## Summary

- **K‑Dense‑AI/scientific‑agent‑skills** documents 12 GPU acceleration libraries in its [`docs/scientific-skills.md`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/docs/scientific-skills.md) file, ranging from drop‑in NumPy replacements to specialized physics engines.
- All libraries expose the **CUDA Array Interface**, enabling zero‑copy data sharing and end‑to‑end GPU pipelines.
- **CuPy**, **cuDF**, **cuML**, and **cuCIM** offer immediate migration paths for existing CPU code by mirroring familiar APIs.
- **Numba CUDA**, **Warp**, and **RAFT** provide lower‑level control for custom kernels and simulation workloads.
- **KvikIO** and **cuxfilter** address infrastructure needs for I/O and visualization without CPU round‑trips.
- Runtime availability is verified via [`scientific-skills/get-available-resources/scripts/detect_resources.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/get-available-resources/scripts/detect_resources.py), which checks for CUDA, ROCm, or Metal backends.

## Frequently Asked Questions

### What is the CUDA Array Interface and why is it important for GPU acceleration libraries?

The **CUDA Array Interface** is a standard protocol that allows GPU libraries to share device memory buffers without copying data through the CPU. As implemented in the K‑Dense‑AI/scientific‑agent‑skills framework, it means a CuPy array can be passed directly to cuML or cuGraph, eliminating redundant memory transfers and enabling pipelines like `cuDF → cuML → cuGraph` to execute entirely on the GPU.

### How do I detect which GPU libraries are available in my environment?

The repository includes [`scientific-skills/get-available-resources/scripts/detect_resources.py`](https://github.com/K-Dense-AI/scientific-agent-skills/blob/main/scientific-skills/get-available-resources/scripts/detect_resources.py), which performs runtime detection of CUDA, ROCm, and Metal backends. This script checks which RAPIDS libraries (CuPy, cuDF, cuML, etc.) are importable and reports the available GPU compute capability, allowing your scientific skills to degrade gracefully or select appropriate backends dynamically.

### When should I choose Numba CUDA over CuPy?

Choose **Numba CUDA** when you need fine‑grained control over thread hierarchies, shared memory, or custom algorithms that do not map to NumPy’s vectorized operations. Choose **CuPy** when you want a drop‑in NumPy replacement with minimal code changes. Numba requires manual tuning of block sizes and kernel launch configurations, making it more verbose than CuPy’s high‑level array operations.

### Can I combine multiple GPU acceleration libraries in a single workflow?

Yes. Because all libraries support the CUDA Array Interface, you can compose heterogeneous pipelines. For example, you can read data with **KvikIO** into a CuPy array, wrap it as a **cuDF** DataFrame, train a model with **cuML**, run graph analytics with **cuGraph**, and visualize results with **cuxfilter**, all without transferring data back to host memory.