Building Python Libraries in Rust with PyO3 and Maturin

Python is everywhere in scientific computing. Rust is fast and safe. What if you could have both? That’s the promise of PyO3 and Maturin: write performance-critical code in Rust, expose it to Python, and package everything as a normal pip-installable library.

When building ParGA, this combination let me keep the ergonomic Python API scientists expect while running genetic algorithms at near-native speed. Here’s how it works.

The Stack

Three tools make Rust-backed Python libraries possible:

PyO3 provides Rust bindings for the Python interpreter. It handles type conversions, memory management, and the messy details of Python’s C API. You annotate Rust structs and functions with #[pyclass] and #[pyfunction], and PyO3 generates the glue code.

Maturin builds and packages PyO3 projects. It compiles your Rust code, generates Python wheels, and can publish directly to PyPI. One command: maturin build --release.

NumPy integration through the numpy crate for PyO3 enables zero-copy array sharing between Rust and Python. For numerical work, this is essential.

Project Structure

A typical PyO3 project has both Rust and Python components:

parga/
├── Cargo.toml           # Rust dependencies
├── pyproject.toml       # Python packaging (uses Maturin)
├── src/
│   ├── lib.rs           # Rust library root
│   ├── genome.rs        # Core Rust types
│   └── python/
│       └── mod.rs       # PyO3 bindings
└── python/
    └── parga/
        ├── __init__.py  # Python package
        └── ga.py        # High-level Python API

The key insight is that you can have two layers: low-level Rust types exposed directly via PyO3, and high-level Python wrappers that provide a more Pythonic interface.

Cargo Configuration

In Cargo.toml, you need a few specific settings:

[lib]
crate-type = ["cdylib", "rlib"]

[dependencies]
pyo3 = { version = "0.24", features = ["extension-module"] }
numpy = "0.24"

[features]
default = ["parallel"]
parallel = ["rayon"]
python = ["pyo3", "numpy"]

The cdylib crate type produces a dynamic library Python can load. The rlib lets you also use the crate as a normal Rust library.

Exposing Types to Python

Here’s a simplified example from ParGA. In Rust, we define a genome type:

use pyo3::prelude::*;
use numpy::{PyArray1, IntoPyArray};

#[pyclass]
#[derive(Clone)]
pub struct RealGenome {
    genes: Vec<f64>,
    bounds: (f64, f64),
}

#[pymethods]
impl RealGenome {
    #[new]
    fn new(length: usize, bounds: (f64, f64)) -> Self {
        let genes = vec![0.0; length];
        RealGenome { genes, bounds }
    }

    fn genes<'py>(&self, py: Python<'py>) -> Bound<'py, PyArray1<f64>> {
        self.genes.clone().into_pyarray_bound(py)
    }
}

The #[pyclass] attribute makes RealGenome visible to Python. The #[pymethods] block exposes methods. The #[new] attribute creates the Python __init__ constructor.

The Module Definition

In your lib.rs, you define the Python module:

#[pymodule]
fn _parga(m: &Bound<'_, PyModule>) -> PyResult<()> {
    m.add_class::<RealGenome>()?;
    m.add_class::<GAResult>()?;
    m.add_function(wrap_pyfunction!(run_ga, m)?)?;
    Ok(())
}

This creates a module called _parga (the underscore is a convention for the compiled extension). Your Python package then re-exports from it.

Python Wrapper Layer

The low-level PyO3 bindings are fast but not always ergonomic. A Python wrapper provides a nicer API:

# python/parga/ga.py
from parga._parga import RealGenome, run_ga as _run_ga

class GA:
    def __init__(self, fitness_fn, genome_length, **kwargs):
        self.fitness_fn = fitness_fn
        self.genome_length = genome_length
        self.config = kwargs

    def run(self):
        # Auto-select strategy based on fitness function cost
        strategy = self._select_strategy()

        if strategy == 'rust':
            return _run_ga(self.fitness_fn, self.genome_length, **self.config)
        else:
            return self._run_parallel()

This two-layer approach gives you the best of both worlds: Rust performance for the hot path, Python flexibility for configuration and orchestration.

Building and Publishing

Maturin makes the build process simple:

# Development build (installs to current venv)
maturin develop

# Release build
maturin build --release

# Publish to PyPI
maturin publish

For CI/CD, Maturin can build wheels for multiple platforms. ParGA’s GitHub Actions workflow builds for Linux, macOS, and Windows across multiple Python versions.

Performance Considerations

A few things I learned building ParGA:

Minimize Python-Rust boundary crossings. Each call from Python into Rust has overhead. Batch operations when possible. Instead of evaluating one genome at a time, pass arrays.

Use zero-copy where you can. NumPy arrays can be passed to Rust without copying if you’re careful. The numpy crate provides PyReadonlyArray for this.

Profile before optimizing. Some operations that seem slow in Python are actually fine. The GIL only matters for CPU-bound parallel work. For I/O or calling external libraries, pure Python is often sufficient.

Consider the GIL for parallelism. Rust code releases the GIL by default during computations, but fitness function callbacks into Python re-acquire it. For expensive Python fitness functions, ParGA uses ProcessPoolExecutor instead of threads.

Why This Matters

The Python scientific ecosystem is remarkable, but performance has always been its Achilles’ heel. NumPy and friends paper over the problem for array operations, but algorithmic code in pure Python is slow.

Traditionally, the solution was writing C extensions. This worked but was painful. PyO3 and Maturin make Rust a realistic alternative that’s actually pleasant to use.

For ParGA, the result is a library that’s 5-50x faster than pure Python implementations while maintaining the simple API Python users expect. The complexity of the Rust backend is invisible to anyone who just wants to call minimize().

Resources

Next in the series: the fundamentals of genetic algorithms and why they’re still relevant for hard optimization problems.