malevich

Look up

Performance

M4 to the raster, bucket-exact grids, and a bench behind the numbers.

Fast is a feature, and claims are measured. Every advertised number has a bench behind it, recorded with its machine and date. This file is the public story. BENCHMARKS.md is the authoritative dated record it summarizes.

Speed and honesty coexist because of the oracle. The fast path is proven byte-identical to drawing every point, so there is no fidelity knob to trade away. See The full draw is the oracle.

The mechanisms#

200,000 points through M4 8 ┤ │ 6 ┤ 4 ┤ │ 2 ┤ │ 0 ┤ │ -2 ┤ -4 ┤ └┬────────┬─────────┬────────┬─────────┬────────┬────────┬─────────┬────────┬ 0 25k 50k 75k 100k 125k 150k 175k 200k the same points, every 400th 8 ┤ │ 6 ┤ 4 ┤ │ 2 ┤ │ 0 ┤ │ -2 ┤ -4 ┤ └┬────────┬─────────┬────────┬─────────┬────────┬────────┬─────────┬────────┬ 0 25k 50k 75k 100k 125k 150k 175k 200k
Plate 1. Speed and honesty are the same claim. On the left, M4 keeps the first, last, minimum, and maximum per rendered column, so the three one-sample spikes survive by construction. On the right, a stride sampler is just as fast and silently lost all three.
180,000 cells, max-reduced 3.0 ┤ 1.00 │ 2.5 ┤ │ ▓ 0.75 k 2.0 ┤ ▓ H │ ▓ z 1.5 ┤ ▒ 0.50 1.0 ┤ ▒ │ ▒ 0.25 0.5 ┤ ░ │ ░ 0.0 ┤ ░ 0.00 └┬────────┬─────────┬─────────┬────────┬─────────┬────────┬ 0 1 2 3 4 5 6 s
Plate 2. Cells::extents and reduce max-reduce a grid denser than the raster, so its spikes survive.
  • Lines: M4 to the raster. Large line layers reduce to min/max/first/last per raster column, bucketed by the column each point renders into. O(n) once, then O(width × height), and pixel-identical to the full draw by construction. The reduction is auto-inserted past four points per column.
  • Grids: bucket-exact reduction. Cells grids denser than the raster reduce through the shared Reducer vocabulary. Every screen bucket owns the cells whose centers fall inside it. Cost is linear in the grid, about 10 ns per cell, not in the raster.
  • Geometry without a raster. Plot::mapping, what an interactive host calls between renders, runs the extent probe and the layout pass only. The aggregation is never computed for a raster nobody draws.
  • No allocation per point. Rendering retains labels and identities at construction and borrows them after that. CI enforces allocation ceilings (at most 275 allocations and 64 KiB for the 10k-point render), so a structural regression — an allocation per point, a new large intermediate — fails the build.

Measured#

the recorded baseline, Apple M1 Pro, single thread │ line, 10k points ┤ line, 10M points ┤ cells 2048² → 80×24 ┤ fit, 1M pairs ┤ color_by, 5 groups ┤ │ color_by, 100k groups ┤ mapping, 10M ┤ dashboard 200×50 ┤ zoom 10M, 200×50 ┤ hover snap, 10M ┤ │ └───┬───────────────────┬────────────────────┬──────────── 10⁻⁴ 10⁻³ 10⁻²
Plate 3. The benchmark table, drawn by the library it measures. Horizontal bars on a log axis, with an SI unit, from the 2026-09-24 baseline.
The code that drew it
use malevich::scale::Unit;
use malevich::{Bars, Plot, Scale};
// The benchmark suite, plotted by the library it measures (BENCHMARKS.md, 2026-09-24).
let rows = [
    "line, 10k points", "line, 10M points", "cells 2048² → 80×24", "fit, 1M pairs",
    "color_by, 5 groups", "color_by, 100k groups", "mapping, 10M", "dashboard 200×50",
    "zoom 10M, 200×50", "hover snap, 10M",
];
let seconds = vec![69.3e-6, 30.7e-3, 40.6e-3, 5.13e-3, 2.02e-3, 5.35e-3, 2.08e-3, 1.77e-3, 18.4e-3, 25.4e-3];
Plot::new()
    .y_scale(Scale::bands(rows))
    .layer(Bars::new(rows, seconds).horizontal())
    .log_x()
    .x_unit(Unit::si("s"))
    .title("the recorded baseline, Apple M1 Pro, single thread")

Two machines, single-threaded, end to end: construct, resolve, reduce, rasterize, encode, at 1.23.0. Order of magnitude and which mechanism wins, not a promise for your machine. The same code runs 1× to 20× slower on the virtualized Xeon than on the laptop, and not uniformly.

measurementApple M1 Prox86_64 Xeon (KVM)
line, 10,000 points, 80×2069 µs152 µs
line, 10,000,000 points, 80×2031 ms142 ms
cells, 2048×2048 grid onto 80×2441 ms42 ms
geometry only (Plot::mapping), 10,000,000 points2.1 ms46 ms
streaming least squares, 1M pairs (stat::Fit)5.1 ms6.4 ms
100,000 points, 5 categories via color_by2.0 ms3.4 ms
100,000 points, 100,000 categories5.3 ms8.0 ms
one interactive frame, 10,000,000 points at a 1% zoom, 200×5018 ms166 ms
the same frame with a hover cursor25 ms199 ms
two-pane dashboard, 200×50 (100k-point lines beside a 256×128 heatmap)1.8 ms4.1 ms

The categorical pair is a structural fence. Runtime grows with input plus legend size, not input × category count. The ten-million-point row is the README's "tens of milliseconds" claim, on the M1 Pro record it cites. The cells row is its matrix analog, 4.19 million cells onto ~4k screen buckets, and the one row the two machines agree on. The interactive rows are what a TUI pays per frame while someone pans or zooms.

The Xeon column is the first record on a second architecture, and it says something the laptop alone could not. The per-point walks are where the machines diverge most: M4, the extent probe, the hover cursor's nearest scan. The per-bucket and per-cell work costs about the same on both. On the Xeon the M4 reduction alone (stat/m4_10m_160cols) is 127 of the 142 ms. The record states those ratios. It does not explain them.

Rerun on yours#

cargo bench --bench render -- render/line_10m_80x20
cargo bench --bench render -- render/cells_2048x2048_80x24
cargo bench --bench alloc
cargo run --example bench_record       # the saved results as a record block

The bench suite is the only source docs quote numbers from. bench_record prints its saved results in the record's format: Criterion's own estimates and intervals, with the machine, OS, compiler, and revision they came from. No number in the record is typed by hand. Baselines, machine details, confidence intervals, and the update protocol live in BENCHMARKS.md.