malevich

Look up

The statistics layer

Bins, densities, quartiles, fits, windows, stacks, M4. What runs before the scales see the data.

A stat is a data operation that runs before the scales see the data. The word follows seaborn's objects interface and ggplot's stat_*. It is the module-level umbrella. It is not one execution algebra. A stat may be an online accumulator, a reducer, keyed orchestration, or a batch transform. The layer refuses a fake uniform Stat trait, because pretending every statistic streams and merges is how libraries ship wrong parallel medians.

Everything here lives in malevich::stat. This page walks the module with a plate per idea.

Bins and histograms#

Bins is a histogram accumulator: a start, a width, a count, and the counts. Bins::auto(values, limit) chooses the geometry. The Freedman–Diaconis width comes from type-7 quartiles. The whole finite span is binned whole. There are at most limit bins. add streams values in. Bars::spans draws it.

Bins::auto + Bars::spans + Scale::Integer == hist 200 ┤ │ 150 ┤ │ │ 100 ┤ │ 50 ┤ │ 0 ┤ └┬──────┬──────┬───────┬──────┬──────┬───────┬──────┬──────┬ 2500 3000 3500 4000 4500 5000 5500 6000 6500 body mass, g
Plate 1. Bins::auto and Bars::spans spell out the histogram preset.
The code that drew it
use malevich::stat::Bins;
use malevich::{Bars, Plot, Scale};
use super::penguins;
// The histogram, spelled out: Bins::auto chooses the geometry, Bars::spans draws the counts.
let mass = penguins().mass;
let mut bins = Bins::auto(&mass, 60).expect("finite data");
for value in &mass {
    bins.add(*value);
}
let counts: Vec<f64> = bins.counts().iter().map(|&c| c as f64).collect();
Plot::new()
    .layer(Bars::spans(bins.start(), bins.width(), counts))
    .y_scale(Scale::Integer)
    .title("Bins::auto + Bars::spans + Scale::Integer == hist")
    .x_label("body mass, g")

That is the entire hist preset, and the packaging test asserts it. HistogramOptions adds what a histogram sometimes needs: a bin cap, a Normalization (count, probability, percent, density per unit of x), and cumulative. kaz hist --normalize percent --cumulative is the same options value from the shell.

cumulative share of requests 100% ┤ │ 75% ┤ │ % │ 50% ┤ │ 25% ┤ │ 0% ┤ └┬─────────┬──────────┬─────────┬──────────┬──────────┬── 0 20 40 60 80 100 ms
Plate 2. HistogramOptions rescales the bins to percent and accumulates them left to right.
The code that drew it
use malevich::stat::Normalization;
use malevich::{hist_with, HistogramOptions};
let mut unit = super::noise(3);
// Latencies: log-normal, the long tail every service has.
let latency: Vec<f64> = (0..4000).map(|_| (0.4 * super::gaussian(&mut unit) + 3.2).exp()).collect();
hist_with(latency, HistogramOptions::new(40).normalization(Normalization::Percent).cumulative(true))
    .expect("valid options")
    .title("cumulative share of requests")
    .x_label("ms")
    .y_label("%")

bins2 is the two-dimensional form behind hist2d. calendar_bins counts unix timestamps per hour, day, ISO week, month, or year. The buckets have their true length. Empty ones are kept. They feed Bars::intervals.

stat::calendar_bins — incidents per month │ 40 ┤ │ 30 ┤ │ │ 20 ┤ │ 10 ┤ │ 0 ┤ └┬───────────────┬────────────────┬────────────────┬────────────────┬ 2025 Apr Jul Oct 2026
Plate 3. calendar_bins counts per calendar month and draws each count between that month's true edges.
The code that drew it
use malevich::stat::{calendar_bins, TimeUnit};
use malevich::{Bars, Plot};
use super::day_stamp;
let mut unit = super::noise(29);
// Incidents over a year: counted per calendar month, each bar drawn between
// its month's true edges on the time axis (February is shorter).
let stamps: Vec<f64> = (0..300).map(|_| day_stamp(2025, 1, 1) + unit() * 365.0 * 86_400.0 * (0.4 + 0.6 * unit())).collect();
let months = calendar_bins(&stamps, TimeUnit::Month).expect("data");
let counts: Vec<f64> = months.counts().iter().map(|&c| c as f64).collect();
Plot::new()
    .layer(Bars::intervals(months.starts().to_vec(), months.ends().to_vec(), counts))
    .time_x()
    .title("stat::calendar_bins — incidents per month")

Densities#

kde(values, points) is a Gaussian kernel density estimate at Silverman's bandwidth. kde_with takes KdeOptions. A Bandwidth is Silverman, a scale of it, or a fixed width. bounds reflect the kernels, so a latency density stays above zero. There is a cut in bandwidths past the data, and a cumulative form. The bandwidth is the one honest knob a density has, and the plate shows why it is worth showing three.

kernel density, three bandwidths ── ½ Silverman ── Silverman (default) ── 2× Silverman 600µ ┤ │ 500µ ┤ 400µ ┤ │ 300µ ┤ │ 200µ ┤ │ 100µ ┤ 0 ┤ └┬────────┬────────┬─────────┬────────┬─────────┬────────┬────────┬ 1000 2000 3000 4000 5000 6000 7000 8000 body mass, g
Plate 4. kde_with draws the same data at three bandwidths.
The code that drew it
use malevich::stat::{kde_with, Bandwidth, KdeOptions};
use malevich::{Line, Plot};
use super::penguins;
// The bandwidth is the one honest knob a density has: the same 342 masses
// at Silverman's rule, at half of it, and at twice it.
let mass = penguins().mass;
let mut plot = Plot::new();
for (scale, label) in [(0.5, "½ Silverman"), (1.0, "Silverman (default)"), (2.0, "2× Silverman")] {
    let options = KdeOptions::new().bandwidth(Bandwidth::Scale(scale));
    let (x, y) = kde_with(&mass, 200, options).expect("valid").expect("data");
    plot = plot.layer(Line::xy(x, y).label(label));
}
plot.title("kernel density, three bandwidths").x_label("body mass, g")

The density and violin presets are built on it; the gallery's latency example shows the bounded estimate beside the unbounded one that leaks below zero.

Quartiles, whiskers, and moments#

BoxStats::of(values) computes type-7 quartiles (Hyndman and Fan's estimator, the one R and NumPy default to), Tukey whiskers at 1.5 IQR, and the outliers beyond them. Whiskers::Percentiles(lo, hi) and Whiskers::MinMax are the other rules. BoxOptions::whiskers passes one to the preset.

body mass, whiskers at p5 and p95 │ 6000 ┤ │ │ 5000 ┤ │ ━━━━━━━━ g │ │ 4000 ┤ │ ━━━━━━━━ ━━━━━━━━ │ 3000 ┤ │ └──────────────────────────────────────────────────── Adelie Chinstrap Gentoo
Plate 5. BoxOptions::whiskers uses a percentile rule instead of Tukey's.
The code that drew it
use malevich::stat::Whiskers;
use malevich::{box_plot_with, BoxOptions};
use super::penguins;
let p = penguins();
let groups = ["Adelie", "Chinstrap", "Gentoo"].map(|s| p.of(s, &p.mass));
// Whiskers to the 5th and 95th percentiles instead of Tukey's 1.5 IQR.
box_plot_with(["Adelie", "Chinstrap", "Gentoo"], groups, BoxOptions::new().whiskers(Whiskers::Percentiles(0.05, 0.95)))
    .expect("valid")
    .title("body mass, whiskers at p5 and p95")
    .y_label("g")

Moments is the online accumulator behind describe: count, mean, variance, standard deviation (population and sample), min, max. It updates one observation at a time, and it merges, because its summary state is order-independent. quantiles(values, positions) is the type-7 estimator on its own, shared by the box plot, the reducers, and the Q–Q plot, so there is exactly one quartile implementation in the crate.

flipper length, mm │ Adelie ┤ 151 190.0 6.539 172 186 190 195 210 Chinstrap ┤ 68 195.8 7.132 178 191 196 201 212 Gentoo ┤ 123 217.2 6.485 203 212 216 221 231 │ └───────────────────────────────────────────────────────────────────── count mean sd min p25 p50 p75 max
Plate 6. describe is the summary that usually precedes a chart, drawn as a stat table.
The code that drew it
use super::penguins;
let p = penguins();
let groups = ["Adelie", "Chinstrap", "Gentoo"].map(|s| p.of(s, &p.flipper));
// Moments, quantiles, and extremes as a stat table: text on two band axes.
malevich::describe(["Adelie", "Chinstrap", "Gentoo"], groups).title("flipper length, mm")

The reducer vocabulary#

One Reducer serves every aggregating stat: Count, Sum, Mean, Median, Min, Max, Percentile(q), Deviation, Variance, StdErr, First, Last. Bins, groups, windows, and dense grids all take one, so a rolling p95, a binned median, a group's mean ± standard error, or a max-reduced heatmap is one call and zero new concepts.

Window is the rolling form: a size, an anchor (trailing by default, or Middle, or Start), and strict. strict gaps the positions whose window is incomplete. It does not guess.

Window: one reducer vocabulary ── samples ── rolling mean ── rolling p95 100 ┤ │ │ 80 ┤ │ 60 ┤ m │ s │ 40 ┤ │ │ 20 ┤ │ 0 ┤ └┬────────┬────────┬────────┬────────┬────────┬────────┬────────┬────────┬ 0 50 100 150 200 250 300 350 400
Plate 7. Window draws a rolling mean and a rolling p95 from one reducer vocabulary.
The code that drew it
use malevich::stat::{Reducer, Window};
use malevich::{Line, Plot};
let mut unit = super::noise(5);
// Latency samples with a spiky tail; a rolling mean and a rolling p95 over 50.
let samples: Vec<f64> = (0..400).map(|i| {
    let base = 40.0 + 10.0 * (f64::from(i) * 0.03).sin();
    if unit() < 0.06 { base + 60.0 * unit() } else { base + 6.0 * super::gaussian(&mut unit) }
}).collect();
let window = Window::new(50);
Plot::new()
    .layer(Line::y(samples.clone()).label("samples"))
    .layer(Line::y(window.mean(&samples)).label("rolling mean"))
    .layer(Line::y(window.reduce(&samples, Reducer::Percentile(0.95))).label("rolling p95"))
    .title("Window: one reducer vocabulary")
    .y_label("ms")
window anchors ── signal ── trailing ── centered, strict 10 ┤ │ 8 ┤ │ 6 ┤ │ │ 4 ┤ │ 2 ┤ │ 0 ┤ └┬───────┬────────┬───────┬────────┬───────┬───────┬────────┬───────┬ 0 25 50 75 100 125 150 175 200
Plate 8. WindowAnchor is trailing, centered, or strict.
The code that drew it
use malevich::stat::{Window, WindowAnchor};
use malevich::{Line, Plot};
let mut unit = super::noise(8);
let signal: Vec<f64> = (0..200).map(|i| if (60..70).contains(&i) { 10.0 } else { 2.0 + super::gaussian(&mut unit) * 0.4 }).collect();
// A trailing window lags the step; a centered one straddles it; `strict`
// gaps the positions whose window is incomplete instead of guessing.
Plot::new()
    .layer(Line::y(signal.clone()).label("signal"))
    .layer(Line::y(Window::new(21).mean(&signal)).label("trailing"))
    .layer(Line::y(Window::new(21).anchor(WindowAnchor::Middle).strict().mean(&signal)).label("centered, strict"))
    .title("window anchors")

binned(x, y, &bins, reducer) reduces y over the bins of x — the reliability diagram of a classifier is Reducer::Mean over 0/1 outcomes binned by confidence.

stat::binned with Reducer::Mean accuracy per bin ── perfect calibration 1.00 ┤ a │ c 0.75 ┤ c │ u │ r 0.50 ┤ a │ c │ y 0.25 ┤ │ 0.00 ┤ └┬──────────┬──────────┬──────────┬──────────┬──────────┬ 0.0 0.2 0.4 0.6 0.8 1.0 confidence
Plate 9. binned with Reducer::Mean builds a reliability diagram from 0/1 outcomes.
The code that drew it
use malevich::stat::{binned, Bins, Reducer};
use malevich::{Bars, Line, Plot, Points};
let mut unit = super::noise(19);
// Confidence versus accuracy: 0/1 outcomes, binned by confidence, mean per bin.
let confidence: Vec<f64> = (0..2000).map(|_| unit()).collect();
let correct: Vec<f64> = confidence.iter().map(|c| if unit() < c * 0.8 + 0.1 { 1.0 } else { 0.0 }).collect();
let bins = Bins::new(0.0, 0.1, 10);
let accuracy = binned(&confidence, &correct, &bins, Reducer::Mean);
let centers: Vec<f64> = (0..10).map(|i| 0.05 + 0.1 * i as f64).collect();
Plot::new()
    .layer(Bars::spans(0.0, 0.1, accuracy.clone()).label("accuracy per bin"))
    .layer(Points::xy(centers, accuracy))
    .layer(Line::xy(vec![0.0, 1.0], vec![0.0, 1.0]).label("perfect calibration"))
    .title("stat::binned with Reducer::Mean")
    .x_label("confidence")
    .y_label("accuracy")

Agg::by(keys, values) groups by a categorical key and reduces each group. Cells::reduce applies the same vocabulary to grids denser than the raster.

Agg::by(...).reduce(Reducer::Median) 5000 ┤ 4000 ┤ │ 3000 ┤ g │ 2000 ┤ │ 1000 ┤ 0 ┤ └──────────────────────────────────────────────── Adelie Chinstrap Gentoo
Plate 10. Agg::by and a reducer draw one bar per group.
The code that drew it
use malevich::stat::{Agg, Reducer};
use malevich::{Bars, Plot};
use super::penguins;
// Group by a key, reduce each group: the median mass per species.
let p = penguins();
let (species, medians) = Agg::by(p.species, p.mass).reduce(Reducer::Median);
Plot::new()
    .layer(Bars::new(species, medians))
    .title("Agg::by(...).reduce(Reducer::Median)")
    .y_label("g")

Smoothing, explicitly#

ewma(values, alpha) is a debiased exponential moving average. lttb(x, y, target) is Largest-Triangle-Three-Buckets. Both are inexact transforms, and both exist only as explicit stats the caller applies. They are never a silent default reduction. The default reduction is M4, below, which is exact.

topos bigram model, 1,000 steps ── per-step loss ── ewma α = 0.05 │ 3.2 ┤ │ l 3.0 ┤ o │ s 2.8 ┤ s │ │ 2.6 ┤ │ 2.4 ┤ └┬────────────┬────────────┬───────────┬────────────┬────────────┬ 0 200 400 600 800 1000 step
Plate 11. ewma runs over a real training log.
The code that drew it
use malevich::stat::ewma;
use malevich::{Line, Plot};
use super::loss_log;
// A real training log, and its debiased exponential moving average.
let loss = loss_log();
Plot::new()
    .layer(Line::y(loss.clone()).label("per-step loss"))
    .layer(Line::y(ewma(&loss, 0.05)).label("ewma α = 0.05").glow())
    .title("topos bigram model, 1,000 steps")
    .x_label("step")
    .y_label("loss")
stat::lttb — an opt-in approximation ── 3,000 points •• 120 kept by LTTB 3 ┤ │ 2 ┤ │ 1 ┤ 0 ┤ │ -1 ┤ │ -2 ┤ │ -3 ┤ └┬──────────┬──────────┬───────────┬──────────┬──────────┬──────────┬ 0 5 10 15 20 25 30
Plate 12. lttb is an explicit, disclosed approximation the caller applies.
The code that drew it
use malevich::stat::lttb;
use malevich::{Line, Plot, Points};
let x: Vec<f64> = (0..3000).map(|i| f64::from(i) * 0.01).collect();
let y: Vec<f64> = x.iter().map(|x| (x * 2.0).sin() * (x * 0.3).cos() * 3.0 + (x * 17.0).sin() * 0.3).collect();
// LTTB keeps 120 visually representative points. It is explicit and inexact —
// a stat the caller applies — never a default reduction.
let (kept_x, kept_y) = lttb(&x, &y, 120);
Plot::new()
    .layer(Line::xy(x, y).label("3,000 points"))
    .layer(Points::xy(kept_x, kept_y).label("120 kept by LTTB"))
    .title("stat::lttb — an opt-in approximation")

Fits#

Fit is streaming least squares: add(x, y) or Fit::xy(&x, &y), then slope, intercept, r_squared, predict(x), standard_error(x). It merges. Two halves of a dataset fitted separately combine exactly, because its summary state is order-independent. The trend preset draws the scatter, the line, and a 95 % confidence band from one accumulator.

Gentoo mass by flipper — slope 54.6 g/mm, R² 0.49 6500 ┤ │ 6000 ┤ │ │ 5500 ┤ g │ │ 5000 ┤ │ 4500 ┤ │ │ 4000 ┤ └────┬───────────┬──────────┬──────────┬───────────┬──────────┬── 205 210 215 220 225 230 flipper, mm
Plate 13. trend and the Fit behind it show the points, the line, a confidence band, and the numbers.
The code that drew it
use malevich::stat::Fit;
use super::penguins;
let p = penguins();
let (x, y) = (p.of("Gentoo", &p.flipper), p.of("Gentoo", &p.mass));
let fit = Fit::xy(&x, &y);
// The preset draws the points, the line, and the band; the accumulator answers the numbers.
malevich::trend(x, y)
    .title(format!("Gentoo mass by flipper — slope {:.1} g/mm, R² {:.2}", fit.slope().unwrap_or(0.0), fit.r_squared().unwrap_or(0.0)))
    .x_label("flipper, mm")
    .y_label("g")

Distributions as curves#

ecdf(values) is the empirical cumulative distribution as a step. The ecdf_with preset adds the Dvoretzky–Kiefer–Wolfowitz band at a chosen level. roc(scores, labels) sweeps the thresholds. auc(x, y) integrates the result.

ECDF with a DKW band 1.0 ┤ │ 0.8 ┤ │ 0.6 ┤ │ │ 0.4 ┤ │ 0.2 ┤ │ 0.0 ┤ └┬────────┬─────────┬─────────┬─────────┬────────┬─────────┬ 0 10 20 30 40 50 60 ms
Plate 14. ecdf_with draws its curve and DKW band.
The code that drew it
use malevich::{ecdf_with, EcdfOptions};
let mut unit = super::noise(13);
let latency: Vec<f64> = (0..600).map(|_| (0.35 * super::gaussian(&mut unit) + 3.0).exp()).collect();
// The empirical CDF, with the Dvoretzky–Kiefer–Wolfowitz 95 % band.
ecdf_with(latency, EcdfOptions::new().band(0.05))
    .expect("valid")
    .title("ECDF with a DKW band")
    .x_label("ms")
ROC, AUC = 0.820 ── classifier ── chance t 1.0 ┤ r │ u 0.8 ┤ e │ │ p 0.6 ┤ o │ s 0.4 ┤ i │ t │ i 0.2 ┤ v │ e 0.0 ┤ └┬──────────┬──────────┬───────────┬──────────┬──────────┬ 0.0 0.2 0.4 0.6 0.8 1.0 false positive rate
Plate 15. roc and auc show the curve, the chance line, and the area in the title.
The code that drew it
use malevich::stat::{auc, roc};
use malevich::{Dash, Line, Plot};
let mut unit = super::noise(17);
// Scores for 400 examples: positives score higher on average, with overlap.
let labels: Vec<bool> = (0..400).map(|i| i % 3 == 0).collect();
let scores: Vec<f64> = labels.iter().map(|&positive| super::gaussian(&mut unit) + if positive { 1.4 } else { 0.0 }).collect();
let (fpr, tpr) = roc(&scores, &labels);
let area = auc(&fpr, &tpr);
Plot::new()
    .layer(Line::xy(fpr, tpr).label("classifier"))
    .layer(Line::xy(vec![0.0, 1.0], vec![0.0, 1.0]).dash(Dash::Dotted).label("chance"))
    .title(format!("ROC, AUC = {area:.3}"))
    .x_label("false positive rate")
    .y_label("true positive rate")

Stack and dodge#

stack(&[&a, &b, &c]) returns a (low, high) pair per series, positives above the baseline and negatives below it, for Area::between or Bars::base. stack_with takes StackOptions. StackOffset::Normalize is the 100 % stack. Center is the streamgraph silhouette. StackOrder::Sum piles the largest series first.

stat::stack — a stacked area coal gas wind solar 160 ┤ │ │ 120 ┤ T │ W 80 ┤ h │ │ 40 ┤ │ 0 ┤ └┬────────────┬────────────┬───────────┬────────────┬────────────┬ 2000 2005 2010 2015 2020 2025 year
Plate 16. stack sits each area on the sum of the ones below.
The code that drew it
use malevich::stat::stack;
use malevich::{Area, Plot};
let years: Vec<f64> = (2000..2026).map(f64::from).collect();
let coal: Vec<f64> = years.iter().map(|y| 40.0 - (y - 2000.0) * 1.1).collect();
let gas: Vec<f64> = years.iter().map(|y| 20.0 + (y - 2000.0) * 0.4).collect();
let wind: Vec<f64> = years.iter().map(|y| 1.0 + ((y - 2000.0) * 0.35).powi(2)).collect();
let solar: Vec<f64> = years.iter().map(|y| 0.2 + ((y - 2000.0) * 0.28).powi(2)).collect();
// stack returns (low, high) per series; each Area sits on the sum of the ones below.
let mut plot = Plot::new();
for ((low, high), label) in stack(&[&coal, &gas, &wind, &solar]).into_iter().zip(["coal", "gas", "wind", "solar"]) {
    plot = plot.layer(Area::between(years.clone(), low, high).label(label));
}
plot.title("stat::stack — a stacked area").x_label("year").y_label("TWh")
StackOffset::Normalize — shares coal gas wind solar 1.00 ┤ │ 0.75 ┤ │ │ 0.50 ┤ │ │ 0.25 ┤ │ 0.00 ┤ └┬────────────┬────────────┬────────────┬────────────┬────────────┬ 2000 2005 2010 2015 2020 2025 year
Plate 17. StackOffset::Normalize makes the 100 % stack.
The code that drew it
use malevich::stat::{stack_with, StackOffset, StackOptions};
use malevich::{Area, Plot};
let years: Vec<f64> = (2000..2026).map(f64::from).collect();
let coal: Vec<f64> = years.iter().map(|y| 40.0 - (y - 2000.0) * 1.1).collect();
let gas: Vec<f64> = years.iter().map(|y| 20.0 + (y - 2000.0) * 0.4).collect();
let wind: Vec<f64> = years.iter().map(|y| 1.0 + ((y - 2000.0) * 0.35).powi(2)).collect();
let solar: Vec<f64> = years.iter().map(|y| 0.2 + ((y - 2000.0) * 0.28).powi(2)).collect();
// Normalize: every x sums to one, so the chart reads as shares.
let options = StackOptions::new().offset(StackOffset::Normalize);
let mut plot = Plot::new();
for ((low, high), label) in stack_with(&[&coal, &gas, &wind, &solar], options).into_iter().zip(["coal", "gas", "wind", "solar"]) {
    plot = plot.layer(Area::between(years.clone(), low, high).label(label));
}
plot.title("StackOffset::Normalize — shares").x_label("year")

dodge(&[&a, &b], step) is stack's sibling for bars beside each other: side-by-side positions within each band, one series per value series, fed to Bars::at. Grouped bars are therefore a composition. They are not a preset — the presets principle applied to the request the field files most.

stat::dodge — grouped bars morning evening 30 ┤ │ │ 20 ┤ │ │ 10 ┤ │ │ 0 ┤ └──────────────────────────────────────────────────────── Oslo Lima Osaka Cairo
Plate 18. dodge sets side-by-side positions within each band.
The code that drew it
use malevich::stat::dodge;
use malevich::{Bars, Plot, Scale};
let cities = ["Oslo", "Lima", "Osaka", "Cairo"];
let morning = vec![14.0, 22.0, 9.0, 31.0];
let evening = vec![11.0, 19.0, 13.0, 27.0];
// dodge computes side-by-side positions within each band; Bars::at draws there.
let width = 0.38;
let positions = dodge(&[&morning, &evening], width);
Plot::new()
    .x_scale(Scale::bands(cities))
    .layer(Bars::at(positions[0].clone(), width, morning).label("morning"))
    .layer(Bars::at(positions[1].clone(), width, evening).label("evening"))
    .title("stat::dodge — grouped bars")

Positions and shapes#

jitter(positions, width) spreads a strip of points evenly across its band with van der Corput offsets: no seed, the same every time, no clumps. steps(x, y, direction) is the piecewise-constant expansion the stairs preset draws, changing after, before, or midway between samples. nearest is the crosshair-snapping lookup interactive hosts use: the index of the closest finite value. The readout shows a datum that exists. It does not show an interpolation.

stat::jitter │ 6000 ┤ │ │ 5000 ┤ │ g │ │ 4000 ┤ │ │ 3000 ┤ │ └──────────────────────────────────────────────────────── Adelie Chinstrap Gentoo
Plate 19. jitter spreads every measurement evenly across its band.
The code that drew it
use malevich::stat::jitter;
use malevich::{Plot, Points, Scale};
use super::penguins;
// A strip plot: every measurement, spread across its band by van der Corput
// offsets — even, seedless, the same every time.
let p = penguins();
let species = ["Adelie", "Chinstrap", "Gentoo"];
let mut plot = Plot::new().x_scale(Scale::bands(species));
for (index, name) in species.iter().enumerate() {
    let masses = p.of(name, &p.mass);
    let positions = jitter(&vec![index as f64; masses.len()], 0.7);
    plot = plot.layer(Points::xy(positions, masses));
}
plot.title("stat::jitter").y_label("g")
stat::steps — where a value changes •• samples ── Post ── Pre (+6) ── Mid (+12) 20 ┤ │ │ 15 ┤ │ │ 10 ┤ │ 5 ┤ │ │ 0 ┤ └┬────────┬─────────┬─────────┬────────┬─────────┬─────────┬────────┬ 0 1 2 3 4 5 6 7
Plate 20. steps shows the three places a step can change.
The code that drew it
use malevich::stat::{steps, StepDirection};
use malevich::{Line, Plot, Points};
let x: Vec<f64> = (0..8).map(f64::from).collect();
let y = vec![1.0, 3.0, 2.0, 5.0, 4.0, 6.0, 3.0, 4.0];
let (ax, ay) = steps(&x, &y, StepDirection::Post);
let (bx, by) = steps(&x, &y, StepDirection::Pre);
let (mx, my) = steps(&x, &y, StepDirection::Mid);
Plot::new()
    .layer(Points::xy(x, y).label("samples"))
    .layer(Line::xy(ax, ay).label("Post"))
    .layer(Line::xy(bx, by.iter().map(|y| y + 6.0).collect::<Vec<_>>()).label("Pre (+6)"))
    .layer(Line::xy(mx, my.iter().map(|y| y + 12.0).collect::<Vec<_>>()).label("Mid (+12)"))
    .title("stat::steps — where a value changes")

The series maps#

cumsum, diff, rank, and normalize(values, reducer) — a running total, first differences, ranks, and division by a reducer of the whole series. Index charts and percent-of-peak are one call.

cumsum and normalize ── daily, as a share of the peak ── cumulative, share of the total 1.0 ┤ │ │ │ 0.5 ┤ │ │ │ 0.0 ┤ └┬──────┬───────┬──────┬──────┬───────┬──────┬──────┬───────┬──────┬ 0 10 20 30 40 50 60 70 80 90 day
Plate 21. cumsum and normalize are the series maps.
The code that drew it
use malevich::stat::{cumsum, normalize, Reducer};
use malevich::{Line, Plot};
let mut unit = super::noise(23);
let daily: Vec<f64> = (0..90).map(|i| 3.0 + (f64::from(i) * 0.15).sin() * 2.0 + unit()).collect();
// The series maps: a running total, and a series divided by a reducer of itself.
Plot::new()
    .layer(Line::y(normalize(&daily, Reducer::Max)).label("daily, as a share of the peak"))
    .layer(Line::y(normalize(&cumsum(&daily), Reducer::Max)).label("cumulative, share of the total"))
    .title("cumsum and normalize")
    .x_label("day")

Contours#

contours(values, columns, levels) runs marching squares over a grid and returns iso-lines as paths. The contour preset chooses its levels the way an axis chooses ticks. contourf is heatmap under a colormap split at those levels.

contour — peaks ── -6 ── -4 ── -2 ── 0 ── 2 ── 4 ── 6 ── 8 60 ┤ │ 50 ┤ │ │ 40 ┤ │ │ 30 ┤ │ 20 ┤ │ │ 10 ┤ │ 0 ┤ └┬─────────┬─────────┬─────────┬────────┬─────────┬─────────┬ 0 10 20 30 40 50 60
Plate 22. contour runs marching squares over a grid, with levels chosen like ticks.
The code that drew it
// The MATLAB peaks function as iso-lines: marching squares, tick-chosen levels.
let n = 60;
let mut values = Vec::with_capacity(n * n);
for row in 0..n {
    for column in 0..n {
        let (x, y) = (column as f64 / n as f64 * 6.0 - 3.0, 3.0 - row as f64 / n as f64 * 6.0);
        values.push(3.0 * (1.0 - x).powi(2) * (-x * x - (y + 1.0).powi(2)).exp() - 10.0 * (x / 5.0 - x.powi(3) - y.powi(5)) * (-x * x - y * y).exp() - (-(x + 1.0).powi(2) - y * y).exp() / 3.0);
    }
}
malevich::contour(n, values).title("contour — peaks")

M4: the reduction that is not a stat#

m4(x, y, columns) keeps the first, last, minimum, and maximum point of every raster column. Because points are bucketed by the column they render into, the reduction reproduces the full draw's pixels exactly (Jugel et al., PVLDB 2014). The plot inserts it automatically past four points per column. It is public so a host can pre-reduce a series it will render many times.

200,000 points through M4 8 ┤ │ 6 ┤ 4 ┤ │ 2 ┤ │ 0 ┤ │ -2 ┤ -4 ┤ └┬────────┬─────────┬────────┬─────────┬────────┬────────┬─────────┬────────┬ 0 25k 50k 75k 100k 125k 150k 175k 200k
Plate 23. M4 reduces 200,000 points, and the three one-sample spikes survive by construction.
The code that drew it
use malevich::{Line, Plot};
let mut unit = super::noise(21);
// 200,000 points with three one-sample spikes. Past four points per column
// the plot reduces by M4 automatically: first, last, min, and max per
// rendered column — the spikes cannot vanish, by construction.
let n = 200_000;
let signal: Vec<f64> = (0..n).map(|i| {
    let t = i as f64 / n as f64;
    let wave = (t * 40.0).sin() * (t * 3.0).cos() * 2.0 + super::gaussian(&mut unit) * 0.2;
    if i == 31_337 || i == 99_999 || i == 150_001 { wave + 7.0 } else { wave }
}).collect();
Plot::new().layer(Line::y(signal)).title("200,000 points through M4").y_min(-4.0).y_max(8.0)

The alternative every sampling library ships, drawn honestly for comparison — the same series, every 400th point:

the same points, every 400th 8 ┤ │ 6 ┤ 4 ┤ │ 2 ┤ │ 0 ┤ │ -2 ┤ -4 ┤ └┬────────┬─────────┬────────┬─────────┬────────┬────────┬─────────┬────────┬ 0 25k 50k 75k 100k 125k 150k 175k 200k
Plate 24. The same series sampled every 400th point. The spikes are gone, and nothing says so.
The code that drew it
use malevich::{Line, Plot};
let mut unit = super::noise(21);
// The same series, "reduced" the way a sampler would: every 400th point.
// The spikes are gone, and nothing on the chart says so.
let n = 200_000;
let signal: Vec<f64> = (0..n).map(|i| {
    let t = i as f64 / n as f64;
    let wave = (t * 40.0).sin() * (t * 3.0).cos() * 2.0 + super::gaussian(&mut unit) * 0.2;
    if i == 31_337 || i == 99_999 || i == 150_001 { wave + 7.0 } else { wave }
}).collect();
let x: Vec<f64> = (0..n).step_by(400).map(|i| i as f64).collect();
let sampled: Vec<f64> = signal.iter().step_by(400).cloned().collect();
Plot::new().layer(Line::xy(x, sampled)).title("the same points, every 400th").y_min(-4.0).y_max(8.0)

Three one-sample spikes are gone, and nothing on the chart says so. The oracle test in the crate asserts raw raster equals reduced raster at several frame sizes; a reduction change that breaks it is a different chart, not an optimization (the full draw is the oracle).

Execution contracts#

The stats keep distinct contracts, stated per type:

kindexamplescontract
online accumulatorMoments, Fit, Bins, M4bounded state, one observation at a time. Moments and Fit merge in any order. Bins needs identical geometry. M4 needs chunks in series order
reducerReducer::*a result for one collection. No public merge
batch transformWindow, kde, ecdf, roc, ewma, lttb, stack, contours, bins2, BoxStatsconsumes a complete ordered collection and emits another. It may use accumulators inside. That does not make it mergeable

Merge results are understood over a fixed reduction tree. Reassociation is not bitwise-independent. Every type states its own identity and preconditions in its docs.