Benchmarks

Brook 0.1.0 on an NVIDIA GeForce RTX 4090 against Kimimaro 5.8.1 on an Intel Core i9-14900KF, Linux. Every time is a complete call, NumPy labels in, osteoid skeletons out.

Twelve datasets

Kimimaro’s default TEASAR parameters for each dataset’s voxel size (its benchmark parameters for the Kimimaro benchmark volume), dust threshold 1000, border and branching correction on. Kimimaro ran with 8 workers on 8 cores, one run per dataset; Brook’s time is the median of three runs after one warmup.

Brook and Kimimaro on 12 datasets: time per volume, skeleton length and distance between skeletons12 of 12 datasets complete. Speedup 9.9× to 134.3× (median 38.5×). Total cable length -7.93% to -0.49%; mean nearest-vertex distance at most 0.24 voxels. Kimimaro 5.8.1 on 8 CPU cores, one run; Brook median of three.DatasetTime per volume, log scaleSpeedupCable ΔNearest vertexBrook, median of 3Kimimaro 5.8.1, 8 CPU coresCREMI B22.5×-3.03%0.11CREMI C20.9×-1.33%0.09CREMI A34.0×-0.49%0.07Kasthuri1135.0×-1.90%0.12MICrONS66.8×-1.21%0.07H0145.3×-1.10%0.09FAFB FFN131.2×-1.55%0.07FIB-2555.3×-0.77%0.04Hemibrain123.0×-1.41%0.09Scroll fibres 16 µm42.1×-7.93%0.24Scroll fibres 8 µm9.9×-5.26%0.17Kimimaro benchmark134.3×-1.45%0.091 s10 s100 s1,000 s−8%000.3 voxCable Δ: Brook’s total skeleton length relative to Kimimaro’s. Nearest vertex: average distance between the two skeletons, in voxels.Kimimaro 5.8.1 on 8 CPU cores (one run per dataset); Brook on one NVIDIA RTX 4090 (median of three).
Brook and Kimimaro on 12 datasets: time per volume, skeleton length and distance between skeletons12 of 12 datasets complete. Speedup 9.9× to 134.3× (median 38.5×). Total cable length -7.93% to -0.49%; mean nearest-vertex distance at most 0.24 voxels. Kimimaro 5.8.1 on 8 CPU cores, one run; Brook median of three.Brook, median of 3Kimimaro, 8 corestime per volume, log scaleCREMI B22.5×cable -3.03% · nearest 0.11 vox · same labelsCREMI C20.9×cable -1.33% · nearest 0.09 vox · same labelsCREMI A34.0×cable -0.49% · nearest 0.07 vox · same labelsKasthuri1135.0×cable -1.90% · nearest 0.12 vox · same labelsMICrONS66.8×cable -1.21% · nearest 0.07 vox · same labelsH0145.3×cable -1.10% · nearest 0.09 vox · same labelsFAFB FFN131.2×cable -1.55% · nearest 0.07 vox · same labelsFIB-2555.3×cable -0.77% · nearest 0.04 vox · same labelsHemibrain123.0×cable -1.41% · nearest 0.09 vox · same labelsScroll fibres 16 µm42.1×cable -7.93% · nearest 0.24 vox · same labelsScroll fibres 8 µm9.9×cable -5.26% · nearest 0.17 vox · same labelsKimimaro benchmark134.3×cable -1.45% · nearest 0.09 vox · same labels1 s10 s100 s1,000 sKimimaro on 8 CPU cores, one run per dataset
DatasetShapeObjectsBrook sKimimaro sSpeedupSame objectsIdenticalLength ΔLength |Δ| median / p95Distance mean / p95EndpointsPieces
CREMI B EM640×640×1253241.03 s (1.03, 1.03, 1.00)23.0 s22.5×Same (324 / 324)39-3.03%0.95% / 7.5%0.11 / 0.311,221 → 1,235393 → 393
CREMI C EM640×640×1255581.08 s (1.12, 1.08, 1.03)22.6 s20.9×Same (558 / 558)112-1.33%0.78% / 6.6%0.09 / 0.242,182 → 2,178821 → 821
CREMI A EM640×640×1251281.05 s (1.05, 1.05, 1.06)35.9 s34.0×Same (128 / 128)32-0.49%0.06% / 1.7%0.07 / 0.06727 → 728156 → 156
Kasthuri11 EM512×512×2569191.06 s (1.08, 1.05, 1.06)37.2 s35.0×Same (919 / 919)128-1.90%1.29% / 5.8%0.12 / 0.262,882 → 2,882972 → 972
MICrONS EM512×512×2561,1981.56 s (1.56, 1.56, 1.55)103.9 s66.8×Same (1,198 / 1,198)192-1.21%0.68% / 3.8%0.07 / 0.185,578 → 5,5951,405 → 1,405
H01 EM512×512×2561,6831.19 s (1.19, 1.22, 1.18)54.0 s45.3×Same (1,683 / 1,683)308-1.10%0.61% / 4.2%0.09 / 0.204,840 → 4,8421,788 → 1,788
FAFB FFN1 EM512×512×2565,3641.18 s (1.19, 1.18, 1.18)37.0 s31.2×Same (5,364 / 5,364)1,707-1.55%0.52% / 5.0%0.07 / 0.1916,411 → 16,4405,476 → 5,476
FIB-25 EM512×512×5122,7672.38 s (2.44, 2.38, 2.38)131.6 s55.3×Same (2,767 / 2,767)1,412-0.77%0.00% / 3.1%0.04 / 0.147,329 → 7,3162,925 → 2,925
Hemibrain EM512×512×5123,6242.69 s (2.74, 2.69, 2.69)331.1 s123.0×Same (3,624 / 3,624)1,298-1.41%0.60% / 5.3%0.09 / 0.2813,318 → 13,2873,808 → 3,808
Scroll fibres 16 µm X-ray CT512×512×51230.63 s (0.64, 0.63, 0.63)26.7 s42.1×Same (3 / 3)0-7.93%7.62% / 8.2%0.24 / 0.316,622 → 6,6232,022 → 2,022
Scroll fibres 8 µm X-ray CT512×512×51230.57 s (0.57, 0.57, 0.57)5.6 s9.9×Same (3 / 3)0-5.26%5.38% / 5.6%0.17 / 0.224,313 → 4,3181,825 → 1,825
Kimimaro benchmark EM512×512×5121,6673.08 s (3.14, 3.08, 3.03)413.4 s134.3×Same (1,667 / 1,667)272-1.45%0.96% / 4.3%0.09 / 0.199,342 → 9,3512,124 → 2,125

Length Δ comes mostly from how equal-cost routes are broken: Brook takes straight axial steps where Kimimaro’s order alternates between adjacent slices, so routes of equal cost, with nearly the same number of points, measure shorter. Thin fibres with cubic voxels have the most such ties.

Kimimaro benchmark volume, 512³

Kimimaro’s 512³ benchmark cutout of mouse visual cortex and a 256³ crop, against Kimimaro with 32 workers on all 32 threads and with one worker pinned to one core. Medians of three calls after one warmup. Its branching-correction case is the Kimimaro benchmark row above, timed separately.

Volume Branching correctionBrookKimimaro 32 workersKimimaro 1 coreSpeedup vs 32 workersSpeedup vs 1 core
256³ On0.396 s24.5 s28.0 s61.7×70.6×
256³ Off0.397 s22.4 s25.7 s56.5×64.8×
512³ On3.115 s415.1 s858.8 s133.3×275.7×
512³ Off2.674 s108.4 s232.8 s40.5×87.1×
Skeletons of the 512-cubed volume from Brook and from Kimimaro, drawn with the same camera and colours.
The 512³ volume, skeletonized by each tool. The straight lines at the lower right join the branches of one cell to the centre of its cell body.
Peak memory during one call on the 512³ volume with branching correction, sampled about every 0.1 s (host: proportional set size).
Main memoryGPU memory
Brook · NVIDIA RTX 40900.7 GiB5.6 GiB
Kimimaro · 32 workers12.7 GiBnot used
Kimimaro · 1 core3.6 GiBnot used

Batches

25 volumes of 1 × 1972 × 2024 voxels, median of three runs each: 4.60 s as separate calls, 1.30 s as one batch, identical results. The batch uses more GPU memory (+3,986 MB against +845 MB).

Cite

@software{brook,
  author = {Angelotti, Giorgio},
  title = {Brook: TEASAR skeletonization on NVIDIA GPUs},
  version = {0.1.0},
  year = {2026},
  license = {GPL-3.0-only},
  url = {https://github.com/giorgioangel/brook},
}

Brook implements the TEASAR algorithm [1] as Kimimaro does [2]:

  1. M. Sato, I. Bitter, M. A. Bender, A. E. Kaufman and M. Nakajima. “TEASAR: tree-structure extraction algorithm for accurate and robust skeletons”. In Proceedings of the Eighth Pacific Conference on Computer Graphics and Applications, Hong Kong, pp. 281–449. IEEE Computer Society, 2000. doi:10.1109/PCCGA.2000.883951
  2. W. Silversmith, J. A. Bae, P. H. Li and A. M. Wilson. Kimimaro: Skeletonize densely labeled 3D image segmentations, version 3.0.0. Zenodo, 2021. doi:10.5281/zenodo.5539913