There is a newer version of the record available.

Published September 26, 2026 | Version v3.13.0

Bramble.jl

Authors/Creators

Description

Task-based multithreading and parallel reductions (milestone v3.13.0): CpuThreaded now threads inner products, warmed assembly refills, the stencil and vector-calculus operators and broadcast assignment, and CpuPolyester batches the same paths, all without new dependencies and with results equal to CpuSerial bit for bit (reductions to rounding). This minor release has one breaking change, listed first.

Breaking

VectorElement integer multi-indices follow AbstractVector rules (#348)

  • On a 2D or 3D grid function, uₕ[i, j] no longer means grid point (i, j). It now follows AbstractVector's trailing-index semantics: uₕ[k, 1] == parent(uₕ)[k], and any other trailing index is a BoundsError. This is what generic code (sparse assembly, LinearAlgebra) already assumed.
  • Index the grid with a CartesianIndex: uₕ[CartesianIndex(i, j)], or through reshape(uₕ), which returns a grid-shaped view of the same memory.

Performance

Threaded inner products without Polyester (#301)

  • innerₕ, normₕ and the other reductions under CpuThreaded sum Threads.nthreads() static chunks in a fixed order, so a result is the same on every run. Masked reductions walk only the set bits, and separable weights are chunked along the last axis without per-point divrem.
  • Measured crossover against CpuSerial on an Apple M2 at --threads=4: 100,000–300,000 elements.

Threaded refills replay the recorded positions (#338)

  • A warmed assemble! under Parallel() or CpuPolyester writes each entry to the nzval position recorded on the first assembly, instead of searching for it; assemble_add! does too. At 1M DOFs a threaded refill now takes 0.57–0.68 of the serial refill's time in 1D/2D/3D, where it used to be slower than serial in 2D and 3D.

Threaded and batched operators (#356)

  • The difference and average stencil engines (D₋ₓ! and its families, Mₓ!, jumps) and the vector-calculus operators (∇, div, curl, ε families) band along the last axis under CpuThreaded and CpuPolyester. Crossover against CpuSerial for D₋ₓ!: about 300,000 DOFs for Threads, 1,000–10,000 for Polyester.

Threaded broadcast assignment (#357)

  • uₕ .= a .* vₕ .+ wₕ into a VectorElement runs Base's own broadcast loop in bands under CpuThreaded and CpuPolyester, aliasing handled as Base does.

Nested calls fall back to serial

  • A Parallel() operation called from inside your own Threads.@threads loop, or from a task racing another one, runs its chunks serially instead of throwing Base's "@threads :static cannot be used concurrently or nested".

Fixes

  • Every Threads.@threads loop in the assembly sweeps uses the :static scheduler (#342).

Documentation and benchmarks

  • benchmark/policy_crossover.jl reports, on your own machine, where CpuThreaded and CpuPolyester start beating CpuSerial for each workload in 1D, 2D and 3D; the backend tutorial points to it.
  • The CpuThreaded and CpuPolyester docstrings quote the measured crossovers; the internals page describes the threaded replay.

Full changelog: https://github.com/gpena/Bramble.jl/compare/v3.12.0...v3.13.0

Notes

If you use this software, please cite it as below.

Files

gpena/Bramble.jl-v3.13.0.zip

Files (1.1 MB)

Name Size Download all
md5:e7a9aa6759c4a7c070c0cdb279fb1cf2
1.1 MB Preview Download

Additional details

Related works

Software