Bramble.jl
Authors/Creators
Description
Task-based multithreading and parallel reductions (milestone v3.13.0): CpuThreaded now threads inner products, warmed assembly refills, the stencil and vector-calculus operators and broadcast assignment, and CpuPolyester batches the same paths, all without new dependencies and with results equal to CpuSerial bit for bit (reductions to rounding). This minor release has one breaking change, listed first.
Breaking
VectorElement integer multi-indices follow AbstractVector rules (#348)
- On a 2D or 3D grid function,
uₕ[i, j]no longer means grid point(i, j). It now followsAbstractVector's trailing-index semantics:uₕ[k, 1] == parent(uₕ)[k], and any other trailing index is aBoundsError. This is what generic code (sparse assembly,LinearAlgebra) already assumed. - Index the grid with a
CartesianIndex:uₕ[CartesianIndex(i, j)], or throughreshape(uₕ), which returns a grid-shaped view of the same memory.
Performance
Threaded inner products without Polyester (#301)
innerₕ,normₕand the other reductions underCpuThreadedsumThreads.nthreads()static chunks in a fixed order, so a result is the same on every run. Masked reductions walk only the set bits, and separable weights are chunked along the last axis without per-pointdivrem.- Measured crossover against
CpuSerialon an Apple M2 at--threads=4: 100,000–300,000 elements.
Threaded refills replay the recorded positions (#338)
- A warmed
assemble!underParallel()orCpuPolyesterwrites each entry to thenzvalposition recorded on the first assembly, instead of searching for it;assemble_add!does too. At 1M DOFs a threaded refill now takes 0.57–0.68 of the serial refill's time in 1D/2D/3D, where it used to be slower than serial in 2D and 3D.
Threaded and batched operators (#356)
- The difference and average stencil engines (
D₋ₓ!and its families,Mₓ!, jumps) and the vector-calculus operators (∇,div,curl,εfamilies) band along the last axis underCpuThreadedandCpuPolyester. Crossover againstCpuSerialforD₋ₓ!: about 300,000 DOFs for Threads, 1,000–10,000 for Polyester.
Threaded broadcast assignment (#357)
uₕ .= a .* vₕ .+ wₕinto aVectorElementruns Base's own broadcast loop in bands underCpuThreadedandCpuPolyester, aliasing handled as Base does.
Nested calls fall back to serial
- A
Parallel()operation called from inside your ownThreads.@threadsloop, or from a task racing another one, runs its chunks serially instead of throwing Base's "@threads :staticcannot be used concurrently or nested".
Fixes
- Every
Threads.@threadsloop in the assembly sweeps uses the:staticscheduler (#342).
Documentation and benchmarks
benchmark/policy_crossover.jlreports, on your own machine, whereCpuThreadedandCpuPolyesterstart beatingCpuSerialfor each workload in 1D, 2D and 3D; the backend tutorial points to it.- The
CpuThreadedandCpuPolyesterdocstrings quote the measured crossovers; the internals page describes the threaded replay.
Full changelog: https://github.com/gpena/Bramble.jl/compare/v3.12.0...v3.13.0
Notes
Files
gpena/Bramble.jl-v3.13.0.zip
Files
(1.1 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:e7a9aa6759c4a7c070c0cdb279fb1cf2
|
1.1 MB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/gpena/Bramble.jl/tree/v3.13.0 (URL)
Software
- Repository URL
- https://github.com/gpena/Bramble.jl