Matrix Algebra on GPU and Multi-core Architectures. Dense linear algebra routines — LU, QR, Cholesky, eigensolvers — engineered for heterogeneous systems, from laptops to exascale machines.
MAGMA builds with CMake against CUDA 11+, ROCm 5+, or oneAPI. Prebuilt packages are available through Spack and conda-forge.
New in 2.9.0
CUDA 13 support and unified-memory batched routines. See the release notes.
git clone https://bitbucket.org/icl/magma.git
cd magma && mkdir build && cd build
cmake -DMAGMA_ENABLE_CUDA=ON \
-DCMAKE_INSTALL_PREFIX=/opt/magma ..
make -j && make install
| Version | Date | Highlights | |
|---|---|---|---|
| 2.9.0 Latest | May 2026 | CUDA 13, unified-memory batched routines, improved SYCL coverage | tar.gz · 9.5 MB |
| 2.8.0 | Nov 2025 | HIP backend parity, mixed-precision GMRES | tar.gz |
| 2.7.2 | Mar 2025 | Bug fixes, oneAPI experimental support | tar.gz |
Citing MAGMA in your work supports continued development.
Towards Dense Linear Algebra for Hybrid GPU Accelerated Manycore Systems
Parallel Computing, 36(5–6):232–240, 2010
Accelerating Mixed-Precision Iterative Refinement on Exascale GPUs
Proceedings of SC '25, Atlanta, GA, November 2025
Batched One-Sided Factorizations on Unified Memory Architectures
IEEE IPDPS, 2026
Stan Tomov
Principal Investigator
Azzam Haidar
Research Scientist
Piotr Luszczek
Research Scientist
Neil Lindquist
Graduate Researcher
Low-volume list — new releases and critical fixes only.
Join the mailing list