Batched QR and SVD algorithms on GPUs with applications in hierarchical matrix compression
| dc.contributor.author | Boukaram, Wajih Halim | |
| dc.contributor.author | Turkiyyah, George M. | |
| dc.contributor.author | Ltaief, Hatem | |
| dc.contributor.author | Keyes, David E. | |
| dc.contributor.department | Department of Computer Science | |
| dc.contributor.faculty | Faculty of Arts and Sciences (FAS) | |
| dc.contributor.institution | American University of Beirut | |
| dc.date.accessioned | 2025-01-24T11:22:56Z | |
| dc.date.available | 2025-01-24T11:22:56Z | |
| dc.date.issued | 2018 | |
| dc.description.abstract | We present high performance implementations of the QR and the singular value decomposition of a batch of small matrices hosted on the GPU with applications in the compression of hierarchical matrices. The one-sided Jacobi algorithm is used for its simplicity and inherent parallelism as a building block for the SVD of low rank blocks using randomized methods. We implement multiple kernels based on the level of the GPU memory hierarchy in which the matrices can reside and show substantial speedups against streamed cuSOLVER SVDs. The resulting batched routine is a key component of hierarchical matrix compression, opening up opportunities to perform H-matrix arithmetic efficiently on GPUs. © 2017 Elsevier B.V. | |
| dc.identifier.doi | https://doi.org/10.1016/j.parco.2017.09.001 | |
| dc.identifier.eid | 2-s2.0-85029565345 | |
| dc.identifier.uri | http://hdl.handle.net/10938/25567 | |
| dc.language.iso | en | |
| dc.publisher | Elsevier B.V. | |
| dc.relation.ispartof | Parallel Computing | |
| dc.source | Scopus | |
| dc.subject | Batched operations | |
| dc.subject | Compression | |
| dc.subject | Gpu | |
| dc.subject | Hierarchical | |
| dc.subject | Qr | |
| dc.subject | Svd | |
| dc.subject | Compaction | |
| dc.subject | Graphics processing unit | |
| dc.subject | Jacobian matrices | |
| dc.subject | Program processors | |
| dc.subject | Hierarchical matrices | |
| dc.subject | High performance implementations | |
| dc.subject | Inherent parallelism | |
| dc.subject | Multiple kernels | |
| dc.subject | One-sided jacobi algorithms | |
| dc.subject | Randomized method | |
| dc.subject | Singular value decomposition | |
| dc.title | Batched QR and SVD algorithms on GPUs with applications in hierarchical matrix compression | |
| dc.type | Article |
Files
Original bundle
1 - 1 of 1