UT Opens Search for Dongarra Professor and ICL Director

The search is on for ICL’s next director. UT Knoxville invites applications and nominations for the endowed Dongarra Professorship in High Performance Computing, with the opportunity to lead ICL. The appointment is at the associate or full professor level in the Min H. Kao Department of Electrical Engineering and Computer Science, beginning in fall 2027.
The search seeks researchers with a record of external funding, teaching and mentorship, and academic leadership. Areas of interest span numerical algorithms and scientific software, parallel and distributed computing, and high-performance AI, among other fields. Researchers whose work connects algorithms, software, and evolving computing architectures are especially encouraged to apply.
Applications submitted by October 26, 2026, will receive full consideration; review will continue until the position is filled. See the Interfolio posting for full position description and application instructions.
ICL Earns Fifth R&D 100 Award for Fault Tolerance Contributions in Sandia’s FENIX

ICL has earned its fifth R&D 100 Award for its contributions to FENIX, an open-source software library that helps supercomputing applications recover from hardware and process failures. The work includes contributions from George Bosilca and Aurelien Bouteiller, whose research with collaborators at Sandia National Laboratories and partner institutions has helped make fault recovery easier to incorporate into scientific software.
ICL’s contribution is based in User-Level Fault Mitigation (ULFM), technology that lets programs using the Message Passing Interface (MPI) respond to failed processes and repair communication among those that survive. FENIX builds on ULFM to replace failed processes with spares and return execution to a point where the application can recover. This allows a calculation to continue within its existing allocation of computing resources, avoiding a full shutdown and another wait in the queue.
Bosilca and Bouteiller co-authored the 2022 IEEE CLUSTER paper “Integrating process, control-flow, and data resiliency layers using a hybrid Fenix/Kokkos approach,” which explains how process recovery fits into a complete recovery system. The team combined FENIX with Kokkos Resilience, which manages where execution resumes and identifies data to save, and VeloC, which stores and restores checkpoints. By making these tools work together, researchers could adapt recovery to an application’s needs while limiting changes to its code. Sandia’s account of a joint Sandia–UT demonstration describes how FENIX and ULFM support local recovery under repeated failures, allowing unaffected nodes to avoid repeating their work.
ICL’s earlier R&D 100 Awards In Brief
FENIX joins four earlier ICL award winners, with recognition spanning more than three decades. Those projects helped researchers combine computing resources, run numerical software efficiently, and understand application performance.
| 1994 | Parallel Virtual Machine (PVM) enabled researchers to combine different types of networked computers into a single parallel computing resource. It allowed scientific applications to divide work across machines and exchange data despite differences in their hardware. |
| 1999 | Automatically Tuned Linear Algebra Software (ATLAS) used empirical testing to tune linear algebra routines for the hardware on which they ran. This automated approach helped deliver efficient numerical computation across different systems. |
| 1999 | NetSolve gave scientists access to remote computing resources and numerical software through familiar desktop environments. Its client–agent–server design connected local problem-solving tools with the computing power available elsewhere on a network. |
| 2002 | The Performance Application Programming Interface (PAPI) provided a consistent way to collect hardware performance counter data across processors. Still developed at ICL, PAPI now supports performance measurement across CPUs, GPUs, and other hardware and software components. |
Dongarra Opens the Fall 2026 Lunch Talk Series
Jack Dongarra opened the Fall 2026 ICL Friday Lunch Talk series on September 4, 2026, with a talk titled “High Performance Computing in Transition.” Several ICL alumni came back for the occasion: Azzam Haidar (NVIDIA), Sam Crawford and John Batson (both Oak Ridge National Laboratory), and Anthony Danalis (AMD).
Dongarra began with the June 2026 TOP500 list, where five systems now exceed an exaflop and China’s LineShine has taken the top spot from El Capitan. He then turned to how generative AI has reshaped computing in capital, energy, and market terms, and to the hybrid approach in which physics-based models are paired with AI surrogate models that must be verified like any other scientific claim. He closed with co-design, the shift toward reduced-precision hardware, and the case that algorithms and software will keep following the hardware because, to borrow the title of a 2020 Science paper that inverts Feynman’s famous lecture, there is still “plenty of room at the top.”
Five more talks are on the calendar this semester, with speakers from ICL, UT’s Department of Applied Engineering, and King Abdullah University of Science and Technology (KAUST).
| Lunch Talks Fall 2026 | ||
|---|---|---|
| Oct2Fri | ![]() |
Piotr LuszczekInnovative Computing Laboratory |
| Oct9Fri | ![]() |
Brian LaRoseApplied Engineering, University of Tennessee, Knoxville |
| Nov6Fri | ![]() |
Daniel BarryInnovative Computing Laboratory |
| Nov13Fri | ![]() |
Hatem LtaiefKAUST |
| Nov25Wed | ![]() |
Hartwig AnztInnovative Computing Laboratory |
| Dec10Thu | ![]() |
Rabab AlomairyKAUST |
Alumni, collaborators, and friends of ICL are always welcome at the Friday Lunch Talks, in person or online. If you would like to hear about each talk as it comes up, just send a note to Dayle Flanigan to be added to the announcement/RSVP list.
Beams Presents heFFTe at MVAPICH User Group Conference

Natalie Beams presents “Performance and Scalability of 3D FFT using heFFTe and MVAPICH” at MUG’26. Video: MVAPICH.
Natalie Beams presented “Performance and Scalability of 3D FFT using heFFTe and MVAPICH” at the 14th Annual MVAPICH User Group (MUG) Conference, held August 17–19, 2026, in Columbus, Ohio. The Ohio Supercomputer Center hosted the meeting together with the Network-Based Computing Research Group at The Ohio State University, which develops the MVAPICH family of MPI libraries. Beams gave the talk virtually on Wednesday, August 19, with Ahmad Abdelfattah as co-author.
The fast Fourier transform (FFT) is a core step in signal processing, molecular dynamics, particle simulations, and many other applications, so these codes need a fast three-dimensional FFT on GPU-accelerated systems. The talk gave an overview of heFFTe (Highly Efficient FFT for Exascale), ICL’s open-source library for distributed 3D FFTs on GPUs, and showed how it draws on the communication optimizations in MVAPICH. Because the FFT is memory-bound, network communication often dominates its run time, which makes heFFTe a realistic and demanding benchmark for an MPI library. Beams presented results obtained with MVAPICH-Plus on large systems with both NVIDIA and AMD GPUs.
A recording of the talk is available on the MVAPICH YouTube channel.
Conference Reports
ICL Presents PAPI Updates at Scalable Tools Workshop
Daniel Barry and Treece Burgess represented ICL at the 19th Scalable Tools Workshop, held July 26–30, 2026, at Morgridge Hall on the University of Wisconsin–Madison campus. ICL alum Anthony Danalis (AMD) also attended. The workshop paired presentations on tools for high-performance computing with working groups devoted to technical problems.
All three presented on Monday, July 27. Burgess discussed support for emerging AMD and NVIDIA architectures in the Performance Application Programming Interface (PAPI), including the challenges of keeping performance measurement tools current with new hardware and vendor interfaces. Barry covered PAPI support for AI hardware, code annotations, and program counter sampling, which helps developers see where a program spends its execution time.
Danalis reviewed developments in AMD’s ROCm performance analysis tools, including new interfaces for developers building on Rocprofiler-SDK. He also described an experimental profiling skill and tool library intended to help developers and AI agents collect and interpret performance data within their existing workflows.
Recent Advances in PAPI Presented at ADAC PSI Seminar

Heike Jagode’s ADAC PSI Seminar talk covered PAPI support for CPUs, GPUs and accelerators, memory and I/O, interconnects, power and energy, and software-defined events.
Heike Jagode presented “Recent Advances in PAPI for Performance and Power Monitoring on Heterogeneous Systems” at the Accelerated Data Analytics and Computing Institute (ADAC) PSI Seminar, organized by Oak Ridge National Laboratory (ORNL) and RIKEN for the Asian scientific computing community. The presentation covered recent developments in PAPI (Performance Application Programming Interface) for performance, power, and energy monitoring on heterogeneous systems, including CPUs, GPUs, and accelerators from AMD, NVIDIA, and Intel. The talk also described PAPI’s growing support for monitoring memory, interconnects, I/O, and software-defined events, all through a single interface for analyzing complex computing systems.
SPADE Project Highlights Presented at NSF PI Meeting
The SPADE project (Scalable Performance and Accuracy analysis for Distributed and Extreme-scale systems), led by Heike Jagode in collaboration with Shirley Moore and Christoph Lauter (University of Texas at El Paso) and Vince Weaver (University of Maine), took part in the annual NSF CSSI/CyberTraining/SCIPE PI Meeting, held Sept. 13–15, 2026, in Rockville, Maryland. This year, Weaver represented the SPADE team and presented recent developments in PAPI (Performance Application Programming Interface).
SPADE focuses on advancing performance monitoring, optimization, and analysis for distributed and extreme-scale computing systems. In the project’s third year, the team expanded PAPI support for new HPC processors, GPUs, and AI accelerators. This work includes native AMD ROCprofiler-SDK support for modern AMD GPUs and APUs, expanded AMD-SMI and NVIDIA CUDA monitoring, PAPI’s first native support for Intel Gaudi AI accelerators, and portable GPU performance metrics (PAPI Presets) that make performance analysis more consistent across vendors and GPU generations.
The team also improved support for modern CPUs through Intel Topdown Metrics, which help identify performance bottlenecks within complex CPU pipelines. Intel support is now part of PAPI, and work is underway to bring similar capabilities to AMD, ARM, and RISC-V architectures.
Interview
Jeff Larkin

Tell us a little about your background and what years you were at ICL.
I was practically born with a keyboard in my lap and can’t even remember a time before I learned to program, so studying CS just felt like a preordained path. I studied CS at Furman University (yup, I was the guy in purple at the UT season opener) and fully expected I’d be joining the ranks of dotCom web programmers. My undergraduate advisor had a grant to do some scientific computing research for NASA and when I expressed interest in Linux he handed me the original Beowulf paper and asked if I wanted to build something like that for him. A year later I presented that work at an ACM conference in Gatlinburg and was later invited to come visit UT’s CS department. I was initially hired as a TA, but after 1 semester I reached out to Jack and became an RA at ICL. I was at ICL from late 2002 until mid 2005.
What prompted you to join ICL back then and what did you work on during your time at ICL?
I’d gotten a taste for both HPC and sysadmining while doing my undergraduate research, so when I had the opportunity to either join labstaff or ICL I had to decide which path seemed most exciting to me. The opportunity to do research, and for it to apply to science problems, really excited me, so I chose ICL. Initially I was brought on to keep the math software on the SInRG clusters, which were distributed among various departments on campus, up to date and available through NetSolve. I also began to contribute to the NetSolve/GridSolve efforts more directly and later used NetSolve as the jumping off point for my Master’s research project. I also started a project with one of my officemates to make installing software on the different clusters more automated.
Where are you working and what are you working on currently?
I’ve been at NVIDIA since 2013 and currently lead the solutions architecture team that supports major supercomputing sites across North America, particularly the DOE labs, NSF computing centers, and similar. This is a recent role change, but I’m really excited to be getting back to my roots in supercomputing and supporting so many sites doing really interesting (and not just supercomputing) things. In many ways, the classical supercomputing customers are being forced to reinvent themselves and so it’s a really exciting time to be engaging with them.
In what ways did working at ICL prepare you for what you do now, if at all?
First and foremost are the connections that I made during my time at ICL. Every time I’m at an industry conference I know that I’ll encounter someone who has a connection to ICL. It’s been fun keeping up with everyone, even when we just reconnect once a year at the supercomputing dinner, and meeting the new ICLers that have come through since I left. Not to downplay how much my ICL work prepared me technically, because it did, but the network of people I met during my time, whether they were staff, students, or visitors, has helped me immensely over the years.
What are some of your favorite memories from your days in the group?
ICL retreats were always a blast, they were one of my favorite parts of the year. In the theme of connections from the past question, Friday lunches were really impactful and memorable. It was so great to hear what other people were working on, reconnect with whoever you happened to get to sit next to that week, and meet people from outside the group (I got my first job from a visiting speaker). Waiting over the coffee pot that could never stay full. Lastly, whatever the game that Keith and Haihang kept in their office, it was always a good way to take a break to think a bit.
What are some of your interests/hobbies outside of work?
I love traveling with my family, we make the best use we can out of school breaks. We especially love going on cruises or to the Orlando theme parks. I volunteer both with my son’s scout troop and also my church’s student ministry. I’m also an avid banjoist and ukelele player. (I said “avid” not “good.”)
























