AuthorsW. L. Guay, S. Reinemo, O. Lysne, T. Skeie, B. D. Johnsen, and L. Holen
EditorsX. Gu, and X. Ma
TitleHost Side Dynamic Reconfiguration With InfiniBand
AfilliationCommunication Systems, Communication Systems
StatusPublished
Publication TypeProceedings, refereed
Year of Publication2010
Conference Name2010 IEEE International Conference on Cluster Computing
Pagination126-135
PublisherIEEE Computer Society
ISBN Number978-0-7695-4220-1
Abstract

Rerouting around faulty components and migration of jobs both require reconfiguration of data structures in the Queue Pairs residing in the hosts on an InfiniBand cluster. In this paper we report an implementation of dynamic reconfiguration of such host side data-structures. Our implementation preserves the Queue Pairs, and lets the application run without being interrupted. With this implementation, we demonstrate a complete solution to fault tolerance in an InfiniBand network, where dynamic network reconfiguration to a topology-agnostic routing function is used to avoid malfunctioning components. This solution is in principle able to let applications run uninterruptedly on the cluster, as long as the topology is physically connected. Through measurements on our test-cluster we show that the increased cost of our method in setup latency is negligible, and that there is only a minor reduction in throughput during reconfiguration.

Citation KeySimula.netsys.12