Jump to content

Limulus Software

From Cluster Documentation Project
Revision as of 11:12, 29 September 2026 by Deadline (talk | contribs) (added image)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

Limulus Software

The following is a list of basic cluster RPMS that will be included in the software stack. The base distribution will be Scientific Linux. An open up-to-date software stack is used for modern Limulus machines.

  1. Scientific Linux V6.3 - RHEL work alike distribution
  2. Warewulf Cluster Toolkit - Cluster provisioning and administration (V3.1)
  3. PDSH - Parallel Distributed Shell for collective administration
  4. Whatsup - A cluster node up/down detection utility
  5. Open Grid Scheduler - previously Sun Grid Engine Resource Scheduler
  6. Ganglia - Cluster Monitoring System
  7. GNU Compilers (gcc, g++, g77, gdb) - Standard GNU compiler suite
  8. Modules - Manages User Environments
  9. PVM - Parallel Virtual Machine (message passing middleware)
  10. MPICH2 - MPI Library (message passing middleware)
  11. OPEN-MPI - MPI Library (message passing middleware)
  12. Open-MX - Myrinet Express over Ethernet
  13. ATLAS - host tuned BLAS library
  14. OpenBLAS - optimized BLAS library (previously Goto library)
  15. FFTW - Optimized FFT library
  16. FFTPACK - FFT library
  17. LAPACK and BLAS - Linear Algebra library
  18. ScaLAPACK - Scalable Linear Algebra Package
  19. PetSc Scalable PDE solvers
  20. GNU GSL - GNU Scientific Library (over 1000 functions)
  21. PADB - Parallel Application Debugger Inspection Tool
  22. Julia - Easy To Use High Performance Parallel Scientific Language
  23. Userstat - a "top" like job queue/node monitoring application
  24. Beowulf Performance Suite - benchmark and testing suite
  25. relayset - power relay control utility
  26. ssmtp - mail forwarder for nodes

Automatic Power Control

One key design component of the new Limulus Case is software controlled power to the nodes. This feature will allow nodes to be powered-on only when needed. As an experiment, a simple script was written that monitors the Grid Engine queue. If there are jobs waiting in the queue, nodes are powered-on. Once the work is done (i.e. nothing more in the queue) the nodes are powered off). Note: Newer Limulus systems use the SLURM Scheduler which provides hooks to power nodes off and on similar to this Grid Engine example.

As an example, an 8 core job was run on the Norbert cluster (in the Limulus case). The head node has 4 cores and each worker node has 2 cores for a total of 10 cores. An 8 node job was submitted via Grid Engine with only the head node powered-on. The script noticed the job waiting in the queue and turned on a single worker node to give 6 cores total, which were still not enough. Another node was powered-on and the total cores reached 8 and the job started to run. After completion, the script noticed that there was nothing in the queue and shutdown the nodes.

To see how the number of nodes changes with the work in the queue, consider the Ganglia trace below. Note that the number of nodes in the cluster load graph (green line) changes from 1 to 2 then 3 then back down to 1. Similarly the number of CPUs (red line) raises to 8 then back to down to 4 (the original 4 cores in the head node). Similar changes can be seen in the system memory.

Cookies help us deliver our services. By using our services, you agree to our use of cookies.