Back in the machine room
Our cluster is working. Well, almost. The job queueing system is broken, the web interface crashes browsers, MPICH leaves stale shared memory allocations lying around, the filesystem is performing way under spec, several nodes have begun showing the symptoms of a failing hard drive, one of the cooling fans on the fileserver has started to rattle, the root node idles at a load average of about two (over two processors) and I just crashed the whole thing when I tried to run an actual job. I mean crash crash, no ping no nothing.
Good to see that all this time in the server room is paying off. I also have to admit that this is still better than things were a month ago, and it's better by a lot. You see, I'm not complaining when I write this post. I'm merely pointing out the irony in my current happiness.


0 Comments:
Post a Comment
<< Home