Saturday, 22 December 2007

Running Reliable Grid Services?

Well, Since it was Friday Night, just before Xmas, I thought I'd do what ervery sad geek would do, and fire off a batch of grid jobs to last over the shutdown.

However, the voms servers at CERN had other plans:

Contacting lcg-voms.cern.ch:15004 [/DC=ch/DC=cern/OU=computers/CN=lcg-voms.cern.ch] "dteam" Failed
Error: dteam: User unknown to this VO.
Trying next server for dteam.
Creating temporary proxy ...................................... Done
Contacting voms.cern.ch:15004 [/DC=ch/DC=cern/OU=computers/CN=voms.cern.ch] "dteam" Failed
Error: Could not connect to socket. Connection refused
None of the contacted servers for dteam were capable of returning a valid AC for the user.


And I'm not the only one - This morning I saw GGUS tickets in from Atlas, CMS and LHCb with the same problem. Ho Hum. Merry festive season to you too.

Wednesday, 28 November 2007

-x

Been debugging apparently slow transfers using FTS (now a 2.0 application....) and discovered a slight change in the UK setup.

Previously I just set the total no if files per channel, but I noticed that even with that set high, my dteam transfers were still only popping one file at a time off the list.

Turns out that the VO shares now have hard coded limits in which are only visible with the -x flag

ie:
glite-transfer-channel-list -x STAR-UKISCOTGRIDECDF
Channel: STAR-UKISCOTGRIDECDF
Between: * and UKI-SCOTGRID-ECDF
State: Active
Contact: lcg-support@gridpp.rl.ac.uk
Bandwidth: 0
Nominal throughput: 0
Number of files: 9, streams: 1
TCP buffer size: default
Message: no reason was provided
Last modification by: /C=UK/O=eScience/OU=CLRC/L=RAL/CN=lcgfts.gridpp.rl.ac.uk/Email=tier1a-certificate@gridpp.rl.ac.uk
Last modification time: 2007-11-28 15:10:39
Number of VO shares: 6
VO 'alice' share is: 1 and is limited to 1 transfers
VO 'lhcb' share is: 1 and is limited to 1 transfers
VO 'dteam' share is: 5 and is limited to 5 transfers
VO 'atlas' share is: 1 and is limited to 1 transfers
VO 'cms' share is: 0 and is not limited
VO 'ops' share is: 1 and is limited to 1 transfers

Wednesday, 7 November 2007

Virtualisation hassles

Got a shiny new Dell Optiplex 740 (AMD 64 X2 goodness) that I wanted to run request-tracker in a VM on. Installed ubuntu gutsy fine, Installed Xen fine, just wouldn't correctly boot any domU's to completion. I even tried out KVM seeing as it's new hardware that has the 'svm' flags in /proc/cpuinfo. That failed miserably at the modprobe kvm_amd stage with kvm: disabled by bios. Most odd as there wasn't a bios option for v12n[1]. However I (after much hassle finding a bootable flopy) upgraded from 1.1.3 to 1.2.2 BIOS and lo - xen 'Just Works'

it's actually scarily easy to do - "xen-create-image --hostname whatever" and you're off :-)

More postings as I get to grips with this. Initial impressions are good.

Monday, 22 October 2007

cfengine gotcha

Not strictly gridpp, but posting this here may save an administrator some sanity...

I've just spent the day battling against ubuntu on a lab cluster. The server had successfully upgraded from Feisty (7.04) to Gutsy (7.10) and cfengine refused to work. At all. It was only a minor upgrade in cfengine too - 2.1.20 to 2.1.22.

cfagent -qv -d2 hung - even with debugging on cfservd nothing obvious. Finally got a tip on IRC (freenode.net #cfengine) that it's likely to be a Berkeley DB upgrade thats the problem. Blew away all the .db files from state/ (and the __something.db that I found lying around in the cfengine dir too). tada! much stress reduced.

Tuesday, 9 October 2007

Feed Me!

At todays dteam meeting, there was a discussion about informing the PMB of any significant changes to the infrastructure. As the standard SOP is to use the EGEE Broadcast tool on the CIC portal, it'd be nice if there was say an RSS feed that we could parse and extract UKI significant info into say the planet feed. Anyone care to add comments?

Sunday, 30 September 2007

RAL Network performance

Prior to running any transfer tests, I normally check the Gridmon plots for that site, the relevant T2 and RAL T1. Using the (non-default) options of 'Metric data on same graph" and "New graph on new test dest" (and "new test src"), you generally get a good feel for the sites capacity, and any 'slow' links associated with it. Sadly recently the Tier1 seems to be performing dreadfully WRT the other sites.
example to Glasgow from Durham, Edinburgh and RAL:

Wednesday, 11 July 2007

Q3 Transfer Tests

OK, new quarter and time to get down to some testing as I'm well overdue. First up - Lancaster. Apart from some user error at this end (typo in script) it went pretty smoothly and can cope with 25 files in flight happily:

ganglia plot from Glasgow end:

showing 10,15,20,25 file setting in transfer channel

I've also started the Oxford tests, but it seems to be much less happy - When transferring from RAL-T2, even at low nos of files (5) the CPU load on t2se01 seems awfy awfy high.


Hmm. Have mailed Pete, but something doesn't look happy...