Showing posts with label AIX. Show all posts
Showing posts with label AIX. Show all posts

Wednesday, November 4, 2009

How I learned to Stop Worrying and Love the Bomb istat

I was recently tasked with organizing 137k+ small .jpg files into a folder structure based on year and quarter, why? users opening this directory with a ftp client complained that it took a "long time" to get a directory listing... apparently 15 - 20 minutes each time they opened the directory, honestly if a program didn't return anything in 15 minutes I would probably kill it and blame the server!

I really didn't think too much of the problem, in my head I though "i'll just use 'find' and 'stat'", which would have worked perfectly EXCEPT that I had to do this on a AIX 4.3 server and mounting the filesystem remotely was not an option.

A few problems with AIX 4.3 - no 'stat' command, in AIX 5.x you can install the coreutils rpm from the AIX toolbox to overcome this problem but you are up the creek without a paddle in 4.3! Also 'find' doesn't have all of the options you would usually have available on a newer version of linux - this was an issue in my case since I had to put the files into subdirectories (example: /basedirectory/2008/Q3) which meant that when searching for files to process in the basedirectory I did not want to descend into the yearly and quarterly subdirectories, easy with the -maxdepth option - which is not available in 4.3.

I ended up getting around the lack of -maxdepth in the find command by using the -prune option to remove subdirectories from processing, because the basedirectoy did not contain any directories except 200{8,9}/Q{1..4} this task was simplified even further by providing a common directory 'Q*'.

The lack of the 'stat' command had me banging my head against 'ls' for a day or so... The problem I have with 'ls' is in trying to get the year from 'ls -l', it works great for files older then 180 days but files less then 180 days are listed with the file modification timestamp in place of the year. I toyed with awk'ing the year/modifaction time column and checking if the value was an integer, which does work but you run into issues if your script is running within 180 days of the end of the year and examining files from the previous year, all the files will have timestamps which would cause you to examine the value of the current month vs. the month of the file being examined to determine the correct year - logic that I was uninterested in writing out.

Enter in my new most loved command in AIX: 'istat'
I was lucky enough to find a post mentioning 'istat' which "displays the i-node information for a particular file". 'istat' is simliar to the linux 'stat' command although it does not allow you modify the output using command line switches - nothing a little grep and awk won't fix! What 'istat' does do is handily format data about file creation, modification and access in an unambiguous matter - dates are always shown in the same format, unlike 'ls -l'. Without this tool I was writing a longer and longer script to deal with corner cases dealing with files modified 180 days ago and files modified around the last 3 months of the year - with 'istat' I was able to make my script much simpler and rely on the computer to hand me information in an consistent format.

I would be surprised if anyone has to solve this same problem but I will post the script anyways, as a warning this script is slow - 'istat' is not a tool for performance! Also working 'xargs' into the mix would make a more elegant solution in-place of 'find' and 'cat'.

In the following script I have disabled the actual move command - this will only print what would happen! uncomment the line beginning with 'mv' and it will move files.


#!/usr/bin/ksh
#
# organize files ending in $fileextension in $basedir
# by moving them into subdirectories $basedir/$year/$quarter
#


# backdate variable controls how many days old a file must be before
# it is considered for processing, 92 days is approx 3 months
# if you don't believe me ask google "3 months in days"
backdate=92

fileext=YOUR_FILE_EXTENTION
outfile=/tmp/jpg_organizer.out
basedir=YOUR_BASE_DIRECTORY

errors=0

# function to calulate which quarter a month lives in
calculate_quarter() {
case $month in
Jan|Feb|Mar)
quarter="Q1"
;;
Apr|May|Jun)
quarter="Q2"
;;
Jul|Aug|Sep)
quarter="Q3"
;;
Oct|Nov|Dec)
quarter="Q4"
;;
esac
}

# rudimentary error checking
error_check() {
let errors="$errors + $?"
if [[ $errors -gt 0 ]]; then
echo "encountered an error, exiting"
exit $?
fi
}

# find files older then $backdate and move them into $basedir/$year/$quarter directories
find $basedir -name Q\* -prune -o -name \*$fileext -mtime +$backdate -type f -print > $outfile
error_check
for i in `cat $outfile` ; do
filename=$i
fileattrib=`istat $i | grep "Last modified:"`
month=`echo $fileattrib | awk '{print $4}'`
year=`echo $fileattrib | awk '{print $7}'`
calculate_quarter
if [[ ! -d $basedir/$year/$quarter ]]; then
mkdir -p $basedir/$year/$quarter
error_check
fi
echo "moving:$filename to $basedir/$year/$quarter/"
#mv $filename $basedir/$year/$quarter/
error_check
done

rm $outfile

exit 0

Monday, October 26, 2009

AIX syslogd and splunk (and more)

AIX is what I would call a 'batteries not-included' OS; the vanilla DVD install leaves you with a functioning system that has telnet (with root access) enabled, no OpenSSL/OpenSSH, korn shell without autocomplete (must be enables 'set -o vi'), no logging, etc...
Since I work around a lot of RedHat boxes I tend to modify the AIX servers to have a simlar setup to RHEL, here are some of the steps I take:

Install the following rpm's from the aix toolbox:
bash (add /usr/bin/bash to /etc/security/login.cfg)
curl
coreutils
less
lsof
python
rsync
sudo
unzip
wget

Install OpenSSL and OpenSSH

Change root home directory to /root and change shell to bash:
mkdir /root && chuser home=/root shell=/usr/bin/bash root
Modify prompt for all users:
# Set bash prompt to be much more linux like
if [[ "$TERM" == "xterm" ]];then
if [[ "$SHELL" == "/usr/bin/bash" || "$SHELL" == "/bin/bash" ]];then
if [[ "$UID" -eq 0 ]];then
PS1="\[\033]0;\u@\h:\w\007\][\[\033[31;1m\]\u\[\033[0m\]@\h \W]# "
else
PS1="\[\033]0;\u@\h:\w\007\][\u@\h \W]\$ "
fi
fi
fi
Change logging setup:
# Linux-ify the AIX logging setup and enable automagic rotation
# Everything but mail and auth to messages
*.info;mail.none;auth.none /var/log/messages rotate size 10m files 10 compress
# Auth to secure
auth.debug /var/log/secure rotate size 10m files 10 compress
# Mail to maillog
mail.debug /var/log/maillog rotate size 10m files 10 compress
# Emergency messages to all users
*.emerg *
*.info;mail.none @NETWORK_LOG_SERVER
Remove "Message forwarded from hostname:" from remote logging output:
chssys -s syslogd -a "-n" ; stopsrc -s syslogd ; startsrc -s syslogd
Run aixpert to enable a much higher level of security:
aixpert -l high

Friday, June 6, 2008

Rebinding TSM archives so they do not expire

I have some TSM archives on an AIX host with a short expiration period of 14 days that I needed to extend for an unknown amount of time, I thought this would be an easy task but it took me a minute to figure out exactly how to quickly and efficiently get this done. To stop expiration on an archive you have to use the 'set event type=hold' command in from 'dsmc', the documentation on the command is complete but sparse with few examples so I had play with it to understand it. The most important thing I learned was that you cannot use a '*' to specify all files in an archive having a specific description - but you can specify just the base directory (in TSM terms 'filespace_name' from the archives table) and append a '/' (example: '/directory/) and it will pick up all of the files in the archive under the base directory.

First build the file list, this can be done easily a sql select statement from within dsmadmc :
select distinct(filespace_name) from archives where node_name='node name' and description='archive description here' > outfile
Then I needed to add a '/' character to the end of each line from the command line using awk:
# cat outfile | awk '{ print $1"/"}' > filelist.out

Alternately you could have selected filespace_name and hl_name and used awk to print both columns without a space between, either way the results should be the same...
Now I had a file that looked similar to this:
/aaaa/
/bbbb/
/cccc/
/db/abcd/
/db/efgh/
/db/ijkl/
/db/mnop/
/db/qrst/
/db/uvwx/
/db/yz/
/eeee/
/ffff/

each entry is a seperate mount point for a filesystem which is why /db/ would not have worked correctly.

Load this file into dsmc using the set event command:
set event -type=hold -filelist=filelist.out -description="Some unique description here"

output like:
....
ANS1899I ***** Examined 35,000 files *****
ANS1899I ***** Examined 36,000 files *****
ANS1899I ***** Examined 37,000 files *****
ANS1899I ***** Examined 38,000 files *****
ANS1899I ***** Examined 39,000 files *****
....

Total number of objects archived: 0
Total number of objects failed: 0
Total number of objects rebound: 82781
Total number of bytes transferred: 0 B
Data transfer time: 0.00 sec
Network data transfer rate: 0.00 KB/sec
Aggregate data transfer rate: 0.00 KB/sec
Objects compressed by: 0%
Elapsed processing time: 00:04:26
tsm>

After the files were rebound I jumped onto dsmadmc to make sure I had gotten every file rebound, I did this by selecting the number of objects from the archive and comparing it to the number of objects rebound:
select count(*) as "# of archived objects" from archives where description='Some unique description here' and node_name='node_name'

# of archived objects
---------------------
82781

Looks good!
This was by far the fastest way to rebind all these objects. I also tried selecting each individual object from the tsm database and loading that as the filelist into dsmc (a file with 82781 lines), after about 10 hours of processing dsmc core dumped. My other thought was to use the filelist containing every object and running it in a for loop so each object was processed individually - this would have worked but the time for completion would have been much longer.