use #if __GNUC__ < 4
#if __GNUC_MINOR__ < 3
A blog about the system and programming of *nix environment. It covers: * Scripts that I am familiar with include shell scripts, Perl, Python, Tcl/Tk, and PHP. * Utilities includes sed, grep, awk, cat, tac, ... * Kernel development and optimization, device driver * Networking * Toolchain, build environment
Wednesday, April 16, 2008
Tuesday, March 04, 2008
Linker script .lds
A script xxx.lds is used to instruct the linker, how to generate and link the final elf target. I recently solved a bootloader problem, by moving a part of (reginfo) into Data section. The bug in binutil overwrite my target elf file, with the reginfo, which shouldn't be loaded.
Some links about the linker.
http://www.gnu.org/software/binutils/manual/ld-2.9.1/html_mono/ld.html
http://www.embedded.com/2000/0002/0002feat2.htm
Some links about the linker.
http://www.gnu.org/software/binutils/manual/ld-2.9.1/html_mono/ld.html
http://www.embedded.com/2000/0002/0002feat2.htm
Tuesday, January 29, 2008
CGI internal server error
Haven't worked on CGI/perl for a while, recently I am reworking on my perl scripts for star5.ca website. The following items should be checked when you got the 'Internal Server Error" in Perl. GoDaddy has no server log enabled by default, so you have to be very careful in doing CGI/Perl programming.
1. File permission (your perl program should have 755 mode)
2. DOS ending (your perl program should be using unix line ending, instead of dos line ending. A lot system will complain of "file not found", because of the ending issue)
3. Path to perl (usual !/usr/bin/perl)
4. Library path (use "path to your local library"), to support local perl modules.
1. File permission (your perl program should have 755 mode)
2. DOS ending (your perl program should be using unix line ending, instead of dos line ending. A lot system will complain of "file not found", because of the ending issue)
3. Path to perl (usual !/usr/bin/perl)
4. Library path (use "path to your local library"), to support local perl modules.
Wednesday, January 16, 2008
gcc
I am working on toolchain upgrade recently, from gcc 3.4.6 to gcc 4.1.1.
Thing I learned during the upgrade.
1. libgcc contains compiler specific library, usually used for floating point computation. For example, 64 bit floating point has no support on the host system, the libgcc has to convert the computation into some native instructions.
You can use "gcc -v" to figure out the default library path "-L xxx" used by gcc, besides the path passed by your Makefile. "gcc -v" can also give you the exact command called by gcc, as gcc itself is an umbrella program, which calls "cc1", "collect2" etc.
2. -ffreestanding, flag may be used for kernel compilation. It implies that standard library may not exist and the program startup may not necessarily be at "main". To use gcc 4.1.1 to compile linux kernel 2.6.17, we need to add -ffreestanding in our makerules.
3. Optimization, gcc optimize the code differently in each version. Some functions were optimized away in kernel, but called in bootloader. I have to create a dummy function in bootloader to avoid linking problem.
4. -std=gnu99. I've met several preprocessing error when I was building gcc in buildroot. cpp was complaining about the unknown labels in assembly code. It turns out the the cpp using gnu standard is not 100% compatible with the iso c standard. By removing the -std=gnu99 flag, I can get gcc compiled.
5. -sysroot. This option will enable you to use a different set of "/include", "/lib" directories in a different root. I think that it may be useful in cross-compiling environment.
Thing I learned during the upgrade.
1. libgcc contains compiler specific library, usually used for floating point computation. For example, 64 bit floating point has no support on the host system, the libgcc has to convert the computation into some native instructions.
You can use "gcc -v" to figure out the default library path "-L xxx" used by gcc, besides the path passed by your Makefile. "gcc -v" can also give you the exact command called by gcc, as gcc itself is an umbrella program, which calls "cc1", "collect2" etc.
2. -ffreestanding, flag may be used for kernel compilation. It implies that standard library may not exist and the program startup may not necessarily be at "main". To use gcc 4.1.1 to compile linux kernel 2.6.17, we need to add -ffreestanding in our makerules.
3. Optimization, gcc optimize the code differently in each version. Some functions were optimized away in kernel, but called in bootloader. I have to create a dummy function in bootloader to avoid linking problem.
4. -std=gnu99. I've met several preprocessing error when I was building gcc in buildroot. cpp was complaining about the unknown labels in assembly code. It turns out the the cpp using gnu standard is not 100% compatible with the iso c standard. By removing the -std=gnu99 flag, I can get gcc compiled.
5. -sysroot. This option will enable you to use a different set of "/include", "/lib" directories in a different root. I think that it may be useful in cross-compiling environment.
Wednesday, December 12, 2007
Kernel network tune-up
I closed two open tickets related to networking in Linux Kernel. So, I'd write it down in order not to forget.
* broadcast ping problem. After Linux kernel 2.6.14+, the kernel wont' respond to broadcast ping by default. You have to "echo 0 > /proc/sys/net/ipv4/ignore_broadcast_icmp" to enable response to broadcast ping. The default behavior has been changed.
* loopback ping > 32K payload. For small system with <= 16M memory, we use loopback UDP for IPC. But we can't receive packets with 32K+ payload. I wrote a test program to verify the problem. After one day's hack, I found that it is due to a limitation in internal buffer size of skbuff (the internal data structure in Linux network stack). "/proc/sys/net/core/rmem_max", "/proc/sys/net/core/rmem_default". We have to increase the limit on those two entries to allow loopback ping with 32K+ payload. There won't any problem with network ping, as the MTU limitation will eventually split the skbuff and you won't need a big skbuff. The loopback has no MTU limitation, thus it matters.
* broadcast ping problem. After Linux kernel 2.6.14+, the kernel wont' respond to broadcast ping by default. You have to "echo 0 > /proc/sys/net/ipv4/ignore_broadcast_icmp" to enable response to broadcast ping. The default behavior has been changed.
* loopback ping > 32K payload. For small system with <= 16M memory, we use loopback UDP for IPC. But we can't receive packets with 32K+ payload. I wrote a test program to verify the problem. After one day's hack, I found that it is due to a limitation in internal buffer size of skbuff (the internal data structure in Linux network stack). "/proc/sys/net/core/rmem_max", "/proc/sys/net/core/rmem_default". We have to increase the limit on those two entries to allow loopback ping with 32K+ payload. There won't any problem with network ping, as the MTU limitation will eventually split the skbuff and you won't need a big skbuff. The loopback has no MTU limitation, thus it matters.
Friday, November 09, 2007
Kernel programming and driver debugging
Don't forget to kill the klogd and syslogd daemons before you start to debug your driver or kernel codes. Otherwise, your printk will be buffered and you won't know the exact point of crash or hang.
"
killall klogd
killall syslogd
"
For interrupt handler, the first thing in your ISR is to disable or mask out the interrupt, otherwise, the interrupt may keep firing and your system may be locked up.
At the end of your ISR, you can re-enable and enable the mask of the interrupt.
To get a list of interrupt and interrupt handler,
"
cat /proc/interrupts
"
"
killall klogd
killall syslogd
"
For interrupt handler, the first thing in your ISR is to disable or mask out the interrupt, otherwise, the interrupt may keep firing and your system may be locked up.
At the end of your ISR, you can re-enable and enable the mask of the interrupt.
To get a list of interrupt and interrupt handler,
"
cat /proc/interrupts
"
Thursday, November 01, 2007
Tuesday, October 16, 2007
my .vimrc
" LEO's customization
" version 4.0
set nocompatible
" Broadcom coding style
set shiftwidth=3
set tabstop=3
set expandtab
" Nice search
set incsearch
set ignorecase
set smartcase
set ai
set backspace=2
"set backupdir=~/.backup
"set nobackup
" My shortcut key mapping
map gs :%s/
map <c-j> <c-w>j
map <c-k> <c-w>k
map <c-h> <c-w>h
map <c-l> <c-w>l
map <f2> @q
map <f3> @a
map <f10> :set paste<cr>
map <f11> <c-w>-
map <f12> <c-w>+
map :Q :qa!
" For nicer scroll, who knows
set showcmd
set sm
set ss=1
set siso=9
set so=3
" highlight search result
set hls
" syntax highlight
syntax on
colorscheme darkblue
"copy and paste betweeen different vim sessions
nmap <c-y> :!echo ""> ~/.vi_tmp<cr><cr>:w! ~/.vi_tmp<cr>
vmap <c-y> :w! ~/.vi_tmp<cr>
nmap <c-p> :r ~/.vi_tmp<cr>
vmap <c-p> c<esc>:r ~/.vi_tmp<cr>
nmap :Q :qa
"expand the directory with pwd of file under editing
nmap ,e :e <c-r>=expand("%:p:h") . "/" <cr>
nmap ,n :new <c-r>=expand("%:p:h") . "/" <cr>
" Remember the last edit position
set viminfo='10,\"100,:20,%,n~/.viminfo
au BufReadPost * if line("'\"") > 0|if line("'\"") <= line("$")|exe("norm '\"")|else|exe "norm $"|endif|endif
" Command abbreviation for spell checking
cab aspe :w<cr>:!aspell -e -x -c %<cr>:e<cr><cr>
highlight RedundantSpaces term=standout ctermbg=red guibg=red
match RedundantSpaces /\s\+$\| \+\ze\t/
set ruler
set autoindent
set smartindent
"set spell
"set statusline=%F%m%r%h%w\ [FORMAT=%{&ff}]\ [TYPE=%Y]\ [ASCII=\%03.3b]\ [HEX=\%02.2B]\ [POS=%04l,%04v][%p%%]\ [LEN=%L]
"set statusline=%F%m%r%h%w\ [[TYPE=%Y]\ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ [POS=%04l,%04v][%p%%]\ [LEN=%L]
"set laststatus=2
let perl_extended_vars=1
filetype plugin on " load filetype plugins
set visualbell t_vb=
let loaded_matchparen=1
" version 4.0
set nocompatible
" Broadcom coding style
set shiftwidth=3
set tabstop=3
set expandtab
" Nice search
set incsearch
set ignorecase
set smartcase
set ai
set backspace=2
"set backupdir=~/.backup
"set nobackup
" My shortcut key mapping
map gs :%s/
map <c-j> <c-w>j
map <c-k> <c-w>k
map <c-h> <c-w>h
map <c-l> <c-w>l
map <f2> @q
map <f3> @a
map <f10> :set paste<cr>
map <f11> <c-w>-
map <f12> <c-w>+
map :Q :qa!
" For nicer scroll, who knows
set showcmd
set sm
set ss=1
set siso=9
set so=3
" highlight search result
set hls
" syntax highlight
syntax on
colorscheme darkblue
"copy and paste betweeen different vim sessions
nmap <c-y> :!echo ""> ~/.vi_tmp<cr><cr>:w! ~/.vi_tmp<cr>
vmap <c-y> :w! ~/.vi_tmp<cr>
nmap <c-p> :r ~/.vi_tmp<cr>
vmap <c-p> c<esc>:r ~/.vi_tmp<cr>
nmap :Q :qa
"expand the directory with pwd of file under editing
nmap ,e :e <c-r>=expand("%:p:h") . "/" <cr>
nmap ,n :new <c-r>=expand("%:p:h") . "/" <cr>
" Remember the last edit position
set viminfo='10,\"100,:20,%,n~/.viminfo
au BufReadPost * if line("'\"") > 0|if line("'\"") <= line("$")|exe("norm '\"")|else|exe "norm $"|endif|endif
" Command abbreviation for spell checking
cab aspe :w<cr>:!aspell -e -x -c %<cr>:e<cr><cr>
highlight RedundantSpaces term=standout ctermbg=red guibg=red
match RedundantSpaces /\s\+$\| \+\ze\t/
set ruler
set autoindent
set smartindent
"set spell
"set statusline=%F%m%r%h%w\ [FORMAT=%{&ff}]\ [TYPE=%Y]\ [ASCII=\%03.3b]\ [HEX=\%02.2B]\ [POS=%04l,%04v][%p%%]\ [LEN=%L]
"set statusline=%F%m%r%h%w\ [[TYPE=%Y]\ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ \ [POS=%04l,%04v][%p%%]\ [LEN=%L]
"set laststatus=2
let perl_extended_vars=1
filetype plugin on " load filetype plugins
set visualbell t_vb=
let loaded_matchparen=1
Sunday, October 14, 2007
Working on my first driver
I have been working on the touchscreen driver for the past two weeks. In last Friday, it seems that my work had made a milestone. The user land touchscreen calibration and test programs works just fine using tslib.
The touchscreen driver I created conforms to the input layer and event interface of Linux kernel. I have to use "input_register_driver" and "input_register_device" to register my device and the driver. The "probe" function then called.
I have to work out the proc entries and cleanup the code, and write documentation in next week, which should be done without much trouble.
The touchscreen driver I created conforms to the input layer and event interface of Linux kernel. I have to use "input_register_driver" and "input_register_device" to register my device and the driver. The "probe" function then called.
I have to work out the proc entries and cleanup the code, and write documentation in next week, which should be done without much trouble.
Thursday, June 14, 2007
let's discuss Linux kernel programming here
As told by Dave today, the MIPS kernel stack dump doesn't dump the calling stack in order. It just dump all the seemly kernel functions in the stack. They could be functions called before the real BUG happened.
For example, the following call trace dump is not the real calling stack. The bug is in set_mctrl, which called spin_lock_irqsave, but the same lock was acquired by the caller of set_mctrl function already.
------------------
BUG: spinlock recursion on CPU#0, swapper/1
lock: 80302a84, .magic: dead4ead, .owner: swapper/1, .owner_cpu: 0
Call Trace:
[<801f03c8>] _raw_spin_lock+0x4c/0x154
[<802c2ef4>] _spin_lock_irqsave+0x40/0x58
[<80212bfc>] bcm1103serial_set_mctrl+0x70/0xb8
[<8012e934>] printk+0x1c/0x28
[<802123ac>] uart_add_one_port+0x290/0x354
[<80184320>] exact_match+0x0/0x8
[<80184328>] exact_lock+0x0/0x28
[<801ff9d4>] alloc_tty_driver+0x24/0x64
[<80212088>] uart_register_driver+0x194/0x1cc
[<80350000>] keypad_init+0x48/0x14c
[<80351020>] bcm1103serial_init+0x44/0x68
[<80100558>] init+0xc4/0x29c
[<80100558>] init+0xc4/0x29c
[<80109cd4>] kernel_thread_helper+0x10/0x18
[<80109cc4>] kernel_thread_helper+0x0/0x18
-------------------
Remember, the call trace dump of MIPS is not the exact calling trace, you have to find clue in the functions dumped.
For example, the following call trace dump is not the real calling stack. The bug is in set_mctrl, which called spin_lock_irqsave, but the same lock was acquired by the caller of set_mctrl function already.
------------------
BUG: spinlock recursion on CPU#0, swapper/1
lock: 80302a84, .magic: dead4ead, .owner: swapper/1, .owner_cpu: 0
Call Trace:
[<801f03c8>] _raw_spin_lock+0x4c/0x154
[<802c2ef4>] _spin_lock_irqsave+0x40/0x58
[<80212bfc>] bcm1103serial_set_mctrl+0x70/0xb8
[<8012e934>] printk+0x1c/0x28
[<802123ac>] uart_add_one_port+0x290/0x354
[<80184320>] exact_match+0x0/0x8
[<80184328>] exact_lock+0x0/0x28
[<801ff9d4>] alloc_tty_driver+0x24/0x64
[<80212088>] uart_register_driver+0x194/0x1cc
[<80350000>] keypad_init+0x48/0x14c
[<80351020>] bcm1103serial_init+0x44/0x68
[<80100558>] init+0xc4/0x29c
[<80100558>] init+0xc4/0x29c
[<80109cd4>] kernel_thread_helper+0x10/0x18
[<80109cc4>] kernel_thread_helper+0x0/0x18
-------------------
Remember, the call trace dump of MIPS is not the exact calling trace, you have to find clue in the functions dumped.
Tuesday, August 15, 2006
Sort complex dictionary in Python
Use lambda function and sorted function to sort a complex dictionary in Python
>>> a={"a":[1,"a"], "b":[2,"b"], "c":[0,"A"], "d":[-2, "z"]}
>>> a.items()
[('a', [1, 'a']), ('c', [0, 'A']), ('b', [2, 'b']), ('d', [-2, 'z'])]
>>> sorted(a.items(), lambda x, y : cmp(x[1][0], y[1][0]))
[('d', [-2, 'z']), ('c', [0, 'A']), ('a', [1, 'a']), ('b', [2, 'b'])]
>>>sorted(a.items(), lambda x, y : cmp(x[1][1], y[1][1]))
[('c', [0, 'A']), ('a', [1, 'a']), ('b', [2, 'b']), ('d', [-2, 'z'])]
It could be useful.
>>> a={"a":[1,"a"], "b":[2,"b"], "c":[0,"A"], "d":[-2, "z"]}
>>> a.items()
[('a', [1, 'a']), ('c', [0, 'A']), ('b', [2, 'b']), ('d', [-2, 'z'])]
>>> sorted(a.items(), lambda x, y : cmp(x[1][0], y[1][0]))
[('d', [-2, 'z']), ('c', [0, 'A']), ('a', [1, 'a']), ('b', [2, 'b'])]
>>>sorted(a.items(), lambda x, y : cmp(x[1][1], y[1][1]))
[('c', [0, 'A']), ('a', [1, 'a']), ('b', [2, 'b']), ('d', [-2, 'z'])]
It could be useful.
Thursday, March 23, 2006
A Python script to get files from FTP site easily
#!/usr/bin/python
from ftplib import FTP
import netrc
import re
import sys
directory = {"siteid":["hostname", "root directory of your ftp site"], "another_site":["another host", "another root directory"]}
if len(sys.argv) > 2:
host = sys.argv[1]
off_file = 2
else:
host = "hostname"
off_file = 1
if not directory.has_key(host):
print host, "is not currently supported"
sys.exit(1)
# we use .netrc to store userid and password info.
net = netrc.netrc()
(user, acct, passwd) = net.authenticators(directory[host][0])
def main(argv):
print "connecting ...", directory[host][0]
ftp = FTP(directory[host][0])
print "login'ing ...", directory[host][0]
ftp.login(user, passwd)
for i in range(len(argv)):
sys.stdout.write("getting ... " + argv[i])
p = re.compile('(.*)(/[^/]+)')
m = p.match(argv[i])
if m:
path = m.group(1)
filename = m.group(2)[1:]
ftp.cwd(directory[host][1] + path)
if len(path):
file = open(argv[i], 'wb')
else:
file = open(filename, 'wb')
ftp.retrbinary("RETR " + filename, file.write)
print " done"
file.close()
else:
print "no matching .."
ftp.close()
if __name__ == '__main__':
if (len(sys.argv) > 1):
main(sys.argv[off_file:])
else:
print "Usage: get [host] file_to_be_get\nHost: siteid(default), another_site\n"
from ftplib import FTP
import netrc
import re
import sys
directory = {"siteid":["hostname", "root directory of your ftp site"], "another_site":["another host", "another root directory"]}
if len(sys.argv) > 2:
host = sys.argv[1]
off_file = 2
else:
host = "hostname"
off_file = 1
if not directory.has_key(host):
print host, "is not currently supported"
sys.exit(1)
# we use .netrc to store userid and password info.
net = netrc.netrc()
(user, acct, passwd) = net.authenticators(directory[host][0])
def main(argv):
print "connecting ...", directory[host][0]
ftp = FTP(directory[host][0])
print "login'ing ...", directory[host][0]
ftp.login(user, passwd)
for i in range(len(argv)):
sys.stdout.write("getting ... " + argv[i])
p = re.compile('(.*)(/[^/]+)')
m = p.match(argv[i])
if m:
path = m.group(1)
filename = m.group(2)[1:]
ftp.cwd(directory[host][1] + path)
if len(path):
file = open(argv[i], 'wb')
else:
file = open(filename, 'wb')
ftp.retrbinary("RETR " + filename, file.write)
print " done"
file.close()
else:
print "no matching .."
ftp.close()
if __name__ == '__main__':
if (len(sys.argv) > 1):
main(sys.argv[off_file:])
else:
print "Usage: get [host] file_to_be_get\nHost: siteid(default), another_site\n"
Thursday, March 16, 2006
My user script to keep only news content in creaders.net and wenxuecity.com
// ==UserScript==
// @name Keep Only Interested Content
// @namespace http://leochen.net
// @description A script to remove all the un-necessary elements and display plain interested content only (version 0.3)
// @include http://*.wenxuecity.com/*
// @include http://*.creaders.net/*
// @include http://*.bbsland.com/*
// ==/UserScript==
var body = document.body;
var theContent = new Array();
var numContent = 0;
/* for wenxuecity.com */
if (/wenxuecity/.test(document.URL)) {
if (/www/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[3].childNodes[1].childNodes[10].childNodes[1].childNodes[0].childNodes[1].childNodes[0].childNodes[5].childNodes[2].childNodes[1].childNodes[2].childNodes[1].childNodes[0];
} else if (/news/.test(document.URL)) {
var content = document.body.childNodes[3].childNodes[1].childNodes[10].childNodes[1].childNodes[0].childNodes[1].childNodes[0].childNodes[5].childNodes[2].childNodes[1].childNodes[0].childNodes[1].childNodes[6];
theContent[numContent++] = content.childNodes[1].childNodes[2].childNodes[1].childNodes[1].childNodes[1].childNodes[0];
}
}
/* for creaders.net */
if (/creaders.net/.test(document.URL)) {
if (/headline/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[9].childNodes[1].childNodes[0].childNodes[3].childNodes[3];
} else if (/digest/.test(document.URL)) {
if (/pool/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[5].childNodes[13];
theContent[numContent++] = document.body.childNodes[5].childNodes[19].childNodes[1].childNodes[2].childNodes[3];
} else {
theContent[numContent++] = document.body.childNodes[5].childNodes[0].childNodes[8].childNodes[7];
}
} else if (/dailynews/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[5].childNodes[3].childNodes[1].childNodes[0].childNodes[3].childNodes[1];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
theContent[numContent++] = document.body.childNodes[1].childNodes[23];
theContent[numContent++] = document.body.childNodes[1].childNodes[27];
theContent[numContent++] = document.body.childNodes[1].childNodes[31];
}
}
if (/bbsland.com/.test(document.URL)) {
if (/bcchinese/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[25];
theContent[numContent++] = document.body.childNodes[1].childNodes[36];
}
}
if (/life/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[14].childNodes[7];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
}
}
if (/military/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
theContent[numContent++] = document.body.childNodes[1].childNodes[18];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
}
}
if (/general/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
}
}
if (/politics/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
} else {
theContent[numContent++] = document.body.childNodes[11];
}
}
if (/sports/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[14];
} else {
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[33];
}
}
if (/child/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[14];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[5];
theContent[numContent++] = document.body.childNodes[1].childNodes[6];
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[21];
}
}
if (/tea/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[31];
}
}
if (/joke/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
}
}
if (/iwish/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
}
}
if (/education/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[5];
theContent[numContent++] = document.body.childNodes[1].childNodes[6].childNodes[8];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[17];
theContent[numContent++] = document.body.childNodes[1].childNodes[27];
theContent[numContent++] = document.body.childNodes[1].childNodes[41];
}
}
if (/newland/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
theContent[numContent++] = document.body.childNodes[1].childNodes[37];
}
}
}
var len = body.childNodes.length;
/* remove all content */
for (i=0; i< len; i++) {
body.removeChild(body.childNodes[0]);
}
/* shown only interested elements */
for (i=0; i< numContent; i++) {
body.appendChild(theContent[i]);
}
// @name Keep Only Interested Content
// @namespace http://leochen.net
// @description A script to remove all the un-necessary elements and display plain interested content only (version 0.3)
// @include http://*.wenxuecity.com/*
// @include http://*.creaders.net/*
// @include http://*.bbsland.com/*
// ==/UserScript==
var body = document.body;
var theContent = new Array();
var numContent = 0;
/* for wenxuecity.com */
if (/wenxuecity/.test(document.URL)) {
if (/www/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[3].childNodes[1].childNodes[10].childNodes[1].childNodes[0].childNodes[1].childNodes[0].childNodes[5].childNodes[2].childNodes[1].childNodes[2].childNodes[1].childNodes[0];
} else if (/news/.test(document.URL)) {
var content = document.body.childNodes[3].childNodes[1].childNodes[10].childNodes[1].childNodes[0].childNodes[1].childNodes[0].childNodes[5].childNodes[2].childNodes[1].childNodes[0].childNodes[1].childNodes[6];
theContent[numContent++] = content.childNodes[1].childNodes[2].childNodes[1].childNodes[1].childNodes[1].childNodes[0];
}
}
/* for creaders.net */
if (/creaders.net/.test(document.URL)) {
if (/headline/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[9].childNodes[1].childNodes[0].childNodes[3].childNodes[3];
} else if (/digest/.test(document.URL)) {
if (/pool/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[5].childNodes[13];
theContent[numContent++] = document.body.childNodes[5].childNodes[19].childNodes[1].childNodes[2].childNodes[3];
} else {
theContent[numContent++] = document.body.childNodes[5].childNodes[0].childNodes[8].childNodes[7];
}
} else if (/dailynews/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[5].childNodes[3].childNodes[1].childNodes[0].childNodes[3].childNodes[1];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
theContent[numContent++] = document.body.childNodes[1].childNodes[23];
theContent[numContent++] = document.body.childNodes[1].childNodes[27];
theContent[numContent++] = document.body.childNodes[1].childNodes[31];
}
}
if (/bbsland.com/.test(document.URL)) {
if (/bcchinese/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[25];
theContent[numContent++] = document.body.childNodes[1].childNodes[36];
}
}
if (/life/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[14].childNodes[7];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
}
}
if (/military/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
theContent[numContent++] = document.body.childNodes[1].childNodes[18];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
}
}
if (/general/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
}
}
if (/politics/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[19];
} else {
theContent[numContent++] = document.body.childNodes[11];
}
}
if (/sports/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[14];
} else {
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[0].childNodes[1].childNodes[33];
}
}
if (/child/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[14];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[5];
theContent[numContent++] = document.body.childNodes[1].childNodes[6];
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[21];
}
}
if (/tea/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[16];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[31];
}
}
if (/joke/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[20];
theContent[numContent++] = document.body.childNodes[1].childNodes[29];
}
}
if (/iwish/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
}
}
if (/education/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[5];
theContent[numContent++] = document.body.childNodes[1].childNodes[6].childNodes[8];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[13];
theContent[numContent++] = document.body.childNodes[1].childNodes[17];
theContent[numContent++] = document.body.childNodes[1].childNodes[27];
theContent[numContent++] = document.body.childNodes[1].childNodes[41];
}
}
if (/newland/.test(document.URL)) {
if (/messages/.test(document.URL)) {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[15];
} else {
theContent[numContent++] = document.body.childNodes[1].childNodes[11];
theContent[numContent++] = document.body.childNodes[1].childNodes[22];
theContent[numContent++] = document.body.childNodes[1].childNodes[33];
theContent[numContent++] = document.body.childNodes[1].childNodes[37];
}
}
}
var len = body.childNodes.length;
/* remove all content */
for (i=0; i< len; i++) {
body.removeChild(body.childNodes[0]);
}
/* shown only interested elements */
for (i=0; i< numContent; i++) {
body.appendChild(theContent[i]);
}
my .screenrc file
startup_message off # default: on
# Affects the copying of text regions
crlf off # default: off
#vbell off
vbell_msg " __bell__ ! "
defscrollback 3300 # default: 100
#nethack on
bindkey -k kI copy
bindkey "^n" screen bash
bindkey "^b" next
bindkey "^v" prev
#bindkey "^p" prev
#bindkey "^1" select 0
#bindkey "^2" select 1
#bindkey "^3" select 2
hardstatus alwayslastline " %{= wk} %c | %d.%m.%Y | %{= Bw} %w %{= dd} "
screen -t "bash" 0 bash
# Affects the copying of text regions
crlf off # default: off
#vbell off
vbell_msg " __bell__ ! "
defscrollback 3300 # default: 100
#nethack on
bindkey -k kI copy
bindkey "^n" screen bash
bindkey "^b" next
bindkey "^v" prev
#bindkey "^p" prev
#bindkey "^1" select 0
#bindkey "^2" select 1
#bindkey "^3" select 2
hardstatus alwayslastline " %{= wk} %c | %d.%m.%Y | %{= Bw} %w %{= dd} "
screen -t "bash" 0 bash
Tuesday, March 14, 2006
Tuesday, March 07, 2006
CD dos path in unix environment
I am working under Windows, but I use cygwin in most of the time.
Sometimes, I have to change directory in cygwin to a path with UNC format (\\machine\path\to\xx)
I'm tired of typing the path myself, because in Unix, we have to use (//machine/path/to/xx) format.
Here comes a simple bash function to do all the conversion and change directory for me.
------------------------
function cddos () {
dos_path=$1;
cd `echo $dos_path | sed 's/\\\/\//g'`;
}
------------------------
Usage:
cddos '\\your\dos\path'
I love Bash function! It should be better to bash function instead of external bash script for such kind of small function.
Sometimes, I have to change directory in cygwin to a path with UNC format (\\machine\path\to\xx)
I'm tired of typing the path myself, because in Unix, we have to use (//machine/path/to/xx) format.
Here comes a simple bash function to do all the conversion and change directory for me.
------------------------
function cddos () {
dos_path=$1;
cd `echo $dos_path | sed 's/\\\/\//g'`;
}
------------------------
Usage:
cddos '\\your\dos\path'
I love Bash function! It should be better to bash function instead of external bash script for such kind of small function.
Thursday, March 02, 2006
A Perl oneliner to extract opcode from a formatted asm source file.
I created a simple perl script to help my co-worker to extract the opcode from some assembly source files.
-----------------------------
Sample Input:
Reset_Handler
$a
Init
0xc0200000: e59ff190 .... LDR pc,[pc,#400] ; [0xc0200198] = 0xc0200004
Instruct_2
0xc0200004: e59f0190 .... LDR r0,[pc,#400] ; [0xc020019c] = 0xc01e0000
0xc0200008: e321f0d1 ..!. MSR CPSR_c,#0xd1
Sample Output:
90
f1
9f
e5
//
90
01
9f
e5
//
d1
f0
21
e3
//
00
d0
40
e2
//
----------------------------------------------
My oneliner version:
perl -e 'map {print "$4\n$3\n$2\n$1\n//\n" if (/^\s+0x\S+:\s+(\S\S)(\S\S)(\S\S)(\S\S)\s+/);} <>; '
-----------------------------
Sample Input:
Reset_Handler
$a
Init
0xc0200000: e59ff190 .... LDR pc,[pc,#400] ; [0xc0200198] = 0xc0200004
Instruct_2
0xc0200004: e59f0190 .... LDR r0,[pc,#400] ; [0xc020019c] = 0xc01e0000
0xc0200008: e321f0d1 ..!. MSR CPSR_c,#0xd1
Sample Output:
90
f1
9f
e5
//
90
01
9f
e5
//
d1
f0
21
e3
//
00
d0
40
e2
//
----------------------------------------------
My oneliner version:
perl -e 'map {print "$4\n$3\n$2\n$1\n//\n" if (/^\s+0x\S+:\s+(\S\S)(\S\S)(\S\S)(\S\S)\s+/);} <>; '
Wednesday, February 22, 2006
Enhanced webbot to grab ads. info from vansky.com
This is an enhanced version of my webbot script to extract ad. info from vansky.com.
---------------------------------------------
#!/usr/bin/perl -w
# Hao Chen
# The purpose of this script is to extract ad info. from vansky.com website
# and write the data to grab.dat, email.dat files.
#
# grab.log file records the id of ads. grabbed to avoid redundant work.
#
use strict;
use LWP::UserAgent;
my $url_hp = 'http://www.vansky.com/vanphp/gg/newsgroup.php';
my $url_root = 'http://www.vansky.com/vanphp/gg/shownews.php?id=';
# starting id
my $start = 50000;
# ending id
my $end = 0;
# wait seconds
my $wait = 1;
my $ua = LWP::UserAgent->new;
$ua->agent( 'Mozilla/5.0' );
my ( $url, $req, $res );
my $verbose = 1;
$req = HTTP::Request->new( GET => $url_hp );
$res = $ua->request( $req );
if ( $res->is_success )
{
foreach ( split( "\n", $res->content ) )
{
if ( /pageno_c=(.*?)shownews\.php\?id=(\d*?)'\)/ )
{
$end = $2;
last;
}
}
}
open( FILE, 'grab.log' ) or die "Can't open file grab.log\n";
my @log = <FILE>;
close( FILE );
my $num_lines = scalar @log;
if ( $num_lines && $log[ $num_lines - 1 ] =~ / => (\d+) - (\d+)/ )
{
$start = $2;
} else
{
print STDERR "brand new task: start = $start\n";
}
if ( $end > $start )
{
print STDERR "new grab task: $start - $end\n";
} else
{
print STDERR "no new grab tasks!\n";
exit;
}
open( FILE, '>>grab.log' ) or die "Can't open file grab.log\n";
my $currTime = localtime;
print FILE $currTime . ' => ' . $start . ' - ' . $end . "\n";
close( FILE );
print "######## grab $start to $end #########\n";
open( FILE, '>>grab.dat' ) or die "Can't open file grab.dat\n";
open( EMAIL, '>>email.dat' ) or die "Can't open file email.dat\n";
print FILE "##### $currTime => grab ad. $start to $end\n";
print EMAIL "##### $currTime => grab email. $start to $end\n";
for ( my $id = $start; $id <= $end; $id++ )
{
$url = $url_root . $id;
print STDERR $url . "\n" if ( $verbose );
$req = HTTP::Request->new( GET => $url );
$res = $ua->request( $req );
if ( $res->is_success )
{
my @content = split( "\n", $res->content );
my $the_ad;
my $start_ad = 0;
foreach ( @content )
{
chop;
if ( /Address:/ )
{ #found the ad. line
$start_ad = 1;
}
$the_ad .= $_ if ( $start_ad );
if ( /<\/pre>/ )
{ #end of ad.
last;
}
}
if ( $the_ad =~ /<font color=darkblue size=5>(.*?)<\/td>.*Author: <\/b>(.*?)<\/td>.*Email:<\/b> (.*?)<\/td>.*Tel\.:<\/b><\/td><td align=left>(.*?)<\/td>.*Address:<\/b> (.*?)<\/td>.*<pre>(.*?)<\/pre>/ )
{
my $title = $1;
my $author = $2;
my $email = $3;
my $tel = $4;
my $address = $5;
my $ad = $6;
if ( $email =~ /.+\@.+\..+/ )
{
$email =~ s/ //g;
print EMAIL lc( $email ) . "\n";
}
print STDERR $id . ' : ' . $tel . ' : ' . lc( $email ) . "\n" if ( $verbose );
$ad =~ s/[\n|\r]//g;
print FILE $id . ' : ' . $title . ' : ' . $author . ' : ' . $email . ' : ' . $tel . ' : ' . $address . ' : ' . $ad . "\n";
}
sleep $wait;
}
}
close( FILE );
close( EMAIL );
exit;
---------------------------------------------
#!/usr/bin/perl -w
# Hao Chen
# The purpose of this script is to extract ad info. from vansky.com website
# and write the data to grab.dat, email.dat files.
#
# grab.log file records the id of ads. grabbed to avoid redundant work.
#
use strict;
use LWP::UserAgent;
my $url_hp = 'http://www.vansky.com/vanphp/gg/newsgroup.php';
my $url_root = 'http://www.vansky.com/vanphp/gg/shownews.php?id=';
# starting id
my $start = 50000;
# ending id
my $end = 0;
# wait seconds
my $wait = 1;
my $ua = LWP::UserAgent->new;
$ua->agent( 'Mozilla/5.0' );
my ( $url, $req, $res );
my $verbose = 1;
$req = HTTP::Request->new( GET => $url_hp );
$res = $ua->request( $req );
if ( $res->is_success )
{
foreach ( split( "\n", $res->content ) )
{
if ( /pageno_c=(.*?)shownews\.php\?id=(\d*?)'\)/ )
{
$end = $2;
last;
}
}
}
open( FILE, 'grab.log' ) or die "Can't open file grab.log\n";
my @log = <FILE>;
close( FILE );
my $num_lines = scalar @log;
if ( $num_lines && $log[ $num_lines - 1 ] =~ / => (\d+) - (\d+)/ )
{
$start = $2;
} else
{
print STDERR "brand new task: start = $start\n";
}
if ( $end > $start )
{
print STDERR "new grab task: $start - $end\n";
} else
{
print STDERR "no new grab tasks!\n";
exit;
}
open( FILE, '>>grab.log' ) or die "Can't open file grab.log\n";
my $currTime = localtime;
print FILE $currTime . ' => ' . $start . ' - ' . $end . "\n";
close( FILE );
print "######## grab $start to $end #########\n";
open( FILE, '>>grab.dat' ) or die "Can't open file grab.dat\n";
open( EMAIL, '>>email.dat' ) or die "Can't open file email.dat\n";
print FILE "##### $currTime => grab ad. $start to $end\n";
print EMAIL "##### $currTime => grab email. $start to $end\n";
for ( my $id = $start; $id <= $end; $id++ )
{
$url = $url_root . $id;
print STDERR $url . "\n" if ( $verbose );
$req = HTTP::Request->new( GET => $url );
$res = $ua->request( $req );
if ( $res->is_success )
{
my @content = split( "\n", $res->content );
my $the_ad;
my $start_ad = 0;
foreach ( @content )
{
chop;
if ( /Address:/ )
{ #found the ad. line
$start_ad = 1;
}
$the_ad .= $_ if ( $start_ad );
if ( /<\/pre>/ )
{ #end of ad.
last;
}
}
if ( $the_ad =~ /<font color=darkblue size=5>(.*?)<\/td>.*Author: <\/b>(.*?)<\/td>.*Email:<\/b> (.*?)<\/td>.*Tel\.:<\/b><\/td><td align=left>(.*?)<\/td>.*Address:<\/b> (.*?)<\/td>.*<pre>(.*?)<\/pre>/ )
{
my $title = $1;
my $author = $2;
my $email = $3;
my $tel = $4;
my $address = $5;
my $ad = $6;
if ( $email =~ /.+\@.+\..+/ )
{
$email =~ s/ //g;
print EMAIL lc( $email ) . "\n";
}
print STDERR $id . ' : ' . $tel . ' : ' . lc( $email ) . "\n" if ( $verbose );
$ad =~ s/[\n|\r]//g;
print FILE $id . ' : ' . $title . ' : ' . $author . ' : ' . $email . ' : ' . $tel . ' : ' . $address . ' : ' . $ad . "\n";
}
sleep $wait;
}
}
close( FILE );
close( EMAIL );
exit;
Tuesday, February 21, 2006
A script to extract email address from vansky.com website.
This scipt is to demonstrate how to use LWP::UserAgent to extract useful info. such as email address from website.
g.pl
#!/usr/bin/perl -w
# Hao Chen
# The purpose of this script is to extract email info. from vansky.com website
#
use strict;
use LWP::UserAgent;
my $url_root = 'http://www.vansky.com/vanphp/gg/shownews.php?id=';
# starting id
my $start = 10600;
# ending id
my $end = 12000;
# wait seconds
my $wait = 1;
my $ua = LWP::UserAgent->new;
$ua->agent( 'Mozilla/5.0' );
my ( $url, $req, $res, %emails, $email );
my $verbose = 1;
for ( my $id = $start; $id < $end; $id++ )
{
$url = $url_root . $id;
print STDERR $url . "\n" if ( $verbose );
$req = HTTP::Request->new( GET => $url );
$res = $ua->request( $req );
if ( $res->is_success )
{
foreach ( split( "\n", $res->content ) )
{
if ( /Email:<\/b> ([^<>\/]*)<\/td><\/tr><tr>/ )
{
$email = $1;
if ( $email =~ /.+\@.+\..+/ )
{
print STDERR $email . "\n" if ( $verbose );
if ( exists $emails{ $email } )
{
$emails{ $email } = $emails{ $email } + 1;
} else
{
$emails{ $email } = 1;
}
}
last;
}
}
sleep $wait;
}
}
print "######## email from $start to $end #########\n";
foreach my $key ( keys %emails )
{
print STDERR "$key => $emails{$key}\n";
print $key. "\n";
}
Monday, February 20, 2006
Cool Dynamic Bash Prompt
Put the following code in your .bashrc and you will see a lovely prompt.
The only problem is that it is a bit slow with the external "date" program.
You may write your own small apps to print out dynamic Bash prompt.
-----------------------------
function smiley () {
echo -e ":\\$(($??50:51))"
}
function thetime () {
echo -e `date +%H:%M`
}
export PS1="\$(smiley) \$(thetime) \w > "
The only problem is that it is a bit slow with the external "date" program.
You may write your own small apps to print out dynamic Bash prompt.
-----------------------------
function smiley () {
echo -e ":\\$(($??50:51))"
}
function thetime () {
echo -e `date +%H:%M`
}
export PS1="\$(smiley) \$(thetime) \w > "
Subscribe to:
Posts (Atom)