Monday, December 25, 2006

Fishing

Last night I went fishing in Shorncliff with my friends. None of us knew how to cast the net, so we improvised. We caught a fish by surprise. Surprise because we thought the net didn't spread when we cast it. So pulled on the net and up came a little fish (3 inch). No idea what fish it was.

The day ended when we accidentally let go of the net completely when we cast it.

Oh well. Headed home after throwing the fish back in the water.

Sunday, December 24, 2006

Tech Bust Due Soon

The news article Rampant M&A activity may signal peak of global business cycle puts forward a nice justification that the next bust is near. Current M&A activity ($3.6 trillion) has already exceeded the activity ($3.4 trillion) of the previous tech bust.

I, for one cannot understand YouTube's valuation of $1.65 billion or Skype's valuation of over $3 billion. Just numbers of customers cannot justify those valuation. Right now all of them are supported by ads. What will happen to them at the next bust when advertising money dries up. However, its all part of evolution. We need the next bubble bust to kill off the weak, so that the strong can rise from the ashes.

It takes a long time for a big company to die: so long that it's non-obvious that most of them are in fact dying, or at best treading water. We're a hit-driven industry. A few big successes can make it seem like everyone's doing well. But most of them have only had one hit. Go visit most tech companies, and all you'll find is a fussy henhouse parading around an aging goose that laid one or two golden eggs. All their innovation happened in the first act, and now they're focused on "managing for success." But that kind of managing is just staving off insolvency until a real innovator takes their business away. Any tech company overly focused on (or dependent on) its management is probably a good candidate for short-selling.

Friday, December 15, 2006

Graduation

Finaly had my graduation ceremony today. What can I say…

What I remember most are the:
1. professors (especially the most bizarre ones)
2. staying up late working on assignments
3. craming for exams the night before
4. writing a year long research paper in three weeks

Now I am officially an engineer. =)

Tuesday, November 28, 2006

Snippet://java/internet connectivity

The following code checks if internet can be accessed. Note, this is not a ping. The Java Socket class isn't capable of such low level function. However, since JDK5, Java java.net.InetAddress.isReachable(int) can be used to check if a server is reachable or not.

isReachable() will use ICMP ECHO REQUESTs if the privilege can be obtained, otherwise it will try to establish a TCP connection on port 7 (Echo) of the destination host. But most Internet sites have disabled the service or blocked the requests (except some university such as web.mit.edu).



public static boolean checkinternet(String url) {
try {
InetAddress address = InetAddress.getByName(url);
System.out.println("Name: " + address.getHostName());
System.out.println("Addr: " + address.getHostAddress());
System.out.println("Reach: " + address.isReachable(1000));
} catch (UnknownHostException e) {
System.err.println("Unable to lookup " + url);
} catch (IOException e) {
System.err.println("Unable to reach " + url);
}
}

Monday, November 27, 2006

Snippet://Java/InputStream

Java code snippet to read InputStream using buffered reader.

NOTE: StringBuffer won't insert extra \n, so the returned string will be exactly as the InputStream. Also, the unusual for statement in the snippet below is 10 times faster than the traditional while statement.

private static String slurp(InputStream in) throws IOException {
StringBuffer out = new StringBuffer();
byte[] b = new byte[4096];
for (int n; (n = in.read(b)) != -1;) {
out.append(new String(b, 0, n));
}
return out.toString();
}

java.lang.NoClassDefFoundError: org/apache/commons/

I'm working on a project where I needed to make HTTP requests. Instead of reinventing the wheel, I decided to use Apache-Commons-Httpclient library. Upon compiling, the code blew up in my face. The debugger says:

java.lang.NoClassDefFoundError: org/apache/commons/logging/LogFactory

Hmm... so Apache-Commons-Httpclient is dependent upon Apache-Commons-Logging library. After I download and add the required library, I compile the code to have it blow up in my face again. This time the debugger says:

java.lang.NoClassDefFoundError: org/apache/commons/codec/DecoderException

Good god, another dependency. This time Apache-Commons-Codec. That's what I hate about external libraries. None of them is self contained. I need the code small enough to fit in embedded devices like a cell phone. Good thing Apache-Commons library is opensource, so the source code is available. I need to strip it down. So much for reinventing the wheel.

Friday, November 17, 2006

High Avalibility - 5 Nines - 99.999%

Recently machinehead asked me how to build high availibility. For the less technically inclined, high availibility is a measure of the reliability of a system and sometimes indicated by "Five nines" or 99.999% reliability. Basically what it means is the system has a total downtime of no longer than five minutes per year.

Having done a course on distributed systems, I decided to take a crack at it. There are three main issues that needs to be addressed:

  • Hardware

  • Software

  • Data

Hardware
Everyone thinks high availibility means only hardware. Its not. Its only the beginning. Firstly, you need at least two of everything located geographically apart. For example, most people believe RAID is insurance against data loss. But RAID only protects against one or two drive failure. What happens if the PSU shorts and pump 240V instead of 5V. That will fry every disk in the array. Plus, rebuilding disk takes time which isn't high availibility. It should also be geographically located apart to protect against fire, earthquake, tsunami, plane etc.

Secondly, you need smooth, automatic switchover incase of crash. For example, using heartbeat to monitor servers and changing the IP at DNS server to the failover server. This makes it smooth and you don't need to make any other server aware of the server failure.

Software
Software should be built from ground up with high availibility in mind. Meaning they should be scalable and clusterable. The best way to do this is to make them stateless. For example, when you click on "2" to goto the second result page on Google, the second page doesn't necessarily have to be processed by the same server that did the initial first page. It can be done by any server. This is the power of stateless.

Data
Data is the hardest problem to solve. If you fragment and replicate the data for performace and scalability, you need to address sync issues. How would you lock and commit multiple partitions? How would you detect deadlocks? 2 phrase lock and 2 phrase commit is not an easy answer. eBay takes down the site on Monday 12-4AM every week to archive sold items so as to keep the fragments small (smaller fragments means faster searching by the database). Yes, this 4 hours is "planned" downtime and no, planned downtime does count towards high availibility.