Saturday, August 11, 2007

Links - 08/11/2007

Man vs machine, or, from SLA to SLAuto (Isabel Wang): Isabel was kind enough to provide her comments on SLAuto, and--no surprise--she get's it. In fact, her analogy at the bottom of this post is a wonderful one, and I hope she's OK if I use it (with appropriate credit, of course :):

"You don't need no SLAuto, you say, because you've got great customer service reps and data center techs? Well... 10 years ago I used to know people who prided themselves on their ability to dish out web space manually. They could charge credit cards and create customer folders faster than anyone else! Then competitors started using auto-provisioning tools and they went out of business. History will repeat itself."

Is the Tap Dry? (CXOtoday, India: Tahirih Gaur): Tahirih describes India's biggest aparent obsticals to utility computing: storage and inefficient management of outsourced IT. (Does anyone else see an irony in that? James McGovern?) She notes that many companies (banks for instance) have a problem with storing sensitive data on disks shared with competitors. She also quotes a Gartner statistic that 80% of all outsourcing deals are renegotiated within 3 years. I've posted on this before, and Nicholas Carr is writing extensively about it, but make no mistake that the move to utility computing is even more of a cultural shift than a technical shift. My employer is betting on the fact that a large number of organizations will not be comfortable outsourcing their utility computing entirely, and want to create a utility within their own infrastructure.

Utility computing's elusive definition (CNet news.com: David Becker): In searching for more coverage on utility computing, I came across this 2003 article covering a panel discussion on the topic at that year's Comdex. My first reaction in reading it is that the more things change, the more they stay the same. All of the issues presented here remain true today. I don't see one element of this article that doesn't ring true today (other than new marketing names for the vendor products).

I also love the proposition by Tony Siress (then Senior Director of Advanced Services for Sun) of transportation as a better analogy of utility computing than electricity:

Siress maintained that transportation is a better analogy, considering how people employ a combination of owned, leased and rented cars along with taxis to meet their changing transit needs. "Taxi cabs are a good example of a fully outsourced piece of infrastructure, and they're the right approach in some situations" he said. "The trick is understanding the mix of approaches that delivers the highest value and the least risk to you."

This actually highlights something that I have trouble remembering sometimes; that utility computing isn't a one-size-fits-all approach. Not every application is appropriate for managed hosting, nor does every one require a private IT utility. Some "trips" (analogous to either transactions or functions?) require multiple "modes of transportation": a little SaaS, a little hosted virtual server capacity, even a few bare metal servers in a closet thrown in for good measure. The challenge for SLAuto is to provide policy across all of these, or at least provide the building blocks to do so.

Functionality (or "service flow") is the electricity in utility computing; hardware and software are just the generators and transformers. The network is the power line, and SLAuto is the demand management system.

Thursday, August 09, 2007

Links - 08/09/2007

Technology companies tout greener credentials, but significant improvements are well off (Associated Press via Technology Review): This article is a layman's explanation about the energy consumption imposed by data centers, and the reasons behind that consumption. There are some interesting statistics here (most of which are recycled--pardon the pun), but the article is light on possible solutions. Remember, the entire purpose of SLAuto is to deliver the required service levels for your business using the most cost (and energy?) effective resources necessary to do so.

"Why is Amazon Web Services partnering with NaviSite?" (Isabel Wang): An overview of several interesting trends around utility computing in the managed hosting market. My comments:
  1. NaviSite is providing an interesting service here, and one I think will need to evolve to an automation model eventually. Today it looks like your same old monitoring-only "management" environment, but with a few interesting hooks who knows.
  2. I love the PlanetWide Media story only because it is one of the first example of "Web 2.0" infrastructure mashup that I've seen. Why not Web 3.0? Because I would bet right now that Planet Media is economically locked in to LayeredTech for the foreseeable future (i.e. the cost of moving their software would negate the benefit of moving it).
  3. Hmmm. Perhaps the future goes so far as to divide mindshare into recombinable building blocks. (OK, sorry, the BS meter pegged on that one...)
  4. Isabel wraps up with a comment about the long tail of computing itself, which I believe is the real next revolution in IT that will drive new businesses and perhaps even industries. I believe Nicholas Carr agrees, but we shall have to wait an see.

Green data center scuttlebutt from NGDC conference (Server Specs: SearchDataCenter.com): The interesting part of this video to me is the description of the GDC panel discussion as "sniping". I agree, and that's why I left early. Nothing interesting came out of the discussion other than the fact that each major vendor is still months or even years away from actually reducing the energy burden of the data center market (reinforcing the AP article above). I say watch for solutions that may be a little more short term.

Wednesday, August 08, 2007

"Web 3.0" and Infrastructure

In "What is Web 3.0?" Nicholas Carr breaks down various early definitions of Web 3.0 for the reader. In the end, he offers the following:
Web 3.0 involves the disintegration of digital data and software into modular components that, through the use of simple tools, can be reintegrated into new applications or functions on the fly by either machines or people.
Great definition, but again leaves out the importance of infrastructure on the equation. (Wait, wasn't Carr the one who pointed out that it all starts with infrastructure? What happened, Nicholas?) I would modify his statement to read
Web 3.0 involves the disintegration of digital data, software and infrastructure into modular components that, through the use of simple tools, can be reintegrated into new applications or functions on the fly by either machines or people.
To get a sense of how this technology would affect infrastructure, picture a world in which every aspect of the infrastructure stack, from application server to operating system to bare metal server to network fabric to shared storage, etc., can be assembled as necessary to meet the service level needs of an application (or even an application function). Need a J2EE service to run at 4 9's up time? Choose from a smorgasbord of app server vendors running on a selection of Hardware as a Service vendors with access to any number of supporting services from a variety of Software as a Service vendors--or let a service level automation tool (whether an appliance, a software product or a SaaS offering) do it for you. Ideally, let the SLAuto determine the most cost-effective way to deliver your service at the SL's you require.

To be fair, this is a ways down the road, but then so is anything Web 3.0.

Links - 08/08/2007

As you can see on the upper right hand column of my blogspot page, I have added myself to the findtechblogs.com realm. Seems like a very cool service, and its already given Ken Oestreich some visibility. (See below.)

The CMDB - An anemic answer for a deeper crisis (Fountnhead; Ken Oestreich): Ken is, of course, a collegue, so I may be a little biased, but I love what he is saying here. Basically, if you are keeping a CMDB, you are keeping manual records of the state of your data center. If you automate with a system that tracks its actions, then you have an automated way to keep those records. I think he washes over the legal requirements for a CMDB a little bit, but other than that, you need to read this. I learned of this post via Google Alerts, and findtechblogs.com.

Nirvanix To Challenge Amazon S3 (TechCrunch): A San Diego startup is daring to challenge Amazon's S3 dominance of the nacent Storage-as-a-Service market. Judging from the comments, these guys have their work cut out for them.

Green Grid lays out 2007 roadmap (Between the Lines; ZDNet): This was the only interesting piece of news to come out of the Green Data Center panel at LinuxWorld, IMHO. When combined with the EPA study, it looks like a trickle of real science being injected into the "green" hype. Its still all talk at this point, but I expect to see some useful tools and guidance from both sources. I hope that optimizing to service levels is one of the key criteria, though.

Tuesday, August 07, 2007

Links - 08/07/2007

Spent the morning at home waiting to resolve jury duty (I'm free! I'm free!), and the afternoon at LinuxWorld. More on that later.

Scratching itches in the cloud (O'Reilly Radar): O'Reilly and Sriram Krishna describe three key problems being introduced by "the cloud", including the inherent difficulty of working with the software itself (whether its open or closed source), potentials for "data lock-in", and the barrier to entry of building a competitive site. Again, I think this argues for a standards around utility computing portability--not just for data, as O'Reilly suggests here, but also for server images and application images. See the earlier discussion documented here and here.

Web Services war is over: Time to REST (The Future of Software): Puh-lease... Those of us with real experience in developing highly scalable distributed applications have always known that WS-* was more vendor opportunity than great architecture, but I'll believe REST has won when I see Google, PG&E and others convert their web services from SOAP-RPC/JAX-WS/whatever. Amazon is a force to be reckoned with, but there is a lot of war left to be fought. It reminds me of the COM+ (now .NET) vs. J2EE battles fought in the late nineties--any winner there yet? The good news is that none of this matters to SLAuto users, assuming their service level monitors at the service component level understand both REST and WS-*.

Service Must Be Job #1 In The Data Center (InformationWeek): Gee, does that mean that it all comes down to meeting service expectations...as defined in service level agreements...which can, in forward thinking organizations, lead to policies that can automate resource allocation...which, in turn, leads to reliable measurement of resource consumption? Believe me when I tell you that this is exactly why I am passionate about this subject. Best quote in the article:

As eBay's Smart put it: "The next generation data center has to be about changing the nature of the data center and its relationship to the business. Understanding these relationships allows you to go to your boss and say that you know the cost of delivering the value of this service."

Amen.

Monday, August 06, 2007

Links - 08/06/2007

In the interest of increasing my post frequency while not increasing the workload it imposes, I thought I'd take a hint from some of the more popular bloggers in my blogspace. Starting today, I will try to show everyone what I am reading online, and why I think it has importance (if any) to Service Level Automation and utility computing.

This will not replace my longer posts covering key topics in SLAuto or utility computing. I simply intend it to replace my tendency to put an interesting post or article aside saying "I'll blog on that some day", never to return.

Here goes day 1:

Sometimes 69 million > 143 million (Isabel Wang): I've talked before about the great consolidation of computing capacity that is coming our way. (Not a complete consolidation, mind you--there will always be private data centers for highly secure applications, and I believe there will be dozens of "boutique" capacity providers.) Isabel is covering Amazon's new payment service here, but her last paragraph on the winning providers supplying framework and application layers in addition to pure hardware capacity is right on the money.

EPA sends final report on data center energy efficiency to Congress (SearchDataCenter.com): It should be obvious why is important to SLAuto/Utility Computing/whatever. The chart from the executive summary says it all: we are on the hook as an industry to improve our practices as it relates to both utilization and power management. Have you thought about what that will take in your data center?

Virtualization users say, 'Better management tools, please' (SearchServerVirtualization.com): Another Survey reinforcing what we've been saying all along; if you jump into the virtualization bandwagon, that's great, but be prepared for a management nightmare. While I would agree that improving the "virtualization awareness" of traditional management tools might help, I would argue that you still have too much volume to handle without automation. In this context, when you see "management", I recommend that you think automation.

Microsoft Building the Ultimate Spyware System (Jek Hui): Of the many posts covering Ray Ozzie's overview of Microsoft's utility computing play, I chose this one a) for the headline, and b) for the generally succinct coverage of Ray's comments. Believe me, if Microsoft can really overcome the "innovator's dilemma" and make this work, they will knock two or three of the utility computing competitors out of the race. So far, it's Google vs. Amazon (see above), but I like Microsoft's chances here. Will the service be sensitive to the needs and privacy of its users? I know where I'd put my money...

Monday, July 30, 2007

SLAuto vs. SLA

A while back, Eric Westerkamp over at eCloudM asked a simple question that got me thinking. Eric wondered aloud (in print?) whether "using the Term Service Level Automation (SLA) causes confusion when presenting the ideas and topics into the business community. I have most often seen SLA refer to Service Level Agreements. While similar in concept, they are very different in implementation."

So, to keep things clear in my blog, I will now use the acronym SLAuto for Service Level Automation, and retain the SLA moniker for Service Level Agreements. I hope this eliminates confusion and allows the market to talk more freely about Service Level Automation.

Speaking of SLA, though, Steve Jones posted a great example of symbiotic relationship between the customer and the service provider in the SLA equation. To put it in the context of an enterprise data center, you could offer 100% up time, .1 sec response time and a 5 minute turnaround time, but it wouldn't be of any value if the customer's application was buggy, they were on a dial-up network and it took them six weeks to get the requirements right for a build-out.

Now, let's look at that in the context of SLAuto. To my eye, the service provider in an SLAuto environment is the infrastructure. The customer is any component or person that accesses or depends on any piece of that infrastructure. Thus, any SOA service can be a service provider in one context, and a customer in another. Even the policy engine(s) that automate the infrastructure can be thought of as a customer in the context of monitoring and management, and a service provider in terms of an interface for other customers to define service level parameters.

Steve's example hints that I could buy the "bestest", fastest, coolest high tech servers, switches and storage on the planet and it wouldn't increase my service levels if I couldn't deliver that infrastructure to the applications that required it quickly, efficiently and (most importantly) reliably. Or, for that matter, if those applications couldn't take advantage of it. So, if you're going to automate, your policy engine should be (you guessed it) quick, efficient and reliable. If it isn't, then your SLAs are limited by your SLAuto capabilities.

Not what you intended, I would think...

Wednesday, July 25, 2007

Why it all boils down to infrastructure management...

Digital Daily made my day today with the story of data center operator 365 Main's declaration of its amazing up time record followed hours later by complete loss of their San Francisco data center for a good hour--a data center populated by a veritable "who's who" of Silicon Valley tech companies. As a result of the SF power outage, a good portion of the Web 2.0 social network world was also down.

What does this have to do with Utility Computing? Well, as Nicholas Carr points out, it has everything to do with the key problem of computing today: not software, not social networks, but infrastructure. 365 Main is (reportedly) one of the best collocation hosting facilities in the world, and despite a bunch of power failure backup, including two sets of generators, they didn't get the job done. Worse yet, they were bitten by something they can do little about: old and tired power infrastructure.

When we talk about utility computing architectures and standards, we must remember that, despite infrastructure becoming a commodity, it is still the foundation on which all other computing services are reliant. If we can't provide the basic capabilities to keep servers running, or quickly replace them if they fail, then all of our pretty virtual data centers, software frameworks and business applications are worthless.

For 365 Main, their focus initially will probably be on why all that backup equipment didn't save their butts as it was designed (and purchased) to do. I mean, as much as I would like to make this a story about how great Service Level Automation software would have saved the day, I have to admit that even my employer's software can't make software run on a server with no electricity.

However, and perhaps more importantly, the operations teams of those Silicon Valley companies that found themselves without service for this hour or so must ask themselves some key questions:
  1. Why weren't their systems distributed geographically, so that some instances of every service would remain live even if an entire data center were lost? If the problem was the architectures of the applications, that is a huge red flag that they just aren't putting the investment into good architecture that they need to. If the problem is the cost/difficulty of replicating geographically, then I believe Service Level Automation/utility computing software can help here by decoupling software images from the underlying hardware (virtual or physical).
  2. Why couldn't they have failed over to a DR site in a matter of minutes? Using the same utility computing software and some basic data replication technologies, any one of these vendors could have kept their software and data images in sync between live and DR instances, and fired up the whole DR site with one command. Heck, they could even have used another live hosting facility/lab to act as the DR site and just repurposed as much equipment as they needed.
  3. How are they going to increase their adaptability to "environmental" events in the future, especially given the likelihood that they won't even own their infrastructure? Again, this is where various utility computing standards become important, including software and configuration portability.

I recommend reading Om Malik's piece (as referenced by Carr) to get an understanding why this is more important than ever. According to Malik, hosting everywhere is a house of cards, if only because of the aging power infrastructure that they rely on. Some people are commenting that geographic redundancy is unnecessary when using a top-tier hosting provider, but I think yesterday's events prove otherwise.

What I believe the future will bring is commoditized computing capacity that, in turn, will make the cost of distributing your application across the globe almost the same as running it all in one data center. You may be charged extra for "high availaility" service levels, but the cost won't be that significant, with the difference mostly going to the additional monitoring and service provided to guarantee those service levels.

Of course, that will require portability standards...

Friday, July 20, 2007

Where's the standard, bub?

Simon Wardley's Bits and Pieces blog has an interesting post breaking down what he thinks are the three key markets for utility computing:
  • Saas: Software as a Service

  • FaaS: Frameworks as a Service

  • HaaS: Hardware as a Service
Such as it goes, this is OK, but I think his most interesting comments surrounded the need for Common Service Providers in each of these areas, and the need for portability across those providers, and the mechanisms that he thinks will drive the standards that will enable portability. To quote:

The issues and the needs of a competitive utility computing market are also the same at each level - portability, multi-providers and agreed standards and solves the same class of problems - disaster recovery, scalability, efficiency and exist costs.

In today's world the fastest way to achieve a standard is not through committee, conversation or whitepapers but through the release and adoption of not only a standard but also an operational means of achieving a standard.

Hence such utility computing standards will only be achieved through the use of open source, without any one CSP being strategically disadvantaged to any owner of the standard.

To be sure, this is controversial, but it aligns nicely with an observation that I and others have had about Xen, and why it may struggle to supersede VMWare, despite being freely distributed by just about every major OS vendor out there.

One of my colleagues put it best in an email:

As to the larger question of why Xen is failing miserably, I would like to profess this opinion -- Storage. The Xen / KVM / Linux / RedHat community botched storage. Hence, they are failing in the marketplace.

To elaborate:

Virtualization brings two important benefits:
  1. Seperates the OS+App stack from the underlying hardware

  2. Enables you to package the OS+App stack into a VM that you can fling around with ease....this is the storage angle.
Xen accomplished (1) reasonably well.

Xen failed miserably with (2), i.e Storage. VMware solved the storage issues admirably well. So long as Xen solutions do not work well in the "copy virtual disk files around and run them anywhere" model, Xen will not succeed.
In other words, Xen has issues with portability by not providing a file representation of virtual machine storage that can be moved between disparate physical systems with ease. I know there are some virtual appliance vendors out there that do Xen, so maybe the problem is solved with their technologies. However, there is no standard proposed by Xen, and thus there is no portable standard for Xen VMs as files.

Alas, VMWare has a nice portable file representation of a VM. Granted, there are portability issues there, as well, but by and large VMWare has a much better solution to portability--within VMWare hosts. Unfortunately, there is still no solution (that I know of) that will run the same file system on both VMWare and Xen. Thus, no universal portability is coming soon from the VM space.

Recently, I have been telling anyone who will listen that this nascent utility computing market is still searching for a standard for server (VM/framework/application/whatever) portability across disparate utility computing service providers. I like the concept of a virtual appliance, but we need a (non-proprietary) standard, or we need another portability mechanism besides VMs. (As a side note to my new friends at eCloudM--this is definitely an opportunity, though it may not meet your criteria.)

Otherwise, utility computing will be "choose your vendor and build your software accordingly", not "build your software as you like and choose any vendor you want".

Please, if I am way off, correct me with a comment or a blog post. I would love to find out I am wrong about this...

A Helping Hand Comes In Handy Sometimes

You may remember my recent post on how data centers resemble complex adaptive systems. This description of a data center has a glaring difference from a true definition of complex adaptive systems, however; data centers require some form of coordinated management beyond what any single entity can provide. In a truly complex adaptive system, there would be no "policy engines" or even Network Operations Centers. Each server, each switch, each disk farm would attempt to adapt to its surroundings, and either survive or die.

Therein lies the problem, however. Unlike a biological system, or the corporate economy, or even a human society, data centers cannot afford to have one of its individual entities (or "agents" in complex systems parlance) arbitrarily disappear from the computing environment. It certainly cannot rely on "trial and error" to determine what survives and what doesn't. (Of course, in terms of human management of IT, this is often what happens, but never mind...)

Adam Smith called the force that guided selfish individuals to work together for the common benefit of the community the "invisible hand". The metaphor is good for explaining how decentralized adaptive systems can organize for the greater good without a guiding force, but the invisible hand depends on the failure of those agents who don't adapt.

Data centers, however, need a "visible hand" to quickly correct some (most?) agent failures. To automate and scale this, certain omnipotent and omnipresent management systems must be mixed into the data center ecology. These systems are responsible for maintaining the "life" of dying agents, particularly if the agents lose the ability to heal themselves.

Now, a topic for another post is the following: can several individual resource pools, each with their own policy engine, be joined together in a completely decentralized model?

Sunday, July 08, 2007

Web 2.0, Utility Computing and Service Level Automation

Industry Girl has an interesting observation of what it is that is driving utility computing. She believes that it is the drive towards "Web 2.0 sites and applications, like the video on YouTube or the social networking pages on MySpace" is creating huge demand on backend server infrastructure--unpredictable demand, I may add--which, in turn, is creating the need for truly dynamic capacity allocation. Add to that the trend of Web 2.0 technologies being used by more and more commercial and public organizations, and you begin to see why it's time to turn your IT into a utility.

I have to say I agree with her, but I would like to observe that this is only a (large) piece of the overall picture. In reality, most of the sites she specifically mentions are actually "Software as a Service" sites targeted at consumers rather than businesses. Its the trend towards getting your computing technology over the Internet in general that is the real driving need.

For "utility computing" plays such as SaaS companies, managed hosting vendors, booksellers :), and others, the need for utility computing isn't just the need to find capacity, it is also the need to control capacity. This, in turn, means intelligent, policy-based systems, that can deliver capacity to where it is needed, share capacity among all compatible software systems and meter capacity usage in enough detail to allow the utility to constantly optimize "profit" (which may or may not be financial gain for the capacity provider itself).

Service Level Automation, anyone?

Web 2.0 drives utility computing, which in turn drives service level automation. So, Industry Girl, I welcome your interest in utility computing, and offer that the extent to which utility computing is successful is the extent by which such infrastructure delivers the functionality required at the service levels demanded. Welcome to our world...

Sunday, June 24, 2007

God in the Grid...

Utility computing must really be hitting the mainstream when its seen as the future of Christian fellowship...

By the way, I've been meaning to preorder Nicholas Carr's book as well. He has a web site that explains the book's premise. If you want to begin dreaming about a world where computing costs very little, and automation, virtual societies and electronic economies abound, there seems to be plenty of inspiration here.

I just hope it's as good as it sounds...

Wednesday, June 20, 2007

Clarification on James McGovern's recent comments

Man, its been a while since I posted. Sorry to my loyal fans, but my last post should explain where my head has been at the past few weeks.

I wanted to take a moment to clear up some misunderstandings in James McGovern's comment on my recent third-hand coverage of the Forrester Research 2007 IT Forum conference (via Ken Oestreich who actually attended). I have huge respect for James' pragmatic, get to the point style, and he is definitely a leading edge player in both the Enterprise Architecture and Security spaces. However, James mistakenly credits the comment, "There are no IT projects anymore, only business projects", to "Ken Oestreich of Forrester". Ken is not an employee of Forrester Research (he is employed by Cassatt as Director of Product Marketing), and he was just reporting what he heard. I am not aware of the name of the Forrester Analyst who made that statement.

Having said that, clearly this quote can be interpreted in a variety of ways. James seems to have read the quote as "IT is dead" or some such thing, where I think the quote is intended to show that IT projects are now founded initially on business need, not on the whims of IT professionals eager to introduce new technologies. As I read James' blog, this seems right up his alley. However, I would agree that the statement is more hype than substance--thus the lack of a specific interpretation.

One other thing James says bugs me a bit. Perhaps I misinterpreted, but the following statement seems way off:

"Maybe he should acknowledge that the vast majority of CIOs aren't even focusing on data centers as this has been commotitized (sic) and therefore pushed several layers down in the organization. Likewise, infrastructure stuff simply doesn't allow an enterprise to either innovate nor sustain competitive advantage where as software development still has the potential for both."

Really? I think if James were to ask his CIO what his number one expense line item was, it wouldn't be research and development. Wonder why utility computing is a key initiative for CIOs in the coming years? Why is "Green Data Center" all over the press these days? I can tell you: operations (labor, facilities, infrastructure and utilities) accounts for the vast majority of enterprise IT budgets. Even those that "outsource" (domestically or otherwise) to Managed Hosting Providers are paying a princely sum for the service.

Add to that the "siloed" nature of most application deployments in data centers, and the incredible barrier to market agility that causes. A giant "bank of our continent" recently filled an RFP for utility computing (in the "turn our IT into a utility" sense), but cost savings wasn't even the most important factor. The bank can only grow through (international) acquisition, and each acquisition has been burdensome largely due to the service level losses caused by integrating each IT organization and its infrastructure. Agility with guaranteed service levels is the bank's number one operations priority.

That's not to say that software isn't also a priority. It is for all the reasons that James alludes to. It certainly is the quickest (only?) route to new revenue streams, and it can also lead to significant cost savings if done right. Hell, if we can get the cost of operations down, it will free up more funds for this important endeavor! But to say infrastructure has been pushed way down the list just doesn't jibe with the pain we are finding in corporate IT today.

Let me make it clear that I have great respect for James McGovern, and I read his blog every day. I hope that we can continue a conversation about both the role of infrastructure innovation in the future of IT, and the relationship between application architectures and their deployment architectures in the data center.

Sunday, June 03, 2007

Want to save gas? Stop leaving your car idling in the garage!

History doesn't repeat, but it sometimes rhymes...

I remember the seventies, when gas prices skyrocketed (the first time) and there were suddenly all these tiny cars on the road. One member of my mom's congregation even showed up one Sunday with this crazy little car that ran on a motorcycle engine. It was made by some new car company called Honda, and it was one of the first years that Civics were sold in America.

As a nation, we clamored to change our lifestyles--ditching heavy steel muscle cars for sporty (or utilitarian) little "economy cars". Our approach to solving the energy crisis was to increase the efficiency with which our cars consumed energy. Note, however, it was not (by a long shot) to reduce the amount of driving we did.

Now flash forward to today, and take a look at the current energy crisis in America's (and the world's) data centers. Electricity is expensive, and growing more so (except for those lucky enough to have subsidised power). Add to that concerns about global climate change, and you've got company after company scrambling to be "green".

Again, however, note that the target is not to do less computing than we did before. In fact, if anything, the demand is increasing for information technology and business automation. I believe pushing the automation envelope is going to take more computing power than we know.

So, like the automobile vendors of the seventies, today's systems vendors are working hard to release "energy efficient" models of servers, laptops and desktops. They do this ostensibly to give us all a good feeling about what good stewards of our tiny planet we are, but in reality its all about saving money. None of this changes our worst behaviour, however; our tendency to leave as much capacity running as possible at all times, "just in case".

Of course, the server that uses the least amount of power is the one that is turned off. That's where Service Level Automation comes into the picture. As noted in the past, one of the key aspects of a good Service Level Automation platform is the capability of shutting down anything that isn't serving an immediate business need. Traditionally, I've always talked about this in relation to scale-out applications--your SLA platform should shut down servers not needed to meet current demand in such applications. Now, however, I want to talk about three use cases where SLA enhances the day to day power consumption of all applications in the data center.
  1. Job-specific management. OK, think of every server you've touched in the last six months. How many of those served a short term purpose (e.g. getting a software release out the door), but frequently spend days unused for any purpose. I remember going days or even weeks between placing builds on staging servers in my previous life. Service Level Automation should be able to detect unused software payloads, and shut down that equipment until needed again by that or any other payload.
  2. Time-specific management. Almost every data center (especially development and test labs) have systems that are hit hard during some portion of the day, then remain idle for the remainder. SLA should provide the capability to not only schedule system shutdowns, but to actually look at that status of systems to determine which are best candidates for shutdown. In other words, go beyond automating "blind" scheduled events to delivering intelligent management of system power cycles.
  3. Power emergency management. One of the great benefits of living in the San Francisco Bay area is the incredible ingenuity of our power utility in encouraging companies to conserve power and "be good neighbors" in a power emergency. PG&E offers rebates to companies willing to join Demand Response programs, where they agree to voluntarily reduce electric consumption to help the utility avoid the infamous "rolling blackout".

The Silicon Valley Leadership Group has recently been hosting a series of events around "Energy Efficient Data Centers", one of which targeted how SLA could deliver on all three of the above. The response was tremendous--so much so that my employer has asked me to join a team building a simple targeted solution to these problems based on our already innovative SLA platform. I can't say much more right now, but I certainly will communicate all that I can as soon as I can.

By the way, the first lesson I've learned from all of this is that power measuring capabilities varies widely from data center to data center. Some companies can't tell you anything more than their monthly bill, others can show you power consumption over time at the individual server level. Part of the issue is that there are no "simple" power metering solutions at the server level...power controllers (i.e. iLO2) are just now starting to give management systems access to the power measurement tools on Intel and AMD boards. MPDUs have some good features, but they vary widely from vendor to vendor.

You can't control what you can't measure, so get on board system vendors! Give us the tools we need to measure and manage those beautifully efficient next generation servers. Heck, give us the tools we need to measure and manage all those older systems we have out there now. That would be more green by far than just squeezing another milliamp out of a MIP.

Thursday, May 24, 2007

Service Level Automation Deconstructed: Respond

For the third and last in my series breaking down the three key assumptions behind Service Level Automation, I would like to focus on how SLA environments can control data center configuration in response to service level goal violations. These goal violations and the high level actions to be taken are determined by the analysis capability of the environment. Details of how to accomplish those high level actions, however, are decided and executed by the response function.

Essentially, the response function of an SLA environment is very much like the driver set that your operating system uses to translate high level actions (e.g. "store this file") to device specific actions ("Move head 32 steps to center, find block 4D5EF, etc."). The responsibility here is to provide the interface between the SLA analysis engine and specific standard or proprietary interfaces to everything from server hardware to network switches to operating systems and middleware.

I see the following key interface points in today's environments:

  • Power Controllers/MPDUs: Job 1 of a service level automation environment is providing the resources required to meet the needs of the software environment, and only those resources. Turn those servers on when they are needed, and off when they are not. This includes virtual server hosts. (Examples: DRAC, iLO, RSA II,MPDUs)
  • Operating Systems: Before you shut off that server, make sure you've "gently" shut down its software payload. Well written server payloads for automated environments will both start up and aquire intial state (if any), and shut down while preserving any necessary state without human intervention. However, from a communications perspective, each action starts with the OS. (Examples: Red Hat, SuSE, MS Windows, Sun Solaris)
  • Middleware/Virtualization: It is interesting to note that many software payload components (e.g. an application server or a hypervisor) are both software to be managed, and computing capacity themselves. For example, an application server should be managed to specific service levels relating to its relationship with its host server (e.g. CPU utilization, thread counts, etc.), while also treated as a capacity resource for JavaEE applications and services. As such, these software containers should be managed for their own guest payloads much like a physical server would for the overall server payload. (BEA Weblogic, VMWare ESX, XenSource XenEnterprise)
  • Layer 2 Networking: In order to use a server to meet an application's needs, that server must have access to required networks. True automation requires that switch ports be reconfigured as necessary to ensure access to specifically the VLANs required by the payloads they will represent. (Examples: Cisco 3750, Extreme Summit400)
  • Network Attached Storage (NAS): The beauty of NAS devices is that they can be dynamically attached to a software payload at startup, without requiring any hardware configuration beyond the Layer 2 configuration described above. SAN is also useful (and common), but requires hardware configuration to make work. That complicates the role of automation. Part of the problem is the inconsistent remote configurability of fiber switches, which may be mitigated somewhat with iSCSI. However, NAS is quickly becoming the preferred storage mechanism in large data centers. (Examples: NetApp FAS, Adaptec Snap Server)

Over time, I see the industry adding more and more "drivers" to manage more and more data center (and perhaps desktop) resources. Imagine a world in which each software and/or hardware vendor produced standard SLA drivers for each individual component that makes up your data center environment. Every switch, disk and server; every service, container and OS; even every light bulb and air conditioner are connected to a single service level policy engine in which business policy (including cost of operations) drives automated decisions about their use and disuse.

Its not here yet, but you won't have to wait long...

I will use the label "respond" to tag posts related to response interfaces.

Wednesday, May 23, 2007

If your CIO doesn't "get it" yet, he soon will

Ken Oestreich posted a cool review of the Forrester Research 2007 IT Forum in Nashville, TN last week. I was fascinated by Forrester's contention that there are "no IT projects anymore, only business projects" (aka "BT"), but perhaps the most interesting observation of the day was the following:

However, the best talk IMHO came from Robert Beauchamp, CEO of BMC software. He's a very down-to-earth, articulate guy-even in front of 1,000 people. I was most impressed by his Shoemaker's Children analogy... that the IT (alright, BT) organizations in enterprises are arguably the least automated departments around. ERP is automated. Finance is automated. Customer interaction is automated. But IT is still manually glued-together, with operations costs continuing to outpace capital investments.

(Emphasis mine.)

This is a gorgeous observation; so simple, so articulate, and--most importantly--so true! I have always been amazed at the amount of manual labor that goes into delivering technology that makes some other schmuck's life more labor free. Programming is a great example of this. (Even with advances in code building, IDE templating and wizard-based programming, I bet the vast majority of developers out there still shudder at the term "code generation".)

Server provisioning (bare metal or virtual servers) is also a great example. In a prior life, I worked for one of the most forward thinking technology companies out there. However, when it came to pushing code to production, it was still a server-by-server hand install job. Provisioning 4 front end portal servers took anywhere from a couple of hours to a couple of days.

Another example is trouble ticket response. How many of the system operators out there still carry pagers, and are forced to get out of bed in the middle of the night to respond to a system event? If you say "not at my company", I bet you have overseas support to back you up overnight. The response remains completely manual.

That is why I am so excited about the Service Level Automation space, its role in utility computing and its role in automating IT processes. It is time this happens, and I hope your "BT" organization is considering it.

Monday, May 07, 2007

VMWare TSX and Reducing Complexity in the Data Center

I had a busy last few days of the week last week. On Wednesday, I attended VMWare TSX in Las Vegas, and on Friday, I had the chance to hear Bill Coleman speak about utility computing, and the events that have lead up to its sudden resurgence in the market place.

All in all, TSX was one of the most informative VMWare events I have ever attended. I only had the chance to attend three sessions--CPU scheduling, ESX networking and DRS/HA--but all three were packed with useful information. (The slides linked here are from a TSX conference in Nice, Italy, April 3-5, 2007. They are a little different from the slides I saw, but are similar enough to communicate the basic concepts.)

If you don't know much about VMWare CPU scheduling, check out that deck. Sure, its basic scheduling stuff, but it is very helpful when it comes to understanding how VMWare settings affect processor share. The networking deck is also critical if you must deploy network applications to virtual machines.

The DRS/HA deck has some helpful tips, but also clearly demonstrates the limited scale of DRS/HA. A 16 physical node limit per HA cluster, for example, is going to be problematic for most medium to large data centers. Furthermore, these are very server-centric technologies; the concept of Service Level Automation is clearly missing, as there is no concept at all of a service or application to be measured. They are hinting at a few new app-level monitors in a later release, but I just don't think monitoring service levels from a business perspective is very important to VMWare.

Bill's speech to the IT department of a large manufacturer was very interesting, if for no other reason than it clearly spelled out the argument for reducing complexity in the data center. (For a quick and dirty argument, see this article.) We are definitely at a cross roads now; IT can choose to attack complexity with people or technology. Most of us are betting technology will win. Furthermore, Bill told the assembled techies, the early adopters of any platform technology get the best jobs when that platform becomes mainstream. Almost nobody is predicting that utilty computing will fail in the long term, so now is the time to jump aboard and get involved.

I should have some time to complete the Service Level Automation Deconstructed series this week. Stay tuned for more.

Monday, April 30, 2007

What do SOA and EDA have to do with SLA?

I've been launching off of Todd Biske's blog roll into the world of SOA and EDA blogging. I'm actually kind of saddened that my voyage into the world of infrastructure automation has pulled me so far from a world in which I was an early practitioner. (My career at Forte Software introduced me to service-oriented architectures and event-based systems long before even Java took off.) I love what the blogging world is doing for software architecture (and has been doing for some time now), and I feel like a kid in a candy store with all the cool ideas I've been running across.

One blog that has been capturing my interest is Jack van Hoof's "SOA and EDA". I love a blog with real patterns, term definition, and a passion for its subject matter. All put together by someone who can get an article published.

The article is actually very interesting to me from a Service Level Automation perspective. Jack captures his thoughts on the importance of building agile software architectures in the following paragraph:
Everything is moving toward on-demand business where service providers react to impulses - events - from the environment. To excel in a competitive market, a high level of autonomy is required, including the freedom to select the appropriate supporting systems. This magnified degree of separation creates a need for agility; a loose coupling between services so as to support continuous, unimpeded augmentation of business processes in response to the changing composition of the organizational structure.

(Emphasis mine.)

The only thing I would change about Jack's statement above is replacing the words "a loose coupling between services" to "a loose coupling between services and between services and infrastructure" and changing "composition of the organizational structure" to "composition of the organizational structure and infrastructure environment". (Some may have issues with the latter, but I don't mean that services should be written with specific technology in mind--just the opposite; they should be written with an eye towards technology independence.)

This is why I have been emphasizing lately the need to view the measure activity through the lens of both business and technical measures. Some of the business events thrown by an EDA may very well indicate the need to change the infrastructure configuration (e.g. if the stock market sees a 20% rise in volume in the matter of three minutes, someone may want to add capacity to those trading systems). However, the technical events from a software system (e.g. thread counts or I/O latency) may also indicate the need to change infrastructure configuration on the fly.

I wish I could spend more time collaborating with SOA architects and "tacticians". In fact, I have been speaking with Ken Oestreich about exactly this. If you are in the SOA space, and interested in talking about how SOA, EDA and SLA interconnect, let me know by commenting below. (Be sure to let me know how to contact you.) At the very least, think about how infrastructure will measure the performance of your software systems as you start your next development iteration.

Thursday, April 26, 2007

Agile Computing Catches Up to the Data Center

What is the biggest hurdle to adopting Service Level Automation--or even dynamic computing in general--in today's data centers?

Is it technology? Nope. Most servers, storage and even network equipment can be managed reasonably well today by several vendors, with varying degrees of dynamic, policy-based provisioning. Several critical monitoring interfaces are also now standard in everything from power controllers to OSes to applications.

Is it infrastructure architecture? Not really, with one caveat. As long as an architecture has been built from the ground up to be easily managed and changed, with real attention paid to dependency management and virtualization where appropriate, most data centers are excellent candidates for automation. Which is a small leap away from utility computing.

Is it software architecture? Nope. I talked about this before, but SLA systems are just your basic event processing architecture specialized to data center resource optimization. The really good ones (*ahem*) can do this without adding any proprietary agentry to the managed software payload. In other words, what ends up running in your data center is almost exactly what you would have run without automation. There is little evidence on the application host that it is being managed at all.

Then what is it? One word: culture. The overwhelming obstacle that I see in the data center market today is fear of rapid change.

It is true of the sys admins, though they get the value of automation right away. They just need to see everything work before they trust it.

Its true of the storage admins, though storage virtualization is gaining ground. Unfortunately, this doesn't yet translate to accepting constant and sometimes rapid, somewhat arbitrary change within their domain.

It is most true of the network guys. Networks are the last bastion of the relatively static "diagram", mapping each component of the network architecture exactly with an eye to controlling change. The idea of switching VLANs on the fly, reconfiguring firewalls on demand, or even not knowing which server is assigned which IP address without looking at a management UI is scary as hell for the average network administrator.

And who can blame any of them? The history of commodity computing in data centers is littered with bad results from untracked changes, or badly managed application rollouts. Add to that the subconscious (of even conscious) fear that they are being replaced by software, and you get staunch resistance to changing times.

What everyone is missing here, though, is the key differentiation between planning for change, and executing it. No one in the entire industry is arguing that data center administrators should stop understanding exactly how their data centers work, what can go wrong, and how to mitigate risk. Cassatt (and I'm sure its competitors) spends significant time with each customer, even in pilot projects, making sure the data center design, software images, and service level definitions result in well understood behavior in all use cases.

But once those parameters are defined, and the target architectures, resources and service levels are defined, its time to let a policy-based automation environment take over execution. A Service Level Automation environment is going to make optimal decisions about resource allocation, network and storage provisioning and event handling, and do it in a fraction of the time that it would take a single human (let alone a team of humans). And, as noted above, once provisioning takes place, the applications, networks and storage run just as if a human had done the same provisioning.

(By the way, none of this breaks with ITIL standards. It just moves execution of key elements from human hands to digital hands. It also requires real integration between the SLA environment and configuration management, asset management, etc.)

All of this reminds me of the paradigm shift(s) that the software development industry went through from the highly planned, statically defined waterfall development methods of the early years to the always moving, but always well defined world of agile development methodologies. Its been painful to change the software engineering culture, but hasn't it been worth it for those that have found success? And, isn't it absolutely necessary for the highly decoupled and interdependent world of SOA?

Data center operations is about to undergo the same pain and upheaval. Developers, be kind and help your brethren through the cultural shift they are experiencing. Perhaps some of you in the agile methods field can begin to work out variations of your methods for data center planning and execution? Perhaps we should integrate data center planning activities into our "product-based" approaches?

Are you ready for this shift? Is your organization? What can you do today (architecturally and culturally) to ready your team for the coming utility computing revolution?

Friday, April 20, 2007

Service Level Automation Deconstructed: Analyzing Service Levels

This is the second post in my series providing a brief overview of the three critical assumptions of a Service Level Automation environment. Today I want to focus on the ways in which the metrics gathered from the "measure" capabilities of an SLA environment are evaluated to determine if and what action should be taken by the "response" capabilities.

Let me first acknowledge that my discussion of the measure capabilities included some analysis of simple metrics to create complex metrics. This is one piece of the analysis puzzle, and is a critical one to acknowledge. Ideally, all software and hardware systems would be designed to intelligently communicate the metrics that matter most to determine service levels. Where this consolidation occurs depends on the requirements of the environment:
  • Centralized approach: Gather fundamental data from target systems to central metrics processor and consolidate metrics there. The advantage here is having one place to maintain consolidation rules. The disadvantage is increased network traffic.
  • Decentralized approach: Gather fundamental data and do any analysis necessary to consolidate the fundamental data into a simplified composite metric there. Send the composite metric to the core rules engine (which I will discuss next).

Metrics consolidation is not really the core analytics function of a Service Level Automation architecture, however. The key functions are actually the following:

  • Are metrics being received as expected? (A negative response would likely indicate a failure in the target component or the communication chain with that component)
  • Are the metrics within the business and IT service level goals set for that metric
  • If metrics are outside of established service level goals, what response should be taken by the SLA environment

Given my recent reading into complex event processing (CEP), this seems like at least a specialized form of event processing to me. The analysis capabilities of an SLA environment must constantly monitor the incoming metrics data stream, look for patterns of interest (namely goal violations, but who knows...) and trigger a response when conditions dictate.

The great thing about this particular EP problem is that well designed solutions can be replicated to all data centers using similar metrics and response mechanisms (e.g. power controllers, OSes, switch interfaces, etc.). Since there are actually relatively few components in the data center stack to be managed (servers [physical, virtual, application, etc.], network and storage), the rule set required to provide basic SLA capabilities is replicable across a wide variety of customer environments.

(That's not to say the rule set is simple...its actually quite complex, and can be affected by new types of measurement and new technologies to be managed. Buy is definitely preferred over build in this space, but some customizability is always necessary.)

Finally, I'd also like to point out that there is a similar analysis function at the response end as at the measure end. Namely, it is often desirable for the response mechanism to take a composite action request and break it into discrete execution steps. The best example I can think of for this is a "power down" action sent from the SLA analysis environment to a server. Typical power controllers will take such a request, signal to the OS that a shutdown is imminent, whereupon the OS will execute any scripts and actions required before signalling that OS shutdown is complete. At that time, the power controller turns off power to the server.

As with measure, I will use the label "analyze" to reflect future posts expanding on the analysis concept. As always, I welcome your feedback and hope you will join the SLA conversation.

Monday, April 16, 2007

Two articles mentioning Service Level Automation

I recently set up Google Alerts to inform me about references to Service Level Automation on the web. There were many articles returned this week, (many of which involved Cassatt), but I found two additional articles of note. Each makes mention of Service Level Automation, and represents the growing interest in this approach.

The first is from the March 2006 issue of ACM Queue, entitled "Under New Management". The article was written by Duncan Johnston-Watt, the founder of Enigmatec. Johnston-Watt does an excellent job of outlining basic issues around one possible architecture for an autonomic data center. As expected for Enigmatec, its a policy automation focused approach, and is, in fact, one of the few articles from a policy engine vendor that I have see where the term Service Level Automation is used correctly.

Unfortunately, I don't necessarily agree that Johnston-Watt's architecture is optimal enterprise data centers. (It requires development of process automation flows to "optimize operational processes"--a significant amount of work that is prone to introducing new inefficiencies. It is also agent based, which alters the footprint of the software stacks being run in the data center, and can negatively affect the execution and architecture of the applications being managed.) All in all, though, there is some excellent information here for those thinking about Service Level Automation holistically, across the entire data center.

The other article, entitled "Virtually Speaking: Xen Achieves Higher Enterprise Consciousness" was published April 6, 2007 on ServerWatch. In the last few paragraphs of the article, uXcomm's aquisition of Virtugo is covered. In it, uXcomm claims the combination their Xen management tools and Virtugo's VMWare tools "fills a gap not just in uXcomm's portfolio but in the virtual landscape as well. Until now [...] there was a gap between service-level automation offerings and performance management products."

Hmmm. Not sure how providing SLA for only virtual servers counts as filling gaps...but, even so, I hope uXcomm is aware that everyone in this space realizes the need for resource optimization includes VM performance management. I guess my question would be, what is uXcomm doing about marrying Service Level Automation to the rest of the data center?

As a side note, I know that I owe two more articles on my Service Level Automation Deconstructed series. I am working on the "Analyze" overview now, but have discovered some interesting technology to discuss here that I am reading up on now.

Wednesday, April 11, 2007

Complexity and the Data Center

I just finished rereading a science book that has been tremendously influential on how I now think of software development, data center management and how people interact in general. Complexity: The Emerging Science at the Edge of Order and Chaos, by M. Mitchell Waldrop, was originally published in 1992, but remains today the quintessential popular tome on the science of complex systems. (Hint: READ THIS BOOK!)

John Holland (as told in Waldrop's history) defined complex systems as having the following traits:
  • Each complex system is a network of many "agents" acting in parallel
  • Each complex system has many levels of organization, with agents at any one level serving as the building blocks for agents at a higher level
  • Complex systems are constantly revising and rearranging their building blocks as they gain experience
  • All complex adaptive system anticipate the future (though this anticipation is usually mechanical and not conscious)
  • Complex adaptive systems have many niches, each of which can be exploited by an agent adapted to fill that niche

Now, I don't know about you, but this sounds like enterprise computing to me. It could be servers, network components, software service networks, supply chain systems, the entire data center, the entire IT operations and organization, etc. What we are all building here is self organizing...we may think we have control, but we are all acting as agents in response to the actions and conditions imposed by all those other agents out there.

A good point about viewing IT as a complex system can be found in Johna Till Johnson's Networld article, "Complexity, crisis and corporate nets". Johna's article articulates a basic concept that I am still struggling to verbalize regarding the current and future evolution of data centers. We are all working hard to adapt to our environments by building architectures, organizations and processes that are resistant to failure. Unfortunately, entire "ecosystem" is bound to fail from time to time. And there is no way to predict how or when. The best you can do is prepare for the worse.

One of the key reasons that I find Service Level Automation so interesting is that it provides a key "gene" to the increasingly complex IT landscape; the ability to "evolve" and "heal" the physical infrastructure level. Combine this with good, resilient software architectures (e.g. SOA and BPM) and solid feedback loops (e.g. BAM, SNMP, JMX, etc.) and your job as the human "DNA" gets easier. And, as the dynamic and automated nature of these systems gets more sophisticated, our IT environments get more and more self organizing, learning new ways to optimize themselves (often with human help) even as the environment they are adapting to constantly changes.

In the end, I like to think that no matter how many boneheaded decisions corporate IT makes, no matter how many lousy standards or products are introduced to the "ecosystem", the entire system will adjust and continually attempt to correct for our weaknesses. In the end, despite the rise and fall of individual agents (companies, technologies, people, etc.), the system will continually work to serve us better...at least until that unpredictable catastrophic failure tears it all down and we start fresh.

Monday, April 09, 2007

SOA blog recognizes SLA!

I was analyzing the Google Analytics for my blog recently, and noticed a small but measurable amount of traffic originating from "biske.com". Now, my Analytics are really never very exciting, so seeing an unknown domain explicitly called out was too much to ignore. I followed the domain, and arrived (after seeing a very nice family photo) at Todd Biske's Outside the Box blog.

Now, I don't know Todd from Adam, so I took some time to read through a few of his posts. His subtitle explains his focus very well: "SOA, BPM, and other strategic IT initiatives". His recent posts included coverage of the Web Methods acquisition by Software AG (I wonder what he thinks about the BEA/Amberpoint announcement), the vagrancies of developing SOA clients and services in parallel (a problem I know quite well, actually, from my past as an architect with Forte Software and Sun Microsystems), and the rapid integration of management and monitoring across all tiers of enterprise architectures (a post I want to specifically address in a later post of my own). From what I read so far, he is a fairly holistic enterprise architect with a good eye for both functional and infrastructure issues.

But why was I getting references to my site from his? I kept looking, and to my delight rested my eyes on my name in his blogroll...WOO HOO!!! The first such link that I know of!

By the way, follow some of those other blogroll links...there you will find a huge wealth of knowledge about SOA, BPM and the culture shock the distributed systems community is feeling as a result (with some calm, collected voices providing solid advice). Don't think SOA and SLA are related? You'll find lots of evidence to the contrary among the posts of Todd and his colleagues.

Thanks, Todd, and you've been added to my daily RSS feed.

Thursday, March 29, 2007

Service Level Automation Deconstructed: Measuring Service Levels

This is the first of three in my series analyzing the key assumptions behind Service Level Automation. Specifically, today I want to focus on the measurement of business systems, and the concepts behind translating those measurements into service level metrics.

Rather than trying to do an exhaustive coverage of this topic (and the other topics in this series) in a single post, what I am going to do is provide a "first look" post now, then use labels when followup posts have relative information. The label for this topic will be "measure".

In my next installment, I'll introduce analysis of those metrics against service level objectives (SLO) the business requires. That post, and future related posts will be labeled with "analyze".

In the final installment of the series, I'll describe the techniques and technologies available to digitally manipulate these systems so that they run within SLO parameters. Posts related to that topic will be labeled "respond".

As noted earlier, my objective is to survey the technologies, academics, etc., of each of these topics in an attempt to enlighten you about the science and technology that enables service level automation.

How do we measure quality of service?

Measuring quality of service is a complex problem, not so much because it is hard to measure information systems and business functionality. I (and I bet you) could list dozens of technical measurements that can be made on an application or service that would reflect some aspect of its current health. For example:

  • System statistics such as CPU utilization or free disk space, as reported by SNMP
  • Response to a ping or HTTP request
  • Checksum processing on network data transfers
  • Any of dozens of Web Services standards

The real problem is that human perception of quality of service isn't (typically) based on any one of these measurements, but on a combination of measurements, where the specific combination may change based on when and how a given business function is being used.

For example, how do you measure the utilization of a Citrix environment? Measuring sessions/instances is a good start, but--as noted before with WTS--what happens when all sessions consume a large amount of CPU at once? CPU utilization, in turn, could fluctuate wildly as sessions are more or less active. Then again, what about memory utilization or I/O throughput? These could become critical completely independently from the others already mentioned.

No, what is needed is more mathematical--one (or a couple of) index(es) of sorts generated from a combination of the base metrics retrieved from the managed system.

There are tools that do this. They range from the basic capabilities available in a good automation tool, to the sophisticated evaluation and computation available in a more specialized monitoring tool.

What I am still searching for are standard metrics being collected by these tools, especially industry standard metrics and/or indexes that demonstrate the health of a datacenter or its individual components. I'll talk more about what I find in the future, but welcome you to contribute here with links, comments, etc. to point me in the right direction.

Monday, March 26, 2007

Service Level Automation is green--when done right

Vinay Pai has a post today that I think spells out one of the key benefits of Service Level Automation. Remember, the key aspect of SLA is knowing what capacity is required to maintain target operational goals. A basic SLA environment will assure that only the necessary capacity is active at any time.

So, if capacity is not being used at any given point in time, why have it consume power or cooling at all? Some "automation" products require underused physical infrastructure to remain running in order to support their management layers--just in case. This is unfortunate, as an idle hypervisor is just idle capacity. Its not serving a current business need.

A truly efficient SLA platform is aware of the power controller states of each of its physical servers, and can power down unused servers. Servers are only turned on when they are needed to meet some aspect of the system's service level goals.

Note Vinay's description of the QA labs at Cassatt. As you might expect, he pushes the limits of what a SLA datacenter must endure, yet can always scale his power and cooling needs to his current workload. Can you say that about your datacenter?

Thursday, March 22, 2007

Service Level Automation Deconstructed: Introduction

Service Level Automation starts with three simple premises:

* The factors contributing to software service quality can be measured electronically.

* Runtime targets indicating high quality of service can be defined for those measurements.

* Systems involved in delivering software functionality can be manipulated to keep those measurements within the runtime targets.

I think the support for each of these premises should be explored more deeply, so I plan to begin a little survey of the technologies and academics over the next few weeks. The idea is to get a good sense of what standards/technologies/concepts/etc. can be used to meet the requirements of each premise. I also hope to discuss how a system smart enough to take advantage of them(*) can save a large datacenter both in terms of direct costs, as well as in losses due to service level failures.

Why Service Level Automation? I wrote about this earlier. However, as a quick reminder, think of service level automation as meeting this objective:

Delivering the quantity and quality of service flow required by the business using the minimum resources required to do so.

I've been quite busy both at work and at home, so I'm hoping to use this exercise as a way to increase my posting frequency. Stay tuned for more.

Wednesday, March 21, 2007

5 things...

The latest blogosphere social phenomenon has reached me doubly this week. Both Ken Oesterich and Ken Wallich have tagged me in the ongoing "5 things tag" that has been sweeping the blogging community (especially the tech bloggers).

Here are five things most people do not know about me:

  1. I was born in Reading, England.
  2. I play pretty decent guitar. I don't know that many songs by other people (a problem when whipping out the guitar at parties), but I have several original works that I think hold their own very nicely against most pop drivel. Lately, however, I have been working on "Tears in Heaven" by Eric Clapton.
  3. I played Mr Anthrobus in Thornton Wilder's "The Skin of our Teeth" in high school. I was a geeky, awkward teenager trying to play a 40 year old man, and was the only member of the primary cast not to win an award for my performance in that show. Now that I am 40, I wonder what the hell was so hard...
  4. My computing career started in fifth grade in Cedar Rapids, Iowa. I was lucky enough to get in a science focused program at a nearby elementary school, and one of the kids' moms was one of the first BASIC programmers at Rockwell Collins, the aviation electronics firm. She came to our school once a week and taught us the basics of variables, loops, conditional statements and subroutines. Very cool. I got caught a bunch of times programming on the teletype terminal in the back of the classroom while I should have been listening to the teacher. Later, my luck continued as the father of one of my close neighborhood friends bought the fifth (or something like it) Apple II computer in the state of Iowa. We would program in BASIC every day after school, and tried to get into writing games and such.
  5. Later, in college, I was determined to be a Music/Computer Science double major...for all of one semester. I didn't practice the music stuff enough, so I got a low grade there, and I hated my systems organization class, so I lost interest in computer science. (Dumb reason, now that I look back, but it worked out.) Instead, I started taking every math and physics class that I could, and finished with a Mathematics/Physics double. The day of graduation, I swore to my friends "I will NEVER be a computer programmer for a living". Two and a half years later, I was coding C for a small manufacturing company. (Do not try to predict the future, even your own. Its pointless. Setting goals is OK, but be willing to float a bit with the breeze.)

Now, let me please introduce to you five more randomly selected from my blogosphere:

  • My mom.
  • Katie Tierney, a former collegue with excellent technical intuition who is proving herself to be a hell of a "head of household" as well.
  • Rama Roberts, another former collegue whose blog never fails to entertain and enlighten.
  • Management guru, Tom Peters, who reenforces my drive to amaze both my employers and customers by being a service professional first and foremost.
  • Alessandro Perilli, author of the virtualization-focused virtualization.info blog.

Monday, February 26, 2007

When CPU utilization is not enough...

Where have I been, you might ask?

Much to my delight, the last few weeks have been filled with customer activity, ranging from helping a Service Level Automation-enabled appliance for a major software company, to assisting the financial wing of one of the world's largest manufacturers to experience first hand the benefits of utility computing.

The latter runs an application that is highly dependent on Windows Terminal Services to deliver a client-server UI to thousands of retail outlets world-wide. Uptime is critical to this application, as customers will go elsewhere for financing if this application doesn't confirm credit within minutes of a purchase decision.

Unfortunately, WTS is also a very inefficient consumer of server payload. It is a session-based infrastructure, which means that a user will be attached to a specific server for their desktop access until they either log out or are kicked off. If 15 user sessions share a physical server, there is no way to predict the load on the system. All 15 sessions could be idle, or all could quickly start consuming cycles simultaneously.

This gives me my first really good example of when CPU and/or memory utilization are not good Service Level metrics on their own. Imagine an environment using WTS to support hundreds of users. These users use their Windows sessions to run a variety of tasks, much like any Windows user. Some tasks use a high level of CPU and memory, others very little. Quite often, the session will sit idle for several minutes.

Now, if you create lots of sessions because the CPU is idle, you could end up with problems if they all get active at the same time (say right after lunch). If you stop creating sessions on a server because CPU utilization is high, you may end up with a highly under utilized server when one user's game of Quake wraps up.

That's not to say that CPU or memory utilization aren't an important part of the Service Level "equation". The truth is, there are several metrics that apply to WTS capacity: sessions, CPU utilization, memory utilization, licenses, etc. Since determining Service Level compliance probably involves evaluating the relationships between several of these metrics at once, there will most likely be one or two compound metrics based on mathematical equations combining these "root" metrics in a way that reasonable thresholds can be set.

Another interesting observation is that this is a lot like the Java EE Service Level Automation problem. (Thanks to Luis Cuyun for pointing this out.) While most horizontally scalable application tiers can be scaled up and down as a unit (i.e. "add a node/remove any node"), app servers, hypervisors and (now) WTS all must be monitored as a unit, but managed on a per server basis (in this case, "add a node/remove this specific node"). This is because the "instances" that each of these software resources are hosting are "sticky" to a server (VMotion not withstanding), and you don't want to shut down any server when capacity is not longer needed, you want to only remove the specific servers with no live sessions remaining on them.

(Speaking of VMotion, one of the things that both Java EE and WTS will require to be really optimizable is the ability to move live services/sessions from one server to another in real-time. Anyone know of a technology addressing either of these?)

The good news for a good Service Level Automation environment (*ahem*) is that if one of these problems (Java EE, virtual servers or WTS) is solved correctly, the same basic technology can be applied to all of them. That's not to say that anyone is doing this for WTS today (to my knowledge, no one is), but I like the idea that the use cases that apply to Java EE SLA also apply to WTS SLA.

I'm hoping to have more to write about this as this pilot continues. In the meantime, anyone with Windows experience is welcome and encouraged to contribute their two cents to this discussion. In particular, are there any tried and true service level metrics for WTS that are being used out there? In general, there are so many moving parts here, that I am sure there are many critical factors to Service Level Automation of WTS that I have not covered, or even considered.

Thursday, February 01, 2007

Welcome Vinay!

Vinay Pai, Cassatt's truly talented Director of Product Development, has joined the blogging fray. He starts with an excellent post on his recent experiments with completely provisioning 400+ servers in just a few hours without having an extensive system administration background. Vinay's blog should be a fun read for those of you who are considering Cassatt (or, at least, service level automation and/or utility computing infrastructure). His day to day job has him getting "down and dirty" with both live customer installations and extreme test scenarios.

Welcome, Vinay! I look forward to the good read.

Great Blogs Unite!

What's that you say? You want a single URL where all of the best minds blogging about service level automation, dynamic datacenters and utility computing come together in a single, easy to read package? Well, your wish is granted! Most of the blogs referenced in my "Links" section are gathered there in real time via RSS, with the latest post always at the top, so you always get the freshest insight into next generation IT architecture, technology and culture...

Completely unbiased, of course! :D