(Formerly "Service Level Automation in the Datacenter")
Cloud Computing and Utility Computing for the Enterprise and the Individual.
Sunday, July 08, 2007
Web 2.0, Utility Computing and Service Level Automation
I have to say I agree with her, but I would like to observe that this is only a (large) piece of the overall picture. In reality, most of the sites she specifically mentions are actually "Software as a Service" sites targeted at consumers rather than businesses. Its the trend towards getting your computing technology over the Internet in general that is the real driving need.
For "utility computing" plays such as SaaS companies, managed hosting vendors, booksellers :), and others, the need for utility computing isn't just the need to find capacity, it is also the need to control capacity. This, in turn, means intelligent, policy-based systems, that can deliver capacity to where it is needed, share capacity among all compatible software systems and meter capacity usage in enough detail to allow the utility to constantly optimize "profit" (which may or may not be financial gain for the capacity provider itself).
Service Level Automation, anyone?
Web 2.0 drives utility computing, which in turn drives service level automation. So, Industry Girl, I welcome your interest in utility computing, and offer that the extent to which utility computing is successful is the extent by which such infrastructure delivers the functionality required at the service levels demanded. Welcome to our world...
Sunday, June 24, 2007
God in the Grid...
By the way, I've been meaning to preorder Nicholas Carr's book as well. He has a web site that explains the book's premise. If you want to begin dreaming about a world where computing costs very little, and automation, virtual societies and electronic economies abound, there seems to be plenty of inspiration here.
I just hope it's as good as it sounds...
Wednesday, June 20, 2007
Clarification on James McGovern's recent comments
I wanted to take a moment to clear up some misunderstandings in James McGovern's comment on my recent third-hand coverage of the Forrester Research 2007 IT Forum conference (via Ken Oestreich who actually attended). I have huge respect for James' pragmatic, get to the point style, and he is definitely a leading edge player in both the Enterprise Architecture and Security spaces. However, James mistakenly credits the comment, "There are no IT projects anymore, only business projects", to "Ken Oestreich of Forrester". Ken is not an employee of Forrester Research (he is employed by Cassatt as Director of Product Marketing), and he was just reporting what he heard. I am not aware of the name of the Forrester Analyst who made that statement.
Having said that, clearly this quote can be interpreted in a variety of ways. James seems to have read the quote as "IT is dead" or some such thing, where I think the quote is intended to show that IT projects are now founded initially on business need, not on the whims of IT professionals eager to introduce new technologies. As I read James' blog, this seems right up his alley. However, I would agree that the statement is more hype than substance--thus the lack of a specific interpretation.
One other thing James says bugs me a bit. Perhaps I misinterpreted, but the following statement seems way off:
"Maybe he should acknowledge that the vast majority of CIOs aren't even focusing on data centers as this has been commotitized (sic) and therefore pushed several layers down in the organization. Likewise, infrastructure stuff simply doesn't allow an enterprise to either innovate nor sustain competitive advantage where as software development still has the potential for both."
Really? I think if James were to ask his CIO what his number one expense line item was, it wouldn't be research and development. Wonder why utility computing is a key initiative for CIOs in the coming years? Why is "Green Data Center" all over the press these days? I can tell you: operations (labor, facilities, infrastructure and utilities) accounts for the vast majority of enterprise IT budgets. Even those that "outsource" (domestically or otherwise) to Managed Hosting Providers are paying a princely sum for the service.
Add to that the "siloed" nature of most application deployments in data centers, and the incredible barrier to market agility that causes. A giant "bank of our continent" recently filled an RFP for utility computing (in the "turn our IT into a utility" sense), but cost savings wasn't even the most important factor. The bank can only grow through (international) acquisition, and each acquisition has been burdensome largely due to the service level losses caused by integrating each IT organization and its infrastructure. Agility with guaranteed service levels is the bank's number one operations priority.
That's not to say that software isn't also a priority. It is for all the reasons that James alludes to. It certainly is the quickest (only?) route to new revenue streams, and it can also lead to significant cost savings if done right. Hell, if we can get the cost of operations down, it will free up more funds for this important endeavor! But to say infrastructure has been pushed way down the list just doesn't jibe with the pain we are finding in corporate IT today.
Let me make it clear that I have great respect for James McGovern, and I read his blog every day. I hope that we can continue a conversation about both the role of infrastructure innovation in the future of IT, and the relationship between application architectures and their deployment architectures in the data center.
Sunday, June 03, 2007
Want to save gas? Stop leaving your car idling in the garage!
I remember the seventies, when gas prices skyrocketed (the first time) and there were suddenly all these tiny cars on the road. One member of my mom's congregation even showed up one Sunday with this crazy little car that ran on a motorcycle engine. It was made by some new car company called Honda, and it was one of the first years that Civics were sold in America.
As a nation, we clamored to change our lifestyles--ditching heavy steel muscle cars for sporty (or utilitarian) little "economy cars". Our approach to solving the energy crisis was to increase the efficiency with which our cars consumed energy. Note, however, it was not (by a long shot) to reduce the amount of driving we did.
Now flash forward to today, and take a look at the current energy crisis in America's (and the world's) data centers. Electricity is expensive, and growing more so (except for those lucky enough to have subsidised power). Add to that concerns about global climate change, and you've got company after company scrambling to be "green".
Again, however, note that the target is not to do less computing than we did before. In fact, if anything, the demand is increasing for information technology and business automation. I believe pushing the automation envelope is going to take more computing power than we know.
So, like the automobile vendors of the seventies, today's systems vendors are working hard to release "energy efficient" models of servers, laptops and desktops. They do this ostensibly to give us all a good feeling about what good stewards of our tiny planet we are, but in reality its all about saving money. None of this changes our worst behaviour, however; our tendency to leave as much capacity running as possible at all times, "just in case".
Of course, the server that uses the least amount of power is the one that is turned off. That's where Service Level Automation comes into the picture. As noted in the past, one of the key aspects of a good Service Level Automation platform is the capability of shutting down anything that isn't serving an immediate business need. Traditionally, I've always talked about this in relation to scale-out applications--your SLA platform should shut down servers not needed to meet current demand in such applications. Now, however, I want to talk about three use cases where SLA enhances the day to day power consumption of all applications in the data center.
- Job-specific management. OK, think of every server you've touched in the last six months. How many of those served a short term purpose (e.g. getting a software release out the door), but frequently spend days unused for any purpose. I remember going days or even weeks between placing builds on staging servers in my previous life. Service Level Automation should be able to detect unused software payloads, and shut down that equipment until needed again by that or any other payload.
- Time-specific management. Almost every data center (especially development and test labs) have systems that are hit hard during some portion of the day, then remain idle for the remainder. SLA should provide the capability to not only schedule system shutdowns, but to actually look at that status of systems to determine which are best candidates for shutdown. In other words, go beyond automating "blind" scheduled events to delivering intelligent management of system power cycles.
- Power emergency management. One of the great benefits of living in the San Francisco Bay area is the incredible ingenuity of our power utility in encouraging companies to conserve power and "be good neighbors" in a power emergency. PG&E offers rebates to companies willing to join Demand Response programs, where they agree to voluntarily reduce electric consumption to help the utility avoid the infamous "rolling blackout".
The Silicon Valley Leadership Group has recently been hosting a series of events around "Energy Efficient Data Centers", one of which targeted how SLA could deliver on all three of the above. The response was tremendous--so much so that my employer has asked me to join a team building a simple targeted solution to these problems based on our already innovative SLA platform. I can't say much more right now, but I certainly will communicate all that I can as soon as I can.
By the way, the first lesson I've learned from all of this is that power measuring capabilities varies widely from data center to data center. Some companies can't tell you anything more than their monthly bill, others can show you power consumption over time at the individual server level. Part of the issue is that there are no "simple" power metering solutions at the server level...power controllers (i.e. iLO2) are just now starting to give management systems access to the power measurement tools on Intel and AMD boards. MPDUs have some good features, but they vary widely from vendor to vendor.
You can't control what you can't measure, so get on board system vendors! Give us the tools we need to measure and manage those beautifully efficient next generation servers. Heck, give us the tools we need to measure and manage all those older systems we have out there now. That would be more green by far than just squeezing another milliamp out of a MIP.
Thursday, May 24, 2007
Service Level Automation Deconstructed: Respond
For the third and last in my series breaking down the three key assumptions behind Service Level Automation, I would like to focus on how SLA environments can control data center configuration in response to service level goal violations. These goal violations and the high level actions to be taken are determined by the analysis capability of the environment. Details of how to accomplish those high level actions, however, are decided and executed by the response function.
Essentially, the response function of an SLA environment is very much like the driver set that your operating system uses to translate high level actions (e.g. "store this file") to device specific actions ("Move head 32 steps to center, find block 4D5EF, etc."). The responsibility here is to provide the interface between the SLA analysis engine and specific standard or proprietary interfaces to everything from server hardware to network switches to operating systems and middleware.
I see the following key interface points in today's environments:
- Power Controllers/MPDUs: Job 1 of a service level automation environment is providing the resources required to meet the needs of the software environment, and only those resources. Turn those servers on when they are needed, and off when they are not. This includes virtual server hosts. (Examples: DRAC, iLO, RSA II,MPDUs)
- Operating Systems: Before you shut off that server, make sure you've "gently" shut down its software payload. Well written server payloads for automated environments will both start up and aquire intial state (if any), and shut down while preserving any necessary state without human intervention. However, from a communications perspective, each action starts with the OS. (Examples: Red Hat, SuSE, MS Windows, Sun Solaris)
- Middleware/Virtualization: It is interesting to note that many software payload components (e.g. an application server or a hypervisor) are both software to be managed, and computing capacity themselves. For example, an application server should be managed to specific service levels relating to its relationship with its host server (e.g. CPU utilization, thread counts, etc.), while also treated as a capacity resource for JavaEE applications and services. As such, these software containers should be managed for their own guest payloads much like a physical server would for the overall server payload. (BEA Weblogic, VMWare ESX, XenSource XenEnterprise)
- Layer 2 Networking: In order to use a server to meet an application's needs, that server must have access to required networks. True automation requires that switch ports be reconfigured as necessary to ensure access to specifically the VLANs required by the payloads they will represent. (Examples: Cisco 3750, Extreme Summit400)
- Network Attached Storage (NAS): The beauty of NAS devices is that they can be dynamically attached to a software payload at startup, without requiring any hardware configuration beyond the Layer 2 configuration described above. SAN is also useful (and common), but requires hardware configuration to make work. That complicates the role of automation. Part of the problem is the inconsistent remote configurability of fiber switches, which may be mitigated somewhat with iSCSI. However, NAS is quickly becoming the preferred storage mechanism in large data centers. (Examples: NetApp FAS, Adaptec Snap Server)
Over time, I see the industry adding more and more "drivers" to manage more and more data center (and perhaps desktop) resources. Imagine a world in which each software and/or hardware vendor produced standard SLA drivers for each individual component that makes up your data center environment. Every switch, disk and server; every service, container and OS; even every light bulb and air conditioner are connected to a single service level policy engine in which business policy (including cost of operations) drives automated decisions about their use and disuse.
Its not here yet, but you won't have to wait long...
I will use the label "respond" to tag posts related to response interfaces.
Wednesday, May 23, 2007
If your CIO doesn't "get it" yet, he soon will
(Emphasis mine.)However, the best talk IMHO came from Robert Beauchamp, CEO of BMC software. He's a very down-to-earth, articulate guy-even in front of 1,000 people. I was most impressed by his Shoemaker's Children analogy... that the IT (alright, BT) organizations in enterprises are arguably the least automated departments around. ERP is automated. Finance is automated. Customer interaction is automated. But IT is still manually glued-together, with operations costs continuing to outpace capital investments.
This is a gorgeous observation; so simple, so articulate, and--most importantly--so true! I have always been amazed at the amount of manual labor that goes into delivering technology that makes some other schmuck's life more labor free. Programming is a great example of this. (Even with advances in code building, IDE templating and wizard-based programming, I bet the vast majority of developers out there still shudder at the term "code generation".)
Server provisioning (bare metal or virtual servers) is also a great example. In a prior life, I worked for one of the most forward thinking technology companies out there. However, when it came to pushing code to production, it was still a server-by-server hand install job. Provisioning 4 front end portal servers took anywhere from a couple of hours to a couple of days.
Another example is trouble ticket response. How many of the system operators out there still carry pagers, and are forced to get out of bed in the middle of the night to respond to a system event? If you say "not at my company", I bet you have overseas support to back you up overnight. The response remains completely manual.
That is why I am so excited about the Service Level Automation space, its role in utility computing and its role in automating IT processes. It is time this happens, and I hope your "BT" organization is considering it.
Monday, May 07, 2007
VMWare TSX and Reducing Complexity in the Data Center
All in all, TSX was one of the most informative VMWare events I have ever attended. I only had the chance to attend three sessions--CPU scheduling, ESX networking and DRS/HA--but all three were packed with useful information. (The slides linked here are from a TSX conference in Nice, Italy, April 3-5, 2007. They are a little different from the slides I saw, but are similar enough to communicate the basic concepts.)
If you don't know much about VMWare CPU scheduling, check out that deck. Sure, its basic scheduling stuff, but it is very helpful when it comes to understanding how VMWare settings affect processor share. The networking deck is also critical if you must deploy network applications to virtual machines.
The DRS/HA deck has some helpful tips, but also clearly demonstrates the limited scale of DRS/HA. A 16 physical node limit per HA cluster, for example, is going to be problematic for most medium to large data centers. Furthermore, these are very server-centric technologies; the concept of Service Level Automation is clearly missing, as there is no concept at all of a service or application to be measured. They are hinting at a few new app-level monitors in a later release, but I just don't think monitoring service levels from a business perspective is very important to VMWare.
Bill's speech to the IT department of a large manufacturer was very interesting, if for no other reason than it clearly spelled out the argument for reducing complexity in the data center. (For a quick and dirty argument, see this article.) We are definitely at a cross roads now; IT can choose to attack complexity with people or technology. Most of us are betting technology will win. Furthermore, Bill told the assembled techies, the early adopters of any platform technology get the best jobs when that platform becomes mainstream. Almost nobody is predicting that utilty computing will fail in the long term, so now is the time to jump aboard and get involved.
I should have some time to complete the Service Level Automation Deconstructed series this week. Stay tuned for more.
Monday, April 30, 2007
What do SOA and EDA have to do with SLA?
One blog that has been capturing my interest is Jack van Hoof's "SOA and EDA". I love a blog with real patterns, term definition, and a passion for its subject matter. All put together by someone who can get an article published.
The article is actually very interesting to me from a Service Level Automation perspective. Jack captures his thoughts on the importance of building agile software architectures in the following paragraph:
Everything is moving toward on-demand business where service providers react to impulses - events - from the environment. To excel in a competitive market, a high level of autonomy is required, including the freedom to select the appropriate supporting systems. This magnified degree of separation creates a need for agility; a loose coupling between services so as to support continuous, unimpeded augmentation of business processes in response to the changing composition of the organizational structure.
(Emphasis mine.)
The only thing I would change about Jack's statement above is replacing the words "a loose coupling between services" to "a loose coupling between services and between services and infrastructure" and changing "composition of the organizational structure" to "composition of the organizational structure and infrastructure environment". (Some may have issues with the latter, but I don't mean that services should be written with specific technology in mind--just the opposite; they should be written with an eye towards technology independence.)
This is why I have been emphasizing lately the need to view the measure activity through the lens of both business and technical measures. Some of the business events thrown by an EDA may very well indicate the need to change the infrastructure configuration (e.g. if the stock market sees a 20% rise in volume in the matter of three minutes, someone may want to add capacity to those trading systems). However, the technical events from a software system (e.g. thread counts or I/O latency) may also indicate the need to change infrastructure configuration on the fly.
I wish I could spend more time collaborating with SOA architects and "tacticians". In fact, I have been speaking with Ken Oestreich about exactly this. If you are in the SOA space, and interested in talking about how SOA, EDA and SLA interconnect, let me know by commenting below. (Be sure to let me know how to contact you.) At the very least, think about how infrastructure will measure the performance of your software systems as you start your next development iteration.
Thursday, April 26, 2007
Agile Computing Catches Up to the Data Center
Is it technology? Nope. Most servers, storage and even network equipment can be managed reasonably well today by several vendors, with varying degrees of dynamic, policy-based provisioning. Several critical monitoring interfaces are also now standard in everything from power controllers to OSes to applications.
Is it infrastructure architecture? Not really, with one caveat. As long as an architecture has been built from the ground up to be easily managed and changed, with real attention paid to dependency management and virtualization where appropriate, most data centers are excellent candidates for automation. Which is a small leap away from utility computing.
Is it software architecture? Nope. I talked about this before, but SLA systems are just your basic event processing architecture specialized to data center resource optimization. The really good ones (*ahem*) can do this without adding any proprietary agentry to the managed software payload. In other words, what ends up running in your data center is almost exactly what you would have run without automation. There is little evidence on the application host that it is being managed at all.
Then what is it? One word: culture. The overwhelming obstacle that I see in the data center market today is fear of rapid change.
It is true of the sys admins, though they get the value of automation right away. They just need to see everything work before they trust it.
Its true of the storage admins, though storage virtualization is gaining ground. Unfortunately, this doesn't yet translate to accepting constant and sometimes rapid, somewhat arbitrary change within their domain.
It is most true of the network guys. Networks are the last bastion of the relatively static "diagram", mapping each component of the network architecture exactly with an eye to controlling change. The idea of switching VLANs on the fly, reconfiguring firewalls on demand, or even not knowing which server is assigned which IP address without looking at a management UI is scary as hell for the average network administrator.
And who can blame any of them? The history of commodity computing in data centers is littered with bad results from untracked changes, or badly managed application rollouts. Add to that the subconscious (of even conscious) fear that they are being replaced by software, and you get staunch resistance to changing times.
What everyone is missing here, though, is the key differentiation between planning for change, and executing it. No one in the entire industry is arguing that data center administrators should stop understanding exactly how their data centers work, what can go wrong, and how to mitigate risk. Cassatt (and I'm sure its competitors) spends significant time with each customer, even in pilot projects, making sure the data center design, software images, and service level definitions result in well understood behavior in all use cases.
But once those parameters are defined, and the target architectures, resources and service levels are defined, its time to let a policy-based automation environment take over execution. A Service Level Automation environment is going to make optimal decisions about resource allocation, network and storage provisioning and event handling, and do it in a fraction of the time that it would take a single human (let alone a team of humans). And, as noted above, once provisioning takes place, the applications, networks and storage run just as if a human had done the same provisioning.
(By the way, none of this breaks with ITIL standards. It just moves execution of key elements from human hands to digital hands. It also requires real integration between the SLA environment and configuration management, asset management, etc.)
All of this reminds me of the paradigm shift(s) that the software development industry went through from the highly planned, statically defined waterfall development methods of the early years to the always moving, but always well defined world of agile development methodologies. Its been painful to change the software engineering culture, but hasn't it been worth it for those that have found success? And, isn't it absolutely necessary for the highly decoupled and interdependent world of SOA?
Data center operations is about to undergo the same pain and upheaval. Developers, be kind and help your brethren through the cultural shift they are experiencing. Perhaps some of you in the agile methods field can begin to work out variations of your methods for data center planning and execution? Perhaps we should integrate data center planning activities into our "product-based" approaches?
Are you ready for this shift? Is your organization? What can you do today (architecturally and culturally) to ready your team for the coming utility computing revolution?
Friday, April 20, 2007
Service Level Automation Deconstructed: Analyzing Service Levels
Let me first acknowledge that my discussion of the measure capabilities included some analysis of simple metrics to create complex metrics. This is one piece of the analysis puzzle, and is a critical one to acknowledge. Ideally, all software and hardware systems would be designed to intelligently communicate the metrics that matter most to determine service levels. Where this consolidation occurs depends on the requirements of the environment:
- Centralized approach: Gather fundamental data from target systems to central metrics processor and consolidate metrics there. The advantage here is having one place to maintain consolidation rules. The disadvantage is increased network traffic.
- Decentralized approach: Gather fundamental data and do any analysis necessary to consolidate the fundamental data into a simplified composite metric there. Send the composite metric to the core rules engine (which I will discuss next).
Metrics consolidation is not really the core analytics function of a Service Level Automation architecture, however. The key functions are actually the following:
- Are metrics being received as expected? (A negative response would likely indicate a failure in the target component or the communication chain with that component)
- Are the metrics within the business and IT service level goals set for that metric
- If metrics are outside of established service level goals, what response should be taken by the SLA environment
Given my recent reading into complex event processing (CEP), this seems like at least a specialized form of event processing to me. The analysis capabilities of an SLA environment must constantly monitor the incoming metrics data stream, look for patterns of interest (namely goal violations, but who knows...) and trigger a response when conditions dictate.
The great thing about this particular EP problem is that well designed solutions can be replicated to all data centers using similar metrics and response mechanisms (e.g. power controllers, OSes, switch interfaces, etc.). Since there are actually relatively few components in the data center stack to be managed (servers [physical, virtual, application, etc.], network and storage), the rule set required to provide basic SLA capabilities is replicable across a wide variety of customer environments.
(That's not to say the rule set is simple...its actually quite complex, and can be affected by new types of measurement and new technologies to be managed. Buy is definitely preferred over build in this space, but some customizability is always necessary.)
Finally, I'd also like to point out that there is a similar analysis function at the response end as at the measure end. Namely, it is often desirable for the response mechanism to take a composite action request and break it into discrete execution steps. The best example I can think of for this is a "power down" action sent from the SLA analysis environment to a server. Typical power controllers will take such a request, signal to the OS that a shutdown is imminent, whereupon the OS will execute any scripts and actions required before signalling that OS shutdown is complete. At that time, the power controller turns off power to the server.
As with measure, I will use the label "analyze" to reflect future posts expanding on the analysis concept. As always, I welcome your feedback and hope you will join the SLA conversation.
Monday, April 16, 2007
Two articles mentioning Service Level Automation
The first is from the March 2006 issue of ACM Queue, entitled "Under New Management". The article was written by Duncan Johnston-Watt, the founder of Enigmatec. Johnston-Watt does an excellent job of outlining basic issues around one possible architecture for an autonomic data center. As expected for Enigmatec, its a policy automation focused approach, and is, in fact, one of the few articles from a policy engine vendor that I have see where the term Service Level Automation is used correctly.
Unfortunately, I don't necessarily agree that Johnston-Watt's architecture is optimal enterprise data centers. (It requires development of process automation flows to "optimize operational processes"--a significant amount of work that is prone to introducing new inefficiencies. It is also agent based, which alters the footprint of the software stacks being run in the data center, and can negatively affect the execution and architecture of the applications being managed.) All in all, though, there is some excellent information here for those thinking about Service Level Automation holistically, across the entire data center.
The other article, entitled "Virtually Speaking: Xen Achieves Higher Enterprise Consciousness" was published April 6, 2007 on ServerWatch. In the last few paragraphs of the article, uXcomm's aquisition of Virtugo is covered. In it, uXcomm claims the combination their Xen management tools and Virtugo's VMWare tools "fills a gap not just in uXcomm's portfolio but in the virtual landscape as well. Until now [...] there was a gap between service-level automation offerings and performance management products."
Hmmm. Not sure how providing SLA for only virtual servers counts as filling gaps...but, even so, I hope uXcomm is aware that everyone in this space realizes the need for resource optimization includes VM performance management. I guess my question would be, what is uXcomm doing about marrying Service Level Automation to the rest of the data center?
As a side note, I know that I owe two more articles on my Service Level Automation Deconstructed series. I am working on the "Analyze" overview now, but have discovered some interesting technology to discuss here that I am reading up on now.
Wednesday, April 11, 2007
Complexity and the Data Center
John Holland (as told in Waldrop's history) defined complex systems as having the following traits:
- Each complex system is a network of many "agents" acting in parallel
- Each complex system has many levels of organization, with agents at any one level serving as the building blocks for agents at a higher level
- Complex systems are constantly revising and rearranging their building blocks as they gain experience
- All complex adaptive system anticipate the future (though this anticipation is usually mechanical and not conscious)
- Complex adaptive systems have many niches, each of which can be exploited by an agent adapted to fill that niche
Now, I don't know about you, but this sounds like enterprise computing to me. It could be servers, network components, software service networks, supply chain systems, the entire data center, the entire IT operations and organization, etc. What we are all building here is self organizing...we may think we have control, but we are all acting as agents in response to the actions and conditions imposed by all those other agents out there.
A good point about viewing IT as a complex system can be found in Johna Till Johnson's Networld article, "Complexity, crisis and corporate nets". Johna's article articulates a basic concept that I am still struggling to verbalize regarding the current and future evolution of data centers. We are all working hard to adapt to our environments by building architectures, organizations and processes that are resistant to failure. Unfortunately, entire "ecosystem" is bound to fail from time to time. And there is no way to predict how or when. The best you can do is prepare for the worse.
One of the key reasons that I find Service Level Automation so interesting is that it provides a key "gene" to the increasingly complex IT landscape; the ability to "evolve" and "heal" the physical infrastructure level. Combine this with good, resilient software architectures (e.g. SOA and BPM) and solid feedback loops (e.g. BAM, SNMP, JMX, etc.) and your job as the human "DNA" gets easier. And, as the dynamic and automated nature of these systems gets more sophisticated, our IT environments get more and more self organizing, learning new ways to optimize themselves (often with human help) even as the environment they are adapting to constantly changes.
In the end, I like to think that no matter how many boneheaded decisions corporate IT makes, no matter how many lousy standards or products are introduced to the "ecosystem", the entire system will adjust and continually attempt to correct for our weaknesses. In the end, despite the rise and fall of individual agents (companies, technologies, people, etc.), the system will continually work to serve us better...at least until that unpredictable catastrophic failure tears it all down and we start fresh.
Monday, April 09, 2007
SOA blog recognizes SLA!
Now, I don't know Todd from Adam, so I took some time to read through a few of his posts. His subtitle explains his focus very well: "SOA, BPM, and other strategic IT initiatives". His recent posts included coverage of the Web Methods acquisition by Software AG (I wonder what he thinks about the BEA/Amberpoint announcement), the vagrancies of developing SOA clients and services in parallel (a problem I know quite well, actually, from my past as an architect with Forte Software and Sun Microsystems), and the rapid integration of management and monitoring across all tiers of enterprise architectures (a post I want to specifically address in a later post of my own). From what I read so far, he is a fairly holistic enterprise architect with a good eye for both functional and infrastructure issues.
But why was I getting references to my site from his? I kept looking, and to my delight rested my eyes on my name in his blogroll...WOO HOO!!! The first such link that I know of!
By the way, follow some of those other blogroll links...there you will find a huge wealth of knowledge about SOA, BPM and the culture shock the distributed systems community is feeling as a result (with some calm, collected voices providing solid advice). Don't think SOA and SLA are related? You'll find lots of evidence to the contrary among the posts of Todd and his colleagues.
Thanks, Todd, and you've been added to my daily RSS feed.
Thursday, March 29, 2007
Service Level Automation Deconstructed: Measuring Service Levels
Rather than trying to do an exhaustive coverage of this topic (and the other topics in this series) in a single post, what I am going to do is provide a "first look" post now, then use labels when followup posts have relative information. The label for this topic will be "measure".
In my next installment, I'll introduce analysis of those metrics against service level objectives (SLO) the business requires. That post, and future related posts will be labeled with "analyze".
In the final installment of the series, I'll describe the techniques and technologies available to digitally manipulate these systems so that they run within SLO parameters. Posts related to that topic will be labeled "respond".
As noted earlier, my objective is to survey the technologies, academics, etc., of each of these topics in an attempt to enlighten you about the science and technology that enables service level automation.
How do we measure quality of service?
Measuring quality of service is a complex problem, not so much because it is hard to measure information systems and business functionality. I (and I bet you) could list dozens of technical measurements that can be made on an application or service that would reflect some aspect of its current health. For example:
- System statistics such as CPU utilization or free disk space, as reported by SNMP
- Response to a ping or HTTP request
- Checksum processing on network data transfers
- Any of dozens of Web Services standards
The real problem is that human perception of quality of service isn't (typically) based on any one of these measurements, but on a combination of measurements, where the specific combination may change based on when and how a given business function is being used.
For example, how do you measure the utilization of a Citrix environment? Measuring sessions/instances is a good start, but--as noted before with WTS--what happens when all sessions consume a large amount of CPU at once? CPU utilization, in turn, could fluctuate wildly as sessions are more or less active. Then again, what about memory utilization or I/O throughput? These could become critical completely independently from the others already mentioned.
No, what is needed is more mathematical--one (or a couple of) index(es) of sorts generated from a combination of the base metrics retrieved from the managed system.
There are tools that do this. They range from the basic capabilities available in a good automation tool, to the sophisticated evaluation and computation available in a more specialized monitoring tool.
What I am still searching for are standard metrics being collected by these tools, especially industry standard metrics and/or indexes that demonstrate the health of a datacenter or its individual components. I'll talk more about what I find in the future, but welcome you to contribute here with links, comments, etc. to point me in the right direction.
Monday, March 26, 2007
Service Level Automation is green--when done right
So, if capacity is not being used at any given point in time, why have it consume power or cooling at all? Some "automation" products require underused physical infrastructure to remain running in order to support their management layers--just in case. This is unfortunate, as an idle hypervisor is just idle capacity. Its not serving a current business need.
A truly efficient SLA platform is aware of the power controller states of each of its physical servers, and can power down unused servers. Servers are only turned on when they are needed to meet some aspect of the system's service level goals.
Note Vinay's description of the QA labs at Cassatt. As you might expect, he pushes the limits of what a SLA datacenter must endure, yet can always scale his power and cooling needs to his current workload. Can you say that about your datacenter?
Thursday, March 22, 2007
Service Level Automation Deconstructed: Introduction
* The factors contributing to software service quality can be measured electronically.
* Runtime targets indicating high quality of service can be defined for those measurements.
* Systems involved in delivering software functionality can be manipulated to keep those measurements within the runtime targets.
I think the support for each of these premises should be explored more deeply, so I plan to begin a little survey of the technologies and academics over the next few weeks. The idea is to get a good sense of what standards/technologies/concepts/etc. can be used to meet the requirements of each premise. I also hope to discuss how a system smart enough to take advantage of them(*) can save a large datacenter both in terms of direct costs, as well as in losses due to service level failures.
Why Service Level Automation? I wrote about this earlier. However, as a quick reminder, think of service level automation as meeting this objective:
Delivering the quantity and quality of service flow required by the business using the minimum resources required to do so.
I've been quite busy both at work and at home, so I'm hoping to use this exercise as a way to increase my posting frequency. Stay tuned for more.
Wednesday, March 21, 2007
5 things...
Here are five things most people do not know about me:
- I was born in Reading, England.
- I play pretty decent guitar. I don't know that many songs by other people (a problem when whipping out the guitar at parties), but I have several original works that I think hold their own very nicely against most pop drivel. Lately, however, I have been working on "Tears in Heaven" by Eric Clapton.
- I played Mr Anthrobus in Thornton Wilder's "The Skin of our Teeth" in high school. I was a geeky, awkward teenager trying to play a 40 year old man, and was the only member of the primary cast not to win an award for my performance in that show. Now that I am 40, I wonder what the hell was so hard...
- My computing career started in fifth grade in Cedar Rapids, Iowa. I was lucky enough to get in a science focused program at a nearby elementary school, and one of the kids' moms was one of the first BASIC programmers at Rockwell Collins, the aviation electronics firm. She came to our school once a week and taught us the basics of variables, loops, conditional statements and subroutines. Very cool. I got caught a bunch of times programming on the teletype terminal in the back of the classroom while I should have been listening to the teacher. Later, my luck continued as the father of one of my close neighborhood friends bought the fifth (or something like it) Apple II computer in the state of Iowa. We would program in BASIC every day after school, and tried to get into writing games and such.
- Later, in college, I was determined to be a Music/Computer Science double major...for all of one semester. I didn't practice the music stuff enough, so I got a low grade there, and I hated my systems organization class, so I lost interest in computer science. (Dumb reason, now that I look back, but it worked out.) Instead, I started taking every math and physics class that I could, and finished with a Mathematics/Physics double. The day of graduation, I swore to my friends "I will NEVER be a computer programmer for a living". Two and a half years later, I was coding C for a small manufacturing company. (Do not try to predict the future, even your own. Its pointless. Setting goals is OK, but be willing to float a bit with the breeze.)
Now, let me please introduce to you five more randomly selected from my blogosphere:
- My mom.
- Katie Tierney, a former collegue with excellent technical intuition who is proving herself to be a hell of a "head of household" as well.
- Rama Roberts, another former collegue whose blog never fails to entertain and enlighten.
- Management guru, Tom Peters, who reenforces my drive to amaze both my employers and customers by being a service professional first and foremost.
- Alessandro Perilli, author of the virtualization-focused virtualization.info blog.
Monday, February 26, 2007
When CPU utilization is not enough...
Much to my delight, the last few weeks have been filled with customer activity, ranging from helping a Service Level Automation-enabled appliance for a major software company, to assisting the financial wing of one of the world's largest manufacturers to experience first hand the benefits of utility computing.
The latter runs an application that is highly dependent on Windows Terminal Services to deliver a client-server UI to thousands of retail outlets world-wide. Uptime is critical to this application, as customers will go elsewhere for financing if this application doesn't confirm credit within minutes of a purchase decision.
Unfortunately, WTS is also a very inefficient consumer of server payload. It is a session-based infrastructure, which means that a user will be attached to a specific server for their desktop access until they either log out or are kicked off. If 15 user sessions share a physical server, there is no way to predict the load on the system. All 15 sessions could be idle, or all could quickly start consuming cycles simultaneously.
This gives me my first really good example of when CPU and/or memory utilization are not good Service Level metrics on their own. Imagine an environment using WTS to support hundreds of users. These users use their Windows sessions to run a variety of tasks, much like any Windows user. Some tasks use a high level of CPU and memory, others very little. Quite often, the session will sit idle for several minutes.
Now, if you create lots of sessions because the CPU is idle, you could end up with problems if they all get active at the same time (say right after lunch). If you stop creating sessions on a server because CPU utilization is high, you may end up with a highly under utilized server when one user's game of Quake wraps up.
That's not to say that CPU or memory utilization aren't an important part of the Service Level "equation". The truth is, there are several metrics that apply to WTS capacity: sessions, CPU utilization, memory utilization, licenses, etc. Since determining Service Level compliance probably involves evaluating the relationships between several of these metrics at once, there will most likely be one or two compound metrics based on mathematical equations combining these "root" metrics in a way that reasonable thresholds can be set.
Another interesting observation is that this is a lot like the Java EE Service Level Automation problem. (Thanks to Luis Cuyun for pointing this out.) While most horizontally scalable application tiers can be scaled up and down as a unit (i.e. "add a node/remove any node"), app servers, hypervisors and (now) WTS all must be monitored as a unit, but managed on a per server basis (in this case, "add a node/remove this specific node"). This is because the "instances" that each of these software resources are hosting are "sticky" to a server (VMotion not withstanding), and you don't want to shut down any server when capacity is not longer needed, you want to only remove the specific servers with no live sessions remaining on them.
(Speaking of VMotion, one of the things that both Java EE and WTS will require to be really optimizable is the ability to move live services/sessions from one server to another in real-time. Anyone know of a technology addressing either of these?)
The good news for a good Service Level Automation environment (*ahem*) is that if one of these problems (Java EE, virtual servers or WTS) is solved correctly, the same basic technology can be applied to all of them. That's not to say that anyone is doing this for WTS today (to my knowledge, no one is), but I like the idea that the use cases that apply to Java EE SLA also apply to WTS SLA.
I'm hoping to have more to write about this as this pilot continues. In the meantime, anyone with Windows experience is welcome and encouraged to contribute their two cents to this discussion. In particular, are there any tried and true service level metrics for WTS that are being used out there? In general, there are so many moving parts here, that I am sure there are many critical factors to Service Level Automation of WTS that I have not covered, or even considered.
Thursday, February 01, 2007
Welcome Vinay!
Welcome, Vinay! I look forward to the good read.
Great Blogs Unite!
Completely unbiased, of course! :D
Thursday, January 25, 2007
Greasing the skids...Simplifying Datacenter Migration
This is a huge trend amongst Fortune 500 companies. In my work, I keep hearing VPs of Operations/Infrastructure and the like saying things like "we are consolidating from [some large number of] datacenters to [some small number, usually 2 or 3] datacenters." In the course of these migrations, they are rationalizing the need for each application that they must migrate from one datacenter to another.
The cost of these migrations can be staggering. "Fork-lifting" servers from one site to another incurs costs in packaging, shipping and replacing damaged goods (hardware in this case). Copying an installation from one datacenter to another involves the same issues: packaging (how to capture the application at the source site and unpack it at the destination site), shipping (costs around bandwidth use or physical shipping to move the application package between sites) and repair of damaged goods (fixing apps that "break" in the new infrastructure).
What if something could "grease the skids" of these moves--reduce the cost and pain of migrating code from one datacenter to another?
One approach is to package your software payloads as images that are portable between hardware, network and storage implementations. Now the cost of packaging the application is taken care of, the cost of shipping the package remains the same or gets cheaper, and the odds of the software failing to run are greatly reduced because it is already prepared for the changing conditions of a new set of infrastructure.
Admittedly, the solution here is more related to decoupling software from hardware than Service Level Automation, per se. But a good Service Level Automation environment will act as an enabler for this kind of imaging, as it too has to solve the problem of creating generic "golden" images that can boot on a variety of hardware using a variety of network and storage configurations. In fact, I have run into several customers in the last couple of months that have a) recognized this advantage and b) rushed to get a POC going to prove it out.
Of course, if you can easily move software images between datacenters, simpler disaster recovery can't be far behind...
Monday, January 22, 2007
NAS overtaking SAN for automated server virtualization
(As an aside, I also love the quote in the article where EMC Corp. vice president of technology alliances Chuck Hollis pointed out that "To be honest, we're not seeing a whole lot of high performance stuff being put on VMware." Don't be fooled, most large datacenters will always have applications that can not be virtualized without a penalty.)
I would have to say that EMC's observation aligns with my own, as it has been clear for some time that NAS has offered some advantages over SAN for application storage in Cassatt environments. It boils down to accessibility--SAN requires special interface cards, and very few (if any) of the SAN switches today are remotely configurable by an automation environment. There are cool vendors out there (see 3PAR and DataCore, for example) that have tools to increase the dynamic nature of SANs, but NAS tends to rule here.
The article also notes some reasons why performance is overrated in the SAN vs. NAS comparison. Low end (e.g. workgroup class) NASs may suffer from some limitations based on network bandwidth, but TOE NICs and multi-NIC high-end NAS configurations are "widening the highway", allowing NAS performance to catch up to, and even surpass SAN. Cost/performance numbers are still something to consider, but I expect that the only apps that will be using fiber SAN in five years will be extremely high I/O applications, such as OLTP apps.
Let me give you quick reason why all of this is important: multitenancy. The Software as a Service (SaaS) and Managed Hosting Provider spaces have embraced the concept of one infrastructure supporting a large number of unique, individual clients. However, to achieve this, one needs to be able to "virtually" isolate each client from each other for both security and data integrity reasons.
To achieve this isolation, it is necessary to uniquely assign each customer two things: network access and (you guessed it) storage. Managed hosting providers and SaaS vendors are looking for tools that will allow them to dynamically assign a server (and thus, its hosted software) to specific VLANs and LUNs/namespaces. This will be a key focus for automation vendors in the next 2 years or so.
What do you think? How do you plan to address storage in your automated data center?
Thursday, January 18, 2007
Tsunami of Automation
Seems to me that to achieve service levels for an application, each part of the application's infrastructure, from the app itself to the electricity it consumes needs to be measured and adjusted as needed to meet demand. I guarantee that if you do less, you will need to integrate your "policy-based" tools with other "policy-based" tools ad nauseum. And it will take you years to get there. (Note: we need standards here...)
Nonetheless, its good to see all of the market validation going on right now. And I encourage you to read about these vendors and others talking about utility computing, QoS, and automation. There are a lot of cool ideas here, waiting to work together...
Thursday, January 04, 2007
IDC recognizes Service Level Automation!!!
What is really cool about this is that it validates the need for systems that focus not on infrastructure automation per se (e.g. automating deployment processes, automating server creation, etc.), but that focus on the needs of the business and their applications and services. Sure, server virtualization, metering and billing, and so on are still important in a utility computing environment, but the concept is not complete unless something is monitoring your quality of service, and making adjustments as necessary to maintain compliance at minimal cost.
Of course, every so-called "policy-based" automation solution, no matter how single-product focused will probably now claim this title. But since you've been reading this blog, you know better than to fall into that trap...right? :)
Ken Oestreich starts blogging
Tuesday, January 02, 2007
I’m facinated by this concept of the coupling between software (especially web services) and infrastructure (including servers, networks and storage). In fact, Cassatt has done a tremendous amount of thinking around how Service Level Automation and service oriented infrastructure applies to web services, especially in the changing world of the software infrastructure used to host those services. (The hardware evolution is also facinating, but is tangental to the conversation here.)
Dynamically changing the number of physical, virtual or even application servers hosting the service certainly addresses the sticky performance issues surrounding web services, but it does nothing to address the *efficiency* issues, especially with regards to how resources can be pooled to meet the demand of a number of applications and services at the same time. Think “how can I deliver the required service levels for my applications and services using the minimum resources required to do so”.
This is what I am addressing on my blog. I hope you will check it out and comment at will on what you see there. I’m glad to see such interesting discussion about service oriented infrastructure. It is certainly a problem that will be addressed dramatically in the next 5-10 years.
Service Level Automation in 2007
- Service Level Automation will expand from the server-centric approach of today, to a variety of granularities, including the application level, the middleware level, the OS level, the cluster level, etc. This is crucial when it comes to truly optimizing both hardware and software usage, including optimizing license utilization, etc.
- Hardware will play a much more significant role in virtualization in 2007. Check out Intel VT and AMD V on-board virtualization, Xsigo I/O virtualization and the Mazu Networks real-time network discovery and analysis appliance as examples. (Note that none of these technologies are in the automation space; rather, they push the boundries of what is possible the virtualization, dynamic provisioning and monitoring functions.)
- The incentive for larger organizations to move to a true utility computing infrastructure will grow tremendously as initiatives are announced throughout the Fortune 500.
- Successful SLA and utility computing implementations will continue to appear in both commercial and government customers. Unsuccessful implementations will also appear, either due to poorly planned solutions (the "I can build it" syndrome), or poorly planned projects (the "I can convert my entire datacenter at once" syndrome).
- The winners in the utility computing and service level automation space will be defined by successful implementations, strong partnerships and innovation that continues to disrupt the traditional tightly coupled, "silo"-based approach IT uses today.
- Utility computing will appear as a system integrator specialty with increasing frequency over 2007. In fact, specialist boutique firms will start appearing in large cities around the United States and Western Europe, and will be quite profitable.
Let me know what you think. I've only had a few hours to think about this this year, so I'm sure more will occur to me over the next few days.
Monday, December 18, 2006
The Organic Cluster
We've been talking about distributed application architectures, and how their tight coupling to underlying hardware and middleware has left most production applications running in infrastructure "silos", where spare capacity is locked up and unavailable to other systems.
Traditionally, nowhere has this been more obvious that with your database servers. Not designed to be scalable horizontally, these servers have relied on an excessively high amount of overprovisioning to be absolutely sure that performance was consistently high, regardless of actual demand. Do you need more capacity for your database than the current server provides? Then buy a bigger box and figure out how to migrate the database without crippling the business.
The exception to this story (right now) is ORACLE 10g RAC (Real Application Clustering). RAC is a grid-based database engine, which uses clustering technology to allow the RDBMS to be distributed across several servers. This increases availability greatly, and allows for an easier upgrade path when new hardware is required.
Unfortunately, as designed, the ORACLE cluster still requires the administrator to allocate excess capacity in case of high demand. Each node of the cluster must be running at all times (in the default architecture), which means each server must be dedicated to RAC whether it is needed or not.
A good Service Level Automation environment gives the administrator an interesting new capability, however. Because the ORACLE cluster can run with less than the maximum number of servers defined for the cluster (which is how the database keeps running when a server node is lost), it is possible to capture the cluster in an image library, and then allocate nodes only according to actual demand. No need for weird code or configuration changes to ORACLE, and no need to have spare capacity "dedicated" to RAC.
If more capacity is needed, the SLA environment will grab it from the pool of spare capacity available to all applications, not just RAC. When demand is detected to exceed the safety margins of the current "live" set of servers, the SLA environment boots up a new node, and RAC "rediscovers" its "lost" node. When demand falls away, the SLA environment shuts down an unneeded node, and RAC just detects that a node went down, but keeps on chugging.
Lest you think this is a pipe dream, my employer Cassatt has this running in its labs and has indeed provided proof-of-concept to prospective customers. And they like it. Which is another reason why Service Level Automation is changing the way IT runs.
Monday, October 16, 2006
Two important links...
- Me doing the Cassatt schtick for the world to see. (Note the great hair!)
- My new Service Level Automation del.icio.us page, with links to a variety of interesting sites related to service level automation, virtualization and (okay, I have to be loyal to my employers) Cassatt. (I have added this to the links section on the right of the http://servicelevelautomation.blogspot.com landing page)
Thursday, October 12, 2006
The Datacenter is Dead! (Or Just Mutating Badly!)
This computer science professor stood in front of a highly attentive audience one evening and declared "data is dead!"
His point was that if we modified our models of how computers stored data persistently to use a "always executing" approach, the need for databases to manage storage and retrieval of data from block-based storage would be made obsolete. (I view "always executing" systems much like your cell phone or Palm device today; when you turn on your device, applications remain in the state they were in when you last shut it off.)
Its funny to remember how much we thought objects were going to replace everything, given the intense dependency we have on relational databases today. But his arguments forced us to really think about the relation between the RDBMS and object oriented applications. One result of years of this thinking, for example, is Hibernate.
Jonathan Schwartz, my beloved leader in a former life, recently blogged about the future of the datacenter, contending that the need for large, centralized computing facilities are numbered. In other words, "the datacenter is dead".
His contention is that the push towards edge and peer computing with "fail in place" architectures would make central facilities tended by technology priests obsolete. Ultimately, his point is that we should reexamine current enterprise architectures given the growing ubiquity of these new technologies.
I have to say, I think he makes a good argument...up to a point. My problem is that he seems to ignore two things:
- Data has to live somewhere (i.e. data is certainly not dead)
- People expect predictable service levels from shared services--the more critical those service levels, the more critical that those service levels can be guaranteed.
I think its good news that, in order to achieve such a vision, we must take baby steps from the static, siloed, humans-as-service-level-managers approach of today's IT shops.
As you may have guessed from my previous blogs:
- I believe the first of these steps is to shed dependencies between software services and infrastructure components.
- Following that we need to begin to turn monitors into meters, capturing usage data for both real time correction of service level violations, as well as analysis of usage and incident trends.
- Finally, we need the automation tools that guarantee these service levels to operate across organization boundaries, allowing businesses to drive the behavior (and associated cost) of their services wherever they may run in an open computing capacity marketplace.
No, neither data nor the datacenter are dead, they are just evolving quickly enough that they may soon be unrecognizable...
Thursday, October 05, 2006
InfoWorld: ITs Virtual Assett Economy
http://www.infoworld.com/article/06/10/04/41OPcurve_1.html
Hmmmm… Service Level Automation, anyone?
“When money is distributed to managers for IT-related purchases, that capital goes to IT with the investor’s minimum requirements attached. Ideally, those requirements will be expressed in terms that are accessible to the investors…”
Great concept. Almost a "commons" (in the 18th century farming sense) for computing resources. Certainly many simularities to commodity market models as well (e.g. options, trading, etc.).
Tuesday, October 03, 2006
Service Virtualization defined
That's crazy.
My argument starts with the observation that its not server utilization service levels that businesses care about, but the quality and availability of the services that run the business that really matter. From a business perspective--from the view of the CEO and CFO--its not how many servers you use and how you use them, its how many orders you gather and how cheaply you gather them.
So, this focus by the VM companies on hardware and manipulating servers (virtual or not) falls short of meeting the goals of the business. Look closely at what VMWare, Virtual Iron and even XenSource are offering:
- Virtual Servers. This is the core of their value proposition, and its by far the most valuable tool they deliver. As we established earlier, this is needed technology.
- Virtual Server Management. VirtualCenter, Virtualization Manager, etc. provide key tools for managing virtual servers.
- Automation. Tools to provision, expand, move and replace servers based on current observed conditions.
Virtual machine technologies have no concept of a service, or even an application. They barely have the concept of an OS. This is by design; if they focus on the hardware virtualization problem, they have a fairly simply bounded problem--just make some software behave exactly like its emulated hardware platform. That way, you can cover existing software installations with minimal effort, and don't have to worry about the vagrancies of application, network and storage configuration. All of that can happen outside of a simple virtualization wrapper targeted at making the virtual servers work in the physical reality.
The true holy grail for IT, in my opinion, is service virtualization. I will define service virtualization as technology that decouples a set of functionality (a web service, an application, etc.) from any of the computing resources required to execute that functionality, regardless of whether those resources are themselves hardware or software. What we ultimately want to do is to optimize the delivery of this functionality by whatever metric is deemed important by the business.
Thus, I applaud the validation of policy-based automation, decoupling of physical hardware from software and automated response to server load and failure that the VM companies are clearly giving. However, I caution each of you to consider closely whether automating server management is enough, or if service virtualization is the better path.