Welcome!

SDN Journal Authors: Carmen Gonzalez, Trevor Parsons, Rex Morrow, Datical, Elizabeth White, Sandi Mappic

Related Topics: Cloud Expo, Java, SOA & WOA, Virtualization, Big Data Journal, SDN Journal

Cloud Expo: Blog Feed Post

All for the Want of a Global Load Balancer

Thoughts on the AWS / Netflix Christmas Eve Elastic Load Balancer Event

Christmas Eve and the AWS/Netflix outage to me aren't so much about whether or not the Cloud is viable or scary or dangerous. Rather, the event resonated with users across the United States because the Cloud delivers so much utility to each of us. And regardless of who was at fault -- Netflix or Amazon Web Services, the event made it clear that there's no going back and that the Cloud has quickly become a part of our culture and our everyday lives. This is significant because while the Internet itself is a technology consumers have grown to love, Cloud is a way of delivering service that makes a service like Netflix streaming possible, and at a measly $8 bucks a month. The Netflix business model of delivering outsize utility for a low price point makes the business of streaming video all the more difficult. HBO, Cinemax, and the networks for me are unusable. I sense something beyond just making money remains in play at Netflix. Somewhere in that organization seems to beat a heart that quickens for humanity.


While much is made of the sparsity of the Netflix streaming catalog, keep in mind that streaming is the disruptor's second act. The Wicked Witch, aka Blockbuster is dead, I think in part because each of us desired a little bit of payback for all the times we returned videos late and were charged outrageous fees. Rather than punish Netflix Streaming for innovating, and for challenging the iron grip of the legacy media houses, I praise Netflix and personally admire the generous open source contributions Netflix has made to the Cloud Community.

I once tweeted that the value of the code and tools Netflix has shared on GitHub may well be larger than the value of the streaming business as a whole. That may be true for about five minutes. We're at a tipping point in streaming video media similar to that of the music business just before the iPod. Yet each evolution carries forward a little bit of spin from the last disruption. In this iteration all the players know the iPod playbook and so they are trying very very hard to fight the gravity of disruption. Having lost Blockbuster and effective control over the distribution of their content, the old guard hold onto to their catalog stubbornly hoping Netflix will disappear. But we all know how this movie ends. I see a much larger Netflix catalog of content on the horizon and substantial value for consumers. Once the old guard have been disintermediated and a more open, consumer friendly market for content prevails every Christmas will be a little brighter (that's kind of a stretch?). In my opinion what's in play today with respect to old media firms is not so different from the agency model that kept eBook prices artificially high for many consumers.

So what about the Global Load Balancer Already?
Most Cloud Architect and Solution Architects of high availability systems are familiar with Load Balancers. AWS designed a mostly well-intended and useful service around this technology and it's often referred to as the (infamous) ELB (Elastic Load Balancer). Functionally a "classic" load balancer distributes service requests across hosts grouped into a pool of largely identical servers. Spreading requests across multiple servers is a key capability required for the horizontal scale-out favored by Cloud and to a lesser degree Enterprise Architecture.

A Global Load Balancer operates one or more levels above the Local Load Balancer. Similar to a classic DNS server, the Global Load Balancer returns an ip address. Often the IP address is the address of a Local Load Balancer but could be configured to return a public or private ip address. The IP address or VIP returned is determined by the specific solution requirements. Because Global Load Balancer technology stands on the shoulders of a fundamental and ubiquitous technology like DNS, the service can be very resilient if configured that way. If LLB (Local Load Balancer) high-availability solution is designed correctly the Load Balancer monitors the health and responsiveness of a pool of servers and directs requests to nodes of an application. A Global Load Balancer performs a similar function, except at a data center level. A Global Load Balancer can detect when a Local Load Balancer is no longer available, and when this happens can return the ip address of a load balancer or endpoint in another data center anywhere in the world. As it is based on DNS technology the Global Load Balancer is not coupled to a specific network. Moreover, the top-level domain is likely not hosted in any Cloud, and so provides a degree of separation. Further, Global Load Balancers can determine the nearest endpoint for a specific user request, and can return an endpoint closest to a user. Global Load Balancers can be configured to monitor the latency of response times across a pool of endpoints. In times of network failure, the endpoint that is normally "closest" or "lowest latency" frequently becomes a very high latency endpoint. So the ability to monitor latency and to adapt how it responds to requests based on latency, enables a very fine grained capability to direct users around system and network anomalies such as those that continue to plague US-EAST-1.

My observation is simply this: why don't Cloud or Web firms take a lesson from classic availability architecture and use a load balancer to enable endpoints to fail across data centers. For enterprise, why make Cloud an either/or question. And why gamble with the availability of important systems when it's not so difficult (especially if you've figured out the persistence layer) to balance systems across multiple Clouds. Using such an architecture to enable rapid switching of endpoints to other data centers could enable an enterprise to run an application both in the Cloud and in the data center and to balance traffic across them not in the sense of the much hyped "Cloud Bursting" which seems to focus mostly on "bursty capacity" but rather as a strategy for mitigating risk and reducing the correlated risk of running Load Balancers and applications within a single specific Cloud, even in multiple availability zones (as in the case of AWS). Hurricane Sandy in New York City would have been far more disruptive if not for systems built using Global Load Balancers for high availability. While there's no comparison to the complexity of failing over critical, albeit smaller-scale platforms to Netflix, several non-web facing internal applications for which I designed the architecture in 2011-2012 failed over to alternate data centers simply by  changing the endpoint returned by the Global Load Balancer.  In the case of Netflix, the option to migrate all the data was not an option, yet the instances remained healthy while those impacted by the ELB failure received no requests. An alternate strategy for directing traffic, in hindsight, could have made a difference.  It puzzles me to no end why given the extensive focus on Cloud failure modes, the AWS ELB remains a single point of failure for Netflix and many other applications running in AWS whether at modest or extreme scale.

I wonder if Netflix will select an additional Cloud in 2013 and in doing so create some real competition in the Cloud Service Provider space. After such a high profile failure (Christmas Eve, for Christ's Sake), I feel certain the decision has already been made. The impact of such a move would legitimize another Cloud. People who know these things tell me that 80% of Amazon Web Services capacity is in US-EAST-1. To me this suggest Amazon Web Services has fallen victim to it's own success. The nature of Clouds is such that as they grow bigger the network effect driven by the efficiency of locating data and compute capacity as close together as possible becomes overwhelming and the penalty for locating in another region, such as the West Coast, or another Cloud in another data center, becomes higher. While recently AWS has announced some capabilities to make it easier to migrate to the US-WEST-2 (Oregon region, which is priced about the same as US-EAST-1) such capabilities don't really seem to matter.

I spend a great deal of my time learning from web scale best practices such as those co-developed by firms like Netflix, Heroku, Pinterest, Google and redefine what people think they know about distributed computing..  Yet based on the theme of AWS re:invent, and my personal experience in Fortune 500 Cloud, 2012 may not only "not be the end of time and ancient calendars," but 2012 may mark the year when this thinking makes the leap and infects the DNA of Fortune 500 technology. The itchy little problem of Cloud as ready for big business remains the persistent failures, yet the risk of failures in hindsight can be minimized--and using technology commonly used by Enterprise that unlike Oracle RAC clusters and other High Availability technology works just fine in both Cloud and traditional data centers.

Listen closely to what Cloud and web scale practitioners have to teach. Architecture of Cloud applications deviates in ways that will make your database and application engineers pull-out their hair, scream, and storm out meetings (based on what I've seen, except for the hair pulling). For example the application server architecture and scale-out strategy for scaling applications in the Cloud is very different from how most of Fortune 500s do Enterprise Application Server clusters today. And after you fail in the Cloud trying to kick it old school, the enlightenment comes quickly. Yet at the same time, don't get too caught up in the hype and forget everything. The Global Load Balancer commonly used in large enterprise, when deployed effectively, could very well help you build applications which balance the new with the familiar.

Further Reading / Cultural Reception of Cloud

Summary of the December 24, 2012 Amazon ELB Service Event in the US-East Region

Some of the media reception of the events. I find the coverage to be grossly inaccurate, yet it's part of the cultural reception of the Cloud, so I've included some links:

‘The Cloud’ Challenges Amazon http://nyti.ms/ZBXT86

More Stories By Brian McCallion

Brian McCallion, founder of New York City-based consultancy Bronze Drum focuses on the unique challenges of Public Cloud adoption in the Fortune 500. Forged along the fault line of Corporate IT and line of business meet, Brian successfully delivers successful enterprise public cloud solutions that matter to the business. In 2011, while the Cloud was just a gleam in the eye of most Fortune 500 firms Brian designed and proved the often referenced hybrid cloud architecture that enabled McGraw-Hill Education to scale the web and application layer of its $160M revenue, 2M user higher education platform in Amazon Web Services. Brian recently designed and delivered the JD Power and Associates strategic customer facing Next Generation Content Platform, an Alfresco Content Management solution supported by a substantial data warehouse and data mart running in AWS and a batch job that processes over 500M records daily in RDS Oracle.”

Comments (0)

Share your thoughts on this story.

Add your comment
You must be signed in to add a comment. Sign-in | Register

In accordance with our Comment Policy, we encourage comments that are on topic, relevant and to-the-point. We will remove comments that include profanity, personal attacks, racial slurs, threats of violence, or other inappropriate material that violates our Terms and Conditions, and will block users who make repeated violations. We ask all readers to expect diversity of opinion and to treat one another with dignity and respect.


@CloudExpo Stories
Performance is the intersection of power, agility, control, and choice. If you value performance, and more specifically consistent performance, you need to look beyond simple virtualized compute. Many factors need to be considered to create a truly performant environment. In their General Session at 15th Cloud Expo, Phil Jackson, Development Community Advocate at SoftLayer, and Harold Hannon, Sr. Software Architect at SoftLayer, to discuss how to take advantage of a multitude of compute option...
Come learn about what you need to consider when moving your data to the cloud. In her session at 15th Cloud Expo, Skyla Loomis, a Program Director of Cloudant Development at Cloudant, will discuss the security, performance, and operational implications of keeping your data on premise, moving it to the cloud, or taking a hybrid approach. She will use real customer examples to illustrate the tradeoffs, key decision points, and how to be successful with a cloud or hybrid cloud solution.
In today's application economy, enterprise organizations realize that it's their applications that are the heart and soul of their business. If their application users have a bad experience, their revenue and reputation are at stake. In his session at 15th Cloud Expo, Anand Akela, Senior Director of Product Marketing for Application Performance Management at CA Technologies, will discuss how a user-centric Application Performance Management solution can help inspire your users with every appli...
With the explosion of the cloud, more businesses are transitioning to a recurring revenue model to generate reliable sales, grow profits, and open new markets. This opportunity requires businesses to get to market quickly with the pricing and packaging options customers want. In addition, you will want to take advantage of the ensuing tidal wave of data to more effectively upsell, cross-sell and manage your customers. All of this is possible, but only with the right approach. At 15th Cloud Exp...
Planning scalable environments isn't terribly difficult, but it does require a change of perspective. In his session at 15th Cloud Expo, Phil Jackson, Development Community Advocate for SoftLayer, will broaden your views to think on an Internet scale by dissecting a video publishing application built with The SoftLayer Platform, Message Queuing, Object Storage, and Drupal. By examining a scalable modular application build that can handle unpredictable traffic, attendees will able to grow your de...
The cloud provides an easy onramp to building and deploying Big Data solutions. Transitioning from initial deployment to large-scale, highly performant operations may not be as easy. In his session at 15th Cloud Expo, Harold Hannon, Sr. Software Architect at SoftLayer, will discuss the benefits, weaknesses, and performance characteristics of public and bare metal cloud deployments that can help you make the right decisions.
Over the last few years the healthcare ecosystem has revolved around innovations in Electronic Health Record (HER) based systems. This evolution has helped us achieve much desired interoperability. Now the focus is shifting to other equally important aspects – scalability and performance. While applying cloud computing environments to the EHR systems, a special consideration needs to be given to the cloud enablement of Veterans Health Information Systems and Technology Architecture (VistA), i.e....
Cloud and Big Data present unique dilemmas: embracing the benefits of these new technologies while maintaining the security of your organization’s assets. When an outside party owns, controls and manages your infrastructure and computational resources, how can you be assured that sensitive data remains private and secure? How do you best protect data in mixed use cloud and big data infrastructure sets? Can you still satisfy the full range of reporting, compliance and regulatory requirements? I...
Scott Jenson leads a project called The Physical Web within the Chrome team at Google. Project members are working to take the scalability and openness of the web and use it to talk to the exponentially exploding range of smart devices. Nearly every company today working on the IoT comes up with the same basic solution: use my server and you'll be fine. But if we really believe there will be trillions of these devices, that just can't scale. We need a system that is open a scalable and by using...
Is your organization struggling to deal with skyrocketing volumes of digital assets? The amount of data is growing exponentially and organizations are having a hard time managing this growth. In his session at 15th Cloud Expo, Amar Kapadia, Senior Director of Open Cloud Strategy at Seagate, will walk through the essential considerations when developing a cloud storage strategy. In this discussion, you will understand the challenges IT is facing, why companies need to move to cloud, and how the...
If cloud computing benefits are so clear, why have so few enterprises migrated their mission-critical apps? The answer is often inertia and FUD. No one ever got fired for not moving to the cloud – not yet. In his session at 15th Cloud Expo, Michael Hoch, SVP, Cloud Advisory Service at Virtustream, will discuss the six key steps to justify and execute your MCA cloud migration.
The 16th International Cloud Expo announces that its Call for Papers is now open. 16th International Cloud Expo, to be held June 9–11, 2015, at the Javits Center in New York City brings together Cloud Computing, APM, APIs, Security, Big Data, Internet of Things, DevOps and WebRTC to one location. With cloud computing driving a higher percentage of enterprise IT budgets every year, it becomes increasingly important to plant your flag in this fast-expanding business opportunity. Submit your speak...
Most of today’s hardware manufacturers are building servers with at least one SATA Port, but not every systems engineer utilizes them. This is considered a loss in the game of maximizing potential storage space in a fixed unit. The SATADOM Series was created by Innodisk as a high-performance, small form factor boot drive with low power consumption to be plugged into the unused SATA port on your server board as an alternative to hard drive or USB boot-up. Built for 1U systems, this powerful devic...
SYS-CON Events announced today that Gridstore™, the leader in software-defined storage (SDS) purpose-built for Windows Servers and Hyper-V, will exhibit at SYS-CON's 15th International Cloud Expo®, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Gridstore™ is the leader in software-defined storage purpose built for virtualization that is designed to accelerate applications in virtualized environments. Using its patented Server-Side Virtual C...
SYS-CON Events announced today that Cloudian, Inc., the leading provider of hybrid cloud storage solutions, has been named “Bronze Sponsor” of SYS-CON's 15th International Cloud Expo®, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. Cloudian is a Foster City, Calif.-based software company specializing in cloud storage. Cloudian HyperStore® is an S3-compatible cloud object storage platform that enables service providers and enterprises to bui...
SYS-CON Events announced today that TechXtend (formerly Programmer’s Paradise), a leading value-added provider of server and storage virtualization, and r-evolution will exhibit at SYS-CON's 15th International Cloud Expo®, which will take place on November 4–6, 2014, at the Santa Clara Convention Center in Santa Clara, CA. TechXtend (formerly Programmer’s Paradise) is a leading value-added provider of software, systems and solutions for corporations, government organizations, and academic instit...
Every healthy ecosystem is diverse. This is especially true in cloud ecosystems, where portability and interoperability are more important than old enterprise models of proprietary ownership. In his session at 15th Cloud Expo, Mark Baker, Server Product Manager at Canonical/Ubuntu, will discuss how single vendors used to take the lead in creating and delivering technology, but in a cloud economy, where users want tools of their preference, when and where they need them, it makes no sense.
The consumption economy is here and so are cloud applications and solutions that offer more than subscription and flat fee models and at the same time are available on a pure consumption model, which not only reduces IT spend but also lowers infrastructure costs, and offers ease of use and availability. In their session at 15th Cloud Expo, Ermanno Bonifazi, CEO & Founder of Solgenia, and Ian Khan, Global Strategic Positioning & Brand Manager at Solgenia, will discuss this shifting dynamic with ...
The emergence of cloud computing and Big Data warrants a greater role for the PMO to successfully manage enterprise transformation driven by these powerful trends. As the adoption of cloud-based services continues to grow, a governance model is needed to orchestrate enterprise cloud implementations and harness the power of Big Data analytics. In his session at 15th Cloud Expo, Mahesh Singh, President of BigData, Inc., to discuss how the Enterprise PMO takes center stage not only in developing th...
Cloud computing started a technology revolution; now DevOps is driving that revolution forward. By enabling new approaches to service delivery, cloud and DevOps together are delivering even greater speed, agility, and efficiency. No wonder leading innovators are adopting DevOps and cloud together! In his session at DevOps Summit, Andi Mann, Vice President of Strategic Solutions at CA Technologies, will explore the synergies in these two approaches, with practical tips, techniques, research dat...