1. Technical Field
The present invention relates generally to computer systems and in particular to distributed computer systems. Still more particularly, the present invention relates to efficient data collection within distributed computer systems as well as context-/resource-/capacity-based routing and dynamic workload management.
2. Description of the Related Art
Client-server distributed computing is becoming the standard computing topology because of quickly-evolving Internet development and associated e-practices, i.e., e-commerce, e-business, e-health, e-education, e-government, and e-everything practices. Client-server distributed computing is becoming even more important with the expansion of Web Services and grid utility computing.
Workload management is a key activity of distributed computing and the most important part of modern e-infrastructure. Conventional workload management in distributed computing has progressed through three different models. In a first model, a monitoring system (i.e., an attached computer system) collects data and the system administrator reviews the collected data to administrate the various computing devices. For example, a DB2 server stores a large amount of data, including query access plan and statistics for each table. However, with this first model, the DB2 server is not utilized to complete workload management and/or context-sensitive application routing.
In the second model, illustrated by
Finally, in the third model, illustrated by
Each of the above methods exhibit limitations that lead to inefficiency and in some cases bottlenecks in the overall system. For example, the first model is a manual process (i.e., no automatic collection and use of collected data) and is therefore not appropriate for current on-demand computing environments. Further, smart client (utilized by Bea Systems and Microsoft) of the second model utilizes only single client data, which is typically skewed due to problems already existing in the communication channels that are being monitored. Further, in typical systems, a single server may serve millions of clients, and the server life is substantially longer than client life, which is typically very short, (e.g., client life of 10 minutes compared with server life of 365 days). Thus client monitoring includes no history of the external network/system from which smart decisions may be made. Also, smart client is not able to complete context-based routing since smart client does not receive sufficient amounts of context data from across the system.
With the third model, the centralized controller needs to exchange data among all servers managed by the centralized server. The centralized server thus creates huge overhead and occasionally malfunctions when CPU-usage is high (e.g., over 90%) or memory usage is high. A list of limitations of a centralized controller also includes occasional congestion of the network, routing oscillation for dynamic WLM, single-source point of failure across the system (i.e., having one bad server operate as a single point of failure), and long latency when processing data at the central controller before transmitting the result to a requesting client. This last limitation is particularly troublesome in real-time on-demand systems. For example, in some applications, centralized WLM controller always delivers server weight too late in key time when server CPU usage is above 90% and/or server memory usage is above 90%. This late delivery occurs because the server is not able to receive timely weights from centralized WLM controller. Elected central controller produces the same problem, which is yet unresolved. Thus, to prevent this occurrence, server vendors typically advise their customers not to put too much load on servers and maintain 30-80% workload. This restriction/limitation on the server reduces server resource utilization and cause server instability.
With further application of Internet-based methods for more-efficient and expansive distributed computing environments (client-server computing), smart WLM becomes more and more important. However, as described above, previous workload management techniques have various limitations/inaccuracies that reduce the effectiveness of the computing environment. Companies are thus investing large amounts of money and resources on eWLM and other similar products, although the above described problems are yet unresolved.
Disclosed are a method, distributed-computing system, and computer program product for providing efficient workload management within a distributed computing environment. Each device within the distributed-computing environment is enhanced with a workload management controller (WLMC) functionality/utility, designed specifically for the type of device (i.e., client WLMC versus server WLMC) and utilized to collect process data about the particular device (e.g., status information) and about the device's interaction with the network. With the localized device-based WLM Controllers, each device utilizes fully distributed tagged information to accomplish capacity-based routing, context-based routing, and resource-based routing without any overhead or loss of data and without any network congestion. The distributed WLM Controller model enables each device to operate without concern for the level of CPU usage or memory usage of the particular device.
The above as well as additional objectives, features, and advantages of the present invention will become apparent in the following detailed written description.
The invention itself, as well as a preferred mode of use, further objects, and advantages thereof, will best be understood by reference to the following detailed description of an illustrative embodiment when read in conjunction with the accompanying drawings, wherein:
The present invention provides a method, distributed-computing system, and computer program product for providing efficient workload management within a distributed computing environment. Each device within the distributed-computing environment is enhanced with a workload management, controller (WLMC) functionality/utility, designed specifically for the type of device (i.e., client WLMC versus server WLMC) and utilized to collect process data about the particular device (e.g., status information) and about the device's interaction with the network. With the localized device-based WLM Controllers, each device utilizes fully distributed tagged information to accomplish capacity-based routing, context-based routing, and resource-based routing without any overhead or loss of data and without any network congestion. The distributed WLM Controller model enables each device to operate without concern for the level of CPU usage or memory usage of the particular device.
With reference now to
Referring now to
In addition to the above described hardware components of data processing system 100, several software and firmware components are also provided within data processing system 100 to enable the computer system to complete the device monitoring services, data collection, and routing server selection processes described herein. Among these software/firmware components are operating system (OS) 117 and WLMC utility 410. WLMC utility 410 is illustrated within memory 115. However, it is understood that, in alternate embodiments, WLMC utility 410 may be located on a removable computer readable medium or be provided as a component part of OS 117. When executed by processor 105, WLMC utility 410 executes a series of processes that provides the various functions described below, with reference to
Server WLMC 412 comprises server interceptor 460, which is a utility that is registered with each server to perform injection of server data into a response stream being sent to the client following receipt of a client request. Client interceptor 420 on the other hand is found within client WLMC 410 and is utilized to extract client request context, input the client request context to the router (connecting client to the network), and inject the client request context to the client's request stream.
In addition to server interceptor 460, server WLMC 412 also comprises Server Monitor and Local Data Collector (SMLDC) 465, which collects performance, capacity, and content data of the server's processes. Then, these locally collected server data are injected into the response stream through server interceptor 460.
Client Processes 450 may be processes within JVM 440. Additionally, client WLMC 410 may also be a component within JVM 440 or one of client processes 450. When the functions are provided within a Java environment, a process may sometimes be equal to JVM; however, the features of the invention may be implemented in environments utilizing C++ processes or processes provided by other programs.
Finally, client data merger 430 is a utility that merges data from the response streams of all servers to construct the full spectrum data. For example, as shown in
Referring again to
In one embodiment, the subsequent client request may also bring new merged data in the client side through client interceptor 420, and the server interceptor 460 peaks out these client-merged data and merges the client-merged data with the server's local data. This embodiment provides another cluster data merger component (CLDM) 470 in the server WLMC 412, which is not required for the other described embodiments. With this embodiment, however, all servers 415 are provided each others complete information without direct communications among servers 415 and without a centralized controller. For example, after three requests, client 405 has received all information of Server1, Server2, and Server3415 in addition to client 405 information. With the fourth client request, client 405 connects to Server1, and Server1 gets all information of client 405, Server2 (415), and Server3 (415) through this client 405 because the client 405 also brings in client's merged information to Server1 (415). In the next/subsequently-issued two requests, Server2 and Server3 also obtain the complete merge information, similar to Server1.
Accordingly, the fully-distributed design operates as an on-demand system, where servers only push directed information out to the network when the client requests that information. The client, meanwhile, is able to obtain all required content and context information to complete the client's scheduling and other processes by simply issuing a number of requests. Among the information obtained by the client are: (1) server capacity information to complete CPU/memory-based routing and provisioning; (2) server content to complete content-based routing; (3) client context to complete context-based routing; and (4) server resource to complete resource-based routing.
Thus, as depicted, the distributed model does not include a centralized controller and thus avoids overhead and congestion. The distributed model also does not malfunction when CPU/memory usage is high. The distributed WLMC controller model of the illustrative embodiments thus unifies and provides all of the above functions and features in a single utility.
According to the illustrative embodiments, a fully-distributed Client-Server WLMC method is provided that maximizes client-server interactions and resolves a substantial majority of previous problems found with single-WLMcontroller implementations. The fully-distributed WLM model allows the client to receive substantially more information, not only from its own monitoring/sensing of the network, but also from server-collected data. According to the illustrative embodiment, among the additional data retrieved from the server WLMCs are server-context (or resource) and server-content. Since a single server is able to serve millions of clients and the server life is significantly longer than client life, enabling the client to retrieve additional data from the server that have been accumulated for a long time by the server enables the client to perform/calculate more accurate analyses of workload management and routing-server selection.
The distributed WLMC method enables the client to be able to collect various kinds of client request contexts and information in addition to the client itself sensing network data/information. This further enables the client to perform context-based routing, which requires the client possess knowledge of the whole spectrum of request context, which the client is conventionally unable to experience. The use of a fully-distributed WLMC mechanism also provides the functionality of: (1) avoiding the congestion of the network; (2) avoiding routing oscillation with delivery of dynamic WLM; (3) avoiding having a single bad device that provides a single point of failure; (4) avoiding the long latency of having a single server with real-time on demand system; and (5) provide early delivery of server weight even when the server's CPU usage is above 90% or the server's memory usage is above 90%.
The invention does not provide any significant overhead at any one point since the management load is distributed throughout the system. According to the illustrative embodiment, the fully-distributed model substantially eliminates the need to exchange single messages between servers, and thus the model results in substantially no overhead. Thus, the fully-distributed model (design) remains functional even when CPU usage is 100% and/or memory usage is 100%. Further, the fully-distributed model utilizes both server capacity data and server content and client context data. Additionally, the fully-distributed design utilizes millions of other clients' context data in addition to the context data of the requesting client.
In one embodiment, a fully distributed client-server WLM mechanism that maximizes client-server interactions is provided. The client system collects data both from the client's own monitoring function as well as data received from each of the servers (collected at the server). The client pulls data from the server, where the data has accumulated from a long period of time by the server, with server-context and server-content (historical data). The client also collects all kinds of client request contexts and information to enable the client to complete context-based routing (knowing the whole spectrum of existing request context).
The client waits on return of the server(s) response(s) to the client's request, and WLMC utility determines at block 510 whether the server(s) response(s) have been received. When server response(s) are received, the WLMC utility parses the response for server content/data at block 512, and the WLMC utility constructs a full spectrum data by merging/combining the client data with the received server data at block 514. The WLMC utility evaluates the combined data to determine the overall workload of the network systems/devices at block 516, and then the WLMC utility selects (at block 517) one of the servers, using the corresponding workload calculations for each of the multiple servers, as the server to which a next client communication is routed through the network. The process then ends at block 518.
By utilizing the fully-distributed design of the illustrative embodiments, the client is able to intercept and inject and transport WLMC data between client and server. The above described features of the invention may be implemented in one of several ways, depending on the protocol being utilized. These protocols and associated mechanisms include: (1) in http protocol by http filter mechanism; (2) in iiop protocol by CORBA ContextService mechanism; and (3) in straight java socket by writing Request/Response head with tagged information. The client completes these operations by (1) http filter for HTTP protocol; (2) Corba ContextService in iiop protocol; or (3) tagging to payload in any socket communication, each without modifying existing routing protocol. Thus, there is no requirement for changes to any existing protocol, i.e., no additional traffic is needed to transport WLMC data between client and server, since the WLMC data are tagged into the request/response stream as additional data.
As a final matter, it is important that while an illustrative embodiment of the present invention has been, and will continue to be, described in the context of a fully functional computer system with installed management software, those skilled in the art will appreciate that the software aspects of an illustrative embodiment of the present invention are capable of being distributed as a program product in a variety of forms, and that an illustrative embodiment of the present invention applies equally regardless of the particular type of signal bearing media used to actually carry out the distribution. Examples of signal bearing media include recordable type media such as floppy disks, hard disk drives, CD ROMs, and transmission type media such as digital and analogue communication links.
While the invention has been particularly shown and described with reference to a preferred embodiment, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention.