None.
This invention was not developed in conjunction with any Federally sponsored contract.
Not applicable.
None.
1. Field of the Invention
The present invention relates to systems and methods for determining causes and sources of problems, errors, and inefficiencies in service oriented architecture computing environments.
2. Background of the Invention
Whereas the determination of a publication, technology, or product as prior art relative to the present invention requires analysis of certain dates and events not disclosed herein, no statements made within this Background of the Invention shall constitute an admission by the Applicant of prior art unless the term “Prior Art” is specifically stated. Otherwise, all statements provided within this Background section are “other information” related to or useful for understanding the invention.
In today's Information Technology (“IT”) system management environment, basic resource monitoring is becoming a commodity. For systems management companies to remain competitive they must move up the monitoring and management stack. Certain IT products, such as Tivoli's™ IT Service Management (“ITSM”) and IT Infrastructure Library (“ITIL”) provide a mechanism and methodology to achieve this. In the IT industry today, customers are moving their mission critical applications onto the Internet and providing them as services in a service oriented architecture (“SOA”) as to enable tighter integration. SOA is a well known style of computing environments which covers all aspects of developing, deploying, and using business processes which are accessed as “services”, for the entire lifecycle of each service.
The advantage of providing business transactions as SOA services (“SOAs”) is that it de-couples the business transaction from the underlying IT infrastructure and technical implementation that drives the transaction. Unfortunately, this makes management of such an environment even more complicated because the link between the business transaction and the IT resource that is servicing that transaction is not clear.
A challenge in SOA management has become how to determine why a SOA transaction is not available or not performing up to its defined performance level, especially to a contractual service level such as a Service Level Agreement (“SLA”). The question of “What resource is causing the end user problem and why?” can plague administrators, but the complexity and fluidity of the SOA arrangement can make problem source determination incredibly difficult.
This is an industry wide problem and many of the available systems management products attempt to provide solutions. One such example is IBM Tivoli Monitoring (“ITM”) which is a resource monitoring product that monitors the individual servers and the metrics of the applications running on those servers. However, ITM does this without the context of the SOA transaction or the business impact of the application being monitored.
Another specific systems management product is IBM Tivoli Composite Application Manager (“ITCAM”) which monitors SOA transactions and tracks the transaction as it flows across the IT infrastructure. ITCAM dynamically discovers the IT resources involved in a SOA transaction and correlates which physical resource is the root cause of the response time problem, but it does not correlate the business impact to the application specific resource metric that caused the problem.
There are many other products in this market space that attempt to provide solutions to this problem, but do so in fragmented and incomplete ways. Another similar challenge is related to how companies currently attempt to manage this type of problem. Currently, many companies establish large management infrastructures and operations centers where they funnel all system and application events from all monitored applications. When a problem occurs, the Operations staff quickly becomes bombarded with thousands of IT resource system events indicating that there is some type of IT problem. It is up to the Operations staff to filter through these events and to attempt to understand which events impact the business and which are just “noise” (e.g. which events have little or no actual business impact). “Business impact” can be defined in many ways. In the case of the present invention, we are referring to business transaction response time and availability from an end user perspective. If a business cannot provide its online SOA transactions to its customers in a timely fashion, its business is directly impacted. But, the concepts and problems addressed herein are general to many of the broader definitions of business impact. Whether companies are attempting to service release management, configuration management, change control management, etc., there are always a vast set of key business metrics that are impacted by the IT infrastructure. The challenge, therefore, is to discover how to identify which specific IT metrics and resources are impacting the business adversely so that the issues can be effectively addressed.
A potential multicomputer related problem is predicted and reported by determining a set of computer resources and relationships there between needed to complete a multicomputer business transaction, retrieving performance monitoring metrics for the computer resources during executions of the multicomputer transaction, dynamically deriving correlations between the resource relationships and the performance metrics, comparing a trend of the correlations to one or more service level requirements to predict one or more potential future violations of a business transaction requirement, including identification of one or more related resources likely to cause the violation, and reporting such prediction and likely case to an administrator.
The following detailed description when taken in conjunction with the figures presented herein provide a complete disclosure of the invention.
a and 2b show a generalized computing platform architecture, and a generalized organization of software and firmware of such a computing platform architecture.
a and 3b show examples of raw dissimilar metrics data and normalized dissimilar metrics data.
a-4c, illustrate computer readable media of various removable and fixed types, signal transceivers, and parallel-to-serial-to-parallel signal circuits.
a-5c illustrate topologies of business transaction systems and their sources of error messages and monitoring metrics.
The inventor of the present invention have recognized and solved problems previously unrecognized by others in the art of managing SOA-based computing arrangements. The inventor has recognized that a mechanism is needed which provides a complete end-to-end solution that can autonomically link IT resource metrics to SOA business transaction performance problems in order to allow SOA systems owners (“customers”) to quickly identify the root cause of a SOA transaction problem.
Turning to
Finally, a “back tier” (53) includes many larger servers and mainframes, such as large computers running well-known operating systems and applications such as z/OS, CICS, Linux, UNIX, and IMS, as well as many associated databases. Most of the application servers and back tier systems are also outfitted with a messaging queue handler (“MSG-Q”), such as IBM's MQ series messaging product.
Present day correlation technology is event based. Alerts are only generated when a problem is detected. An administrator must define rules to correlate related events to sources or causes of problems. And, according to present day approaches, resource relationships are hard coded and defined by the administrator, in which a user defines systems related by business boundaries. But, these SOA systems frequently change roles and functions, making previous manually-defined relationships obsolete.
Further, present day expert advice and situation based thresholding and alerting are hard coded and based on generalized observations from previous customer experiences and what “should” happen.
Turning to
The present invention provides an autonomic correlation engine that utilizes the existing resource relationships defined in monitoring products such as Tivoli's Configuration Management Database (“CMDB”) and the real time resource metrics defined in a data warehouse, like that provided by the Tivoli DataWarehouse and ITM, to dynamically discover and link SOA business transaction performance with the IT resource metrics that caused the business transaction violation. The invention's method for discovering such relationships is unique and provides significant business advantage to any company or product that can provide this capability as a SOA management solution.
To this end, the present invention provides these advantages and functions:
Turning to
As a result of the analysis and predictive methods of the invention, described in detail in the following paragraphs, it may be signaled to an administrator that because resource monitoring has detected a significant decrease in free tablespace on the sixth server below the normal level of free tablespace, business transactions of type 1 which use the affected database may exceed a performance threshold within 30 minutes according to tablespace usage trends.
According to one embodiment of the invention, the correlation agent provides the following outputs from such an analysis:
A generalized monitoring procedure according to the invention is shown in
As a result, customers are no longer required to set predefined thresholds, the system simply detects & reports abnormalities in the monitored data.
Turning to
In order to effectively compare dissimilar metrics, a method was developed to render the metrics to a form which is readily and meaningfully comparable.
However, by normalizing all three metrics data to a common range, such as 0 to 1, as shown in
One can see that once the data is normalized, it is more straightforward and meaningful to calculate a correlation coefficient to determine if the metrics are directly or inversely related to the key metric of SOA transaction response time. From this example of
Such normalization, as previously disclosed as a step in a larger logical process, is useful for the present invention whereas the many metrics to be compared and monitored are often dissimilar in units and quantity ranges.
As shown in
It will be recognized by those skilled in the art that many schemas may be adopted for use with the logical processes of the present invention. By way of more complete illustration of an example embodiment of the present invention, one possible schema for such a database is (e.g. the column names and data types for the fields in each metric record or row):
According to other embodiments of the invention, such an application may further provide the following functionality:
Whereas at least one embodiment of the present invention incorporates, uses, or operates on, with, or through one or more computing platforms, and whereas many devices, even purpose-specific devices, are actually based upon computing platforms of one type or another, it is useful to describe a suitable computing platform, its characteristics, and its capabilities.
Therefore, it is useful to review a generalized architecture of a computing platform which may span the range of implementation, from a high-end web or enterprise server platform, to a personal computer, to a portable PDA or wireless phone.
In one embodiment of the invention, the functionality including the previously described logical processes are performed in part or wholly by software executed by a computer, such as personal computers, web servers, web browsers, or even an appropriately capable portable computing platform, such as personal digital assistant (“PDA”), web-enabled wireless telephone, or other type of personal information management (“PIM”) device. In alternate embodiments, some or all of the functionality of the invention are realized in other logical forms, such as circuitry.
Turning to
Many computing platforms are also provided with one or more storage drives (29), such as hard-disk drives (“HDD”), floppy disk drives, compact disc drives (CD, CD-R, CD-RW, DVD, DVD-R, etc.), and proprietary disk and tape drives (e.g., I omega Zip™ and Jaz™, Addonics SuperDisk™, etc.). Additionally, some storage drives may be accessible over a computer network.
Many computing platforms are provided with one or more communication interfaces (210), according to the function intended of the computing platform. For example, a personal computer is often provided with a high speed serial port (RS-232, RS-422, etc.), an enhanced parallel port (“EPP”), and one or more universal serial bus (“USB”) ports. The computing platform may also be provided with a local area network (“LAN”) interface, such as an Ethernet card, and other high-speed interfaces such as the High Performance Serial Bus IEEE-1394.
Computing platforms such as wireless telephones and wireless networked PDA's may also be provided with a radio frequency (“RF”) interface with antenna, as well. In some cases, the computing platform may be provided with an infrared data arrangement (“IrDA”) interface, too.
Computing platforms are often equipped with one or more internal expansion slots (211), such as Industry Standard Architecture (“ISA”), Enhanced Industry Standard Architecture (“EISA”), Peripheral Component Interconnect (“PCI”), or proprietary interface slots for the addition of other hardware, such as sound cards, memory boards, and graphics accelerators.
Additionally, many units, such as laptop computers and PDA's, are provided with one or more external expansion slots (212) allowing the user the ability to easily install and remove hardware expansion devices, such as PCMCIA cards, SmartMedia cards, and various proprietary modules such as removable hard drives, CD drives, and floppy drives.
Often, the storage drives (29), communication interfaces (210), internal expansion slots (211) and external expansion slots (212) are interconnected with the CPU (21) via a standard or industry open bus architecture (28), such as ISA, EISA, or PCI. In many cases, the bus (28) may be of a proprietary design.
A computing platform is usually provided with one or more user input devices, such as a keyboard or a keypad (216), and mouse or pointer device (217), and/or a touch-screen display (218). In the case of a personal computer, a full size keyboard is often provided along with a mouse or pointer device, such as a track ball or TrackPoint™. In the case of a web-enabled wireless telephone, a simple keypad may be provided with one or more function-specific keys. In the case of a PDA, a touch-screen (218) is usually provided, often with handwriting recognition capabilities.
Additionally, a microphone (219), such as the microphone of a web-enabled wireless telephone or the microphone of a personal computer, is supplied with the computing platform. This microphone may be used for simply reporting audio and voice signals, and it may also be used for entering user choices, such as voice navigation of web sites or auto-dialing telephone numbers, using voice recognition capabilities.
Many computing platforms are also equipped with a camera device (2100), such as a still digital camera or full motion video digital camera.
One or more user output devices, such as a display (213), are also provided with most computing platforms. The display (213) may take many forms, including a Cathode Ray Tube (“CRT”), a Thin Flat Transistor (“TFT”) array, or a simple set of light emitting diodes (“LED”) or liquid crystal display (“LCD”) indicators.
One or more speakers (214) and/or annunciators (215) are often associated with computing platforms, too. The speakers (214) may be used to reproduce audio and music, such as the speaker of a wireless telephone or the speakers of a personal computer. Annunciators (215) may take the form of simple beep emitters or buzzers, commonly found on certain devices such as PDAs and PIMs.
These user input and output devices may be directly interconnected (28′, 28″) to the CPU (21) via a proprietary bus structure and/or interfaces, or they may be interconnected through one or more industry open buses such as ISA, EISA, PCI, etc. The computing platform is also provided with one or more software and firmware (2101) programs to implement the desired functionality of the computing platforms.
Turning to now
Additionally, one or more “portable” or device-independent programs (224) may be provided, which must be interpreted by an OS-native platform-specific interpreter (225), such as Java™ scripts and programs.
Often, computing platforms are also provided with a form of web browser or micro-browser (226), which may also include one or more extensions to the browser such as browser plug-ins (227).
The computing device is often provided with an operating system (220), such as Microsoft Windows™, UNIX, IBM OS/2™, IBM AIX™, open source LINUX, Apple's MAC OS™, or other platform specific operating systems. Smaller devices such as PDA's and wireless telephones may be equipped with other forms of operating systems such as real-time operating systems (“RTOS”) or Palm Computing's PalmOS™.
A set of basic input and output functions (“BIOS”) and hardware device drivers (221) are often provided to allow the operating system (220) and programs to interface to and control the specific hardware functions provided with the computing platform.
Additionally, one or more embedded firmware programs (222) are commonly provided with many computing platforms, which are executed by onboard or “embedded” microprocessors as part of the peripheral device, such as a micro controller or a hard drive, a communication processor, network interface card, or sound or graphics card.
As such,
In another embodiment of the invention, logical processes according to the invention and described herein are realized in computer program code encoded on or in one or more computer-readable media. Some computer-readable media are read-only (e.g. they must be initially programmed using a different device than that which is ultimately used to read the data from the media), some are write-only (e.g. from the data encoders perspective they can only be encoded, but not read simultaneously), or read-write. Still some other media are write-once, read-many-times.
Some media are relatively fixed in their mounting mechanisms, while others are removable, or even transmittable. All computer-readable media form two types of systems when encoded with data and/or computer software: (a) when removed from a drive or reading mechanism, they are memory devices which generate useful data-driven outputs when stimulated with appropriate electromagnetic, electronic, and/or optical signals; and (b) when installed in a drive or reading device, they form a data repository system accessible by a computer.
a illustrates some computer readable media including a computer hard drive (40) having one or more magnetically encoded platters or disks (41), which may be read, written, or both, by one or more heads (42). Such hard drives are typically semi-permanently mounted into a complete drive unit, which may then be integrated into a configurable computer system such as a Personal Computer, Server Computer, or the like.
Similarly, another form of computer readable media is a flexible, removable “floppy disk” (43), which is inserted into a drive which houses an access head. The floppy disk typically includes a flexible, magnetically encodable disk which is accessible by the drive head through a window (45) in a sliding cover (44).
A Compact Disk (“CD”) (46) is usually a plastic disk which is encoded using an optical and/or magneto-optical process, and then is read using generally an optical process. Some CD's are read-only (“CD-ROM”), and are mass produced prior to distribution and use by reading-types of drives. Other CD's are writable (e.g. “CD-RW”, “CD-R”), either once or many time. Digital Versatile Disks (“DVD”) are advanced versions of CD's which often include double-sided encoding of data, and even multiple layer encoding of data. Like a floppy disk, a CD or DVD is a removable media.
Another common type of removable media are several types of removable circuit-based (e.g. solid state) memory devices, such as Compact Flash (“CF”) (47), Secure Data (“SD”), Sony's MemoryStick, Universal Serial Bus (“USB”) FlashDrives and “Thumbdrives” (49), and others. These devices are typically plastic housings which incorporate a digital memory chip, such as a battery-backed random access chip (“RAM”), or a Flash Read-Only Memory (“FlashROM”). Available to the external portion of the media is one or more electronic connectors (48, 400) for engaging a connector, such as a CF drive slot or a USB slot. Devices such as a USB FlashDrive are accessed using a serial data methodology, where other devices such as the CF are accessed using a parallel methodology. These devices often offer faster access times than disk-based media, as well as increased reliability and decreased susceptibility to mechanical shock and vibration. Often, they provide less storage capability than comparably priced disk-based media.
Yet another type of computer readable media device is a memory module (403), often referred to as a SIMM or DIMM. Similar to the CF, SD, and FlashDrives, these modules incorporate one or more memory devices (402), such as Dynamic RAM (“DRAM”), mounted on a circuit board (401) having one or more electronic connectors for engaging and interfacing to another circuit, such as a Personal Computer motherboard. These types of memory modules are not usually encased in an outer housing, as they are intended for installation by trained technicians, and are generally protected by a larger outer housing such as a Personal Computer chassis.
Turning now to
In general, a microprocessor or microcontroller (406) reads, writes, or both, data to/from storage for data, program, or both (407). A data interface (409), optionally including a digital-to-analog converter, cooperates with an optional protocol stack (408), to send, receive, or transceive data between the system front-end (410) and the microprocessor (406). The protocol stack is adapted to the signal type being sent, received, or transceived. For example, in a Local Area Network (“LAN”) embodiment, the protocol stack may implement Transmission Control Protocol/Internet Protocol (“TCP/IP”). In a computer-to-computer or computer-to-peripheral embodiment, the protocol stack may implement all or portions of USB, “FireWire”, RS-232, Point-to-Point Protocol (“PPP”), etc.
The system's front-end, or analog front-end, is adapted to the signal type being modulated, demodulate, or transcoded. For example, in an RF-based (413) system, the analog front-end comprises various local oscillators, modulators, demodulators, etc., which implement signaling formats such as Frequency Modulation (“FM”), Amplitude Modulation (“AM”), Phase Modulation (“PM”), Pulse Code Modulation (“PCM”), etc. Such an RF-based embodiment typically includes an antenna (414) for transmitting, receiving, or transceiving electromagnetic signals via open air, water, earth, or via RF wave guides and coaxial cable. Some common open air transmission standards are BlueTooth, Global Services for Mobile Communications (“GSM”), Time Division Multiple Access (“TDMA”), Advanced Mobile Phone Service (“AMPS”), and Wireless Fidelity (“Wi-Fi”).
In another example embodiment, the analog front-end may be adapted to sending, receiving, or transceiving signals via an optical interface (415), such as laser-based optical interfaces (e.g. Wavelength Division Multiplexed, SONET, etc.), or infra Red Data Arrangement (“IrDA”) interfaces (416). Similarly, the analog front-end may be adapted to sending, receiving, or transceiving signals via cable (412) using a cable interface, which also includes embodiments such as USB, Ethernet, LAN, twisted-pair, coax, Plain-old Telephone Service (“POTS”), etc.
Signals transmitted, received, or transceived, as well as data encoded on disks or in memory devices, may be encoded to protect it from unauthorized decoding and use. Other types of encoding may be employed to allow for error detection, and in some cases, correction, such as by addition of parity bits or Cyclic Redundancy Codes (“CRC”). Still other types of encoding may be employed to allow directing or “routing” of data to the correct destination, such as packet and frame-based protocols.
c illustrates conversion systems which convert parallel data to and from serial data. Parallel data is most often directly usable by microprocessors, often formatted in 8-bit wide bytes, 6-bit wide words, 32-bit wide double words, etc. Parallel data can represent executable or interpretable software, or it may represent data values, for use by a computer. Data is often serialized in order to transmit it over a media, such as a RF or optical channel, or to record it onto a media, such as a disk. As such, many computer-readable media systems include circuits, software, or both, to perform data serialization and re-parallelization.
Parallel data (421) can be represented as the flow of data signals aligned in time, such that parallel data unit (byte, word, d-word, etc.) (422, 423, 424) is transmitted with each bit D0-Dn being on a bus or signal carrier simultaneously, where the “width” of the data unit is n−1. In some systems, D0 is used to represent the least significant bit (“LSB”), and in other systems, it represents the most significant bit (“MSB”). Data is serialized (421) by sending one bit at a time, such that each data unit (422, 423, 424) is sent in serial fashion, one after another, typically according to a protocol.
As such, the parallel data stored in computer memory (407, 407′) is often accessed by a microprocessor or Parallel-to-Serial Converter (425, 425′) via a parallel bus (421), and exchanged (e.g. transmitted, received, or transceived) via a serial bus (421′). Received serial data is converted back into parallel data before storing it in computer memory, usually. The serial bus (421′) generalized in
In these manners, various embodiments of the invention may be realized by encoding software, data, or both, according to the logical processes of the invention, into one or more computer-readable mediums, thereby yielding a product of manufacture and a system which, when properly read, received, or decoded, yields useful programming instructions, data, or both, including, but not limited to, the computer-readable media types described in the foregoing paragraphs.
While certain examples and details of a preferred embodiment have been disclosed, it will be recognized by those skilled in the art that variations in implementation such as use of different programming methodologies, computing platforms, and processing technologies, may be adopted without departing from the spirit and scope of the present invention. Therefore, the scope of the invention should be determined by the following claims.