Processor power dissipation has become an issue for processors of all types, from low end mobile processers to high end server processors. Among processor components, a cache memory is a major portion of a processor's area and transistor counts, and consumes significant leakage power. For example, for a typical commercially available multicore processor, 40% of total leakage power is due to a last level cache (LLC) and interconnect.
While reducing a cache's leakage power by turning off portions of a cache memory may reduce processor power consumption, it is practically difficult to turn off even portions of a LLC as it is typically implemented as a shared memory structure in which a portion of all memory addresses of a system is statically mapped to each LLC portion. As such, even if one core of a multicore processor is operating, all LLC slices are active to service memory requests mapped to the slices. And thus there are limited power saving opportunities for a cache memory in current processors.
In various embodiments, a multicore processor such as a multi-tile chip multiprocessor (CMP) may be provided with a multi-level cache memory hierarchy. In this hierarchy each of the levels of cache memory may be arranged as private portions each associated with a given core or other caching agent. In this way, a private cache organization is realized to allow embodiments to enable local LLC portions or slices associated with a caching agent to be dynamically power gated in certain circumstances. As will be described herein, this dynamic power gating of an LLC slice may occur when at least one of the following situations is present: (i) the associated core is in a low power state (such as a given sleep state); and (ii) a lower cache level in the hierarchy (e.g., a mid-level cache (MLC)) provides sufficient capacity for execution of an application or other workload running on the core. In this way, total chip power may be significantly reduced. Although the embodiments described herein are with regard to power gating of a LLC, understand that in other embodiments a different level of a multi-level cache hierarchy can be power gated.
Referring now to
As seen, processor 110 may be a single die processor including multiple tiles 120a-120n. Each tile includes a processor core and an associated private cache memory hierarchy that, in some embodiments is a three-level hierarchy with a low level cache, a MLC and an LLC. In addition, each tile may be associated with an individual voltage regulator 125a-125n. Accordingly, a fully integrated voltage regulator (FIVR) implementation may be provided to allow for fine-grained control of voltage and thus power and performance of each individual tile. As such, each tile can operate at an independent voltage and frequency, enabling great flexibility and affording wide opportunities for balancing power consumption with performance. Of course, embodiments apply equally to a processor package not having integrated voltage regulators.
Still referring to
Also shown is a power control unit (PCU) 138, which may include hardware, software and/or firmware to perform power management operations with regard to processor 110. In various embodiments, PCU 138 may include logic to perform adaptive local LLC power control in accordance with an embodiment of the present invention. Furthermore, PCU 138 may be coupled via a dedicated interface to external voltage regulator 160. In this way, PCU 138 can instruct the voltage regulator to provide a requested regulated voltage to the processor.
While not shown for ease of illustration, understand that additional components may be present within processor 100 such as uncore logic, and other components such as internal memories, e.g., an embedded dynamic random access memory (eDRAM), and so forth. Furthermore, while shown in the implementation of
Although the following embodiments are described with reference to energy conservation and energy efficiency in specific integrated circuits, such as in computing platforms or processors, other embodiments are applicable to other types of integrated circuits and logic devices. Similar techniques and teachings of embodiments described herein may be applied to other types of circuits or semiconductor devices that may also benefit from better energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to any particular type of computer systems, and may be also used in other devices, such as handheld devices, systems on chip (SoCs), and embedded applications. Some examples of handheld devices include cellular phones, Internet protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications typically include a microcontroller, a digital signal processor (DSP), network computers (NetPC), set-top boxes, network hubs, wide area network (WAN) switches, or any other system that can perform the functions and operations taught below. Moreover, the apparatus', methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimizations for energy conservation and efficiency. As will become readily apparent in the description below, the embodiments of methods, apparatus', and systems described herein (whether in reference to hardware, firmware, software, or a combination thereof) are vital to a ‘green technology’ future, such as for power conservation and energy efficiency in products that encompass a large portion of the US economy.
Note that the power savings realized by local LLC power control described herein may be independent of and complementary to an operating system (OS)-based mechanism, such as the Advanced Configuration and Platform Interface (ACPI) standard (e.g., Rev. 3.0b, published Oct. 10, 2006). According to ACPI, a processor can operate at various performance states or levels, namely from P0 to PN. In general, the P1 performance state may correspond to the highest guaranteed performance state that can be requested by an OS. In addition to this P1 state, the OS can further request a higher performance state, namely a P0 state. This P0 state may thus be an opportunistic state in which, when power and/or thermal budget is available, processor hardware can configure the processor or at least portions thereof to operate at a higher than guaranteed frequency. In many implementations a processor can include multiple so-called bin frequencies above a guaranteed maximum frequency, also referred to as a P1 frequency. In addition, according to ACPI, a processor can operate at various power states or levels. With regard to power states, ACPI specifies different power consumption states, generally referred to as C-states, C0, C1 to Cn states. When a core is active, it runs at a C0 state, and when the core is idle it may be placed in a core low power state, also called a core non-zero C-state (e.g., C1-C6 states), with each C-state being at a lower power consumption level (such that C6 is a deeper low power state than C1, and so forth). When all cores of a multicore processor are in a core low power state, the processor can be placed in a package low power state, such as a package C6 low power state. This package low power state is a deeper low power state than a core C6 state, since additional processor circuitry can be turned off such as certain clock generation circuitry including phase locked loops (PLLs).
Referring now to
Core 210 may couple to a first level cache 220 that may be a relatively small cache memory closely associated with the core. In turn, cache 220 is coupled to a mid-level cache (MLC) 230 that may include greater storage capacity. In turn, MLC 230 is coupled to a further portion of a cache hierarchy, namely a last level cache (LLC) 240 that may include even greater amounts of storage capacity. In some embodiments, this cache hierarchy formed of cache 220, MLC 230, and LLC 240 may be an inclusive cache hierarchy such that all of the information stored in cache 220 is present in cache 230, and all of the information stored in cache 230 is present in LLC 240. Of course, when LLC 240 is power gated when the core is active, the hierarchy is non-inclusive.
Also understand that this tile-included cache hierarchy includes slices or portions of a shared cache implemented with LLC 240 of a larger cache memory structure. In other words, the LLC may be a portion of a larger cache structure of the level distributed access the tiles. In addition, tile 200 includes a router 250 and a directory 260. In general, router 250 may be used to route communications to and from other components of the CMP such as one or more other tiles and/or other agents of the processor. In turn, directory 260 may be part of a distributed directory protocol to enable cache coherent operation within the system. Although shown at this high level in the embodiment of
Referring now to
As further seen in
If the decision is to power gate the LLC, the information stored in the LLC can be flushed, e.g., to another cache hierarchy of another tile, another storage within the processor package such as an eDRAM or other internal storage. Or the information can be flushed to a system memory, via memory controller 285. Although shown at this high level in the embodiment of
Note that embodiments also apply to other processor architectures such as a clustered cache organization, e.g., a cluster on die, where the unit of LLC power management is a cluster of LLC slices instead of one individual LLC slice. Throughout this discussion, the terms “core” and “a cluster of cores” are used interchangeably.
To determine whether power savings via LLC power gating are possible, embodiments may continuously monitor core activity and cache behavior to detect certain scenarios. In an embodiment, two different scenarios may be detected. First, when a core is to be placed into a certain low power state (e.g., entering a C6 state), potentially its LLC can be flushed and turned off to save power. To determine whether a power savings is possible, one or more run-time conditions such as certain performance metrics may be analyzed to make an intelligent decision on whether an upcoming idle duration can amortize the cost of flushing and turning off the LLC slice. Second, for a core that is in an active state, other performance metrics can be analyzed such as misses being sent out of a MLC. If the number of misses generated by the MLC within a period of time is less than a threshold, the LLC may be flushed and power gated, as this is an indication that the MLC has sufficient cache capacity to handle a given workload.
In an embodiment, the run-time conditions to be analyzed include the workload characteristics, number of ways that are enabled, misses of the MLC per thousand instructions (mpki), and upcoming idle duration estimation. In some embodiments, quality of service (QoS) requirements of an OS/application also may be considered, since power gating the LLC introduces a longer resume latency (e.g., due to loading the data from system memory). Note that in some embodiments, this delay can be mitigated by storing data in other under-utilized LLC slices or in another memory such as an eDRAM.
Embodiments may make a determination as to whether to flush and power gate a LLC slice when a core enters a low power state since flushing/reloading cache contents and entering/exiting a given low power state have certain overhead associated with the actions. Thus, it can be determined whether the LLC slice is to stay in the low power state for at least a minimum amount of time, referred to herein as an energy break even time (EBET), to compensate for the energy overhead for the transitions (in and out). Moreover, flushing a cache slice and entering a low power state may also create a small amount of delay when the core wakes up, and thus may impact responsiveness and performance. As such, embodiments may only flush and turn off a LLC slice when the expected idle duration is long enough to amortize the overhead (as measured by the EBET), and also when a certain performance level is maintained.
Further, if the core is running, the MLC might offer sufficient cache capacity for a short period of time until the application changes its behavior. Hence, this application behavior may also be detected such that the LLC is only flushed if the MLC provides enough capacity and a workload's memory needs are relatively stable, to prevent oscillation.
In an embodiment, EBET is defined as the minimum time for the LLC to transition into and remain in a given lower power state (including the entry and exit transition time), to compensate for the cost of entering and exiting this state from a powered on state. According to an embodiment, this EBET is based both on static and dynamic conditions, including resume and exit latency, power consumption during on, transition and off states, and dynamic run-time conditions, such as cache size used. For example, when an opened cache size is very small, the overhead of flushing and reloading the data is typically small, thus reducing the total overhead.
In one embodiment, the following factors may be considered in determining an adaptive EBET value: (1) enter and exit latency (L); (2) power during enter and exit period (Pt); (3) LLC on power (Pm), and LLC leakage power (Poff) when the LLC is power gated; (4) core power Pc during reloading time, and memory related power Pm during flushing and reloading time; and (5) content flushing overhead Of and content reloading overhead Or, including system memory access overhead, e.g., as obtained from a memory controller. This total overhead is O=Of+Or.
Flushing and reloading cache content take time, introducing overhead. How long it takes, however, depends on the workload characteristic. And the energy consumption overhead also depends on system characteristics such as memory and memory controller power consumption. The first 4 factors are design parameters that may be stored in a configuration storage of the processor, while the fifth factor depends on run time dynamics. If a given workload consumes only a small number of cache ways, the flush time and reload time will be small as well. Moreover, for a workload that does not reuse data frequently (such as a streaming application), the cache flushing and reloading cost is relatively low since new data will be loaded anyway, regardless of flushing or not. This workload determination can be measured, in an embodiment, using LLC miss rate. If the miss rate is high, the application exhibits a streaming characteristic.
After the overhead O is determined, the EBET can be calculated as follows:
EBET*Pon=O*Pon+L*Pt+(EBET−O−L)*Poff+Pc*Or+Pm*O.
And it yields:
EBET=(O*Pon+Pc*Or+Pm*O+L*Pt−(O−L)*Poff)/(Pon−Poff).
In order to simplify the design, Of and Or can be replaced by a design parameter to represent the average flushing and reloading overhead. Note that an EBET value can be updated periodically and dynamically based on the workload characteristics. If the workload behavior is stable, the EBET can be updated less frequently to reduce calculation overhead; otherwise the EBET can be updated more frequently. As an example, for a relatively stable workload, the EBET value can be updated approximately every several minutes. Instead with a dynamically changing workload, the EBET value can be updated approximately every several seconds or even less.
After the break even time EBET is determined, a length of the upcoming idle duration may be compared with the EBET value to determine whether it saves energy to flush and shut down the LLC slice. QoS factors may be considered in this analysis, in an embodiment.
There are multiple techniques to obtain an estimate of the upcoming idle duration. One exemplary technique includes an auto demotion method in which a controller observes a few idle durations in the near past. If these durations are smaller than the adaptive EBET when the LLC was flushed, then the controller automatically demotes the processor to a low power state (such as a core C4 state) in which the cache is not power gated.
Another exemplary technique to estimate upcoming idle duration includes using one or more device idle duration reports and other heuristics. These device idle duration reports may be received from devices coupled to the processor such as one or more peripheral or IO devices. Based on this information and heuristic information, the next idle duration can be predicted, e.g., by a control logic. Or a control logic can predict the idle duration based on a recent history of instances of idle durations such as via a moving filter that tracks the trend and provides a reasonable estimation of the idle duration.
A still further exemplary technique for power management is to power gate a LLC based on a measurement of MLC utilization to determine whether the LLC is needed or not. That is, if the MLC fulfills a workload's memory needs, the number of misses generated for a given number of instructions should be very little. Hence, in order to measure MLC performance, a metric can be monitored, such as the number of misses in the MLC per thousand instructions (mpki), which can be obtained from one or more counters associated with the LLC. If this metric (or another appropriate metric) is consistently below a given threshold (e.g., a th_mpki) for an evaluation period of time that is more than the EBET, then a determination may be made to flush and power gate the LLC. Note that this threshold may be a value above zero. That is, even if MLC is generating misses, but the misses are generated at a very slow pace, it may be more beneficial to power gate the LLC and let those misses be serviced by a further portion of a memory hierarchy instead of keeping the LLC consuming on and leakage power to service such a slow request rate. In some embodiments, this threshold (th_mpki) may depend on the LLC leakage and the dynamic power difference between fetching data from the LLC as compared to fetching data from memory.
Next, a decision may be made as to whether an LLC slice should be power gated when an associated core or cluster of cores enters into a given low power state. Independently, MLC utilization may be monitored to determine whether the LLC can be power gated even when the core is running.
In an embodiment, QoS information may be considered to prevent too frequent deep low power entries to avoid a performance impact. QoS information can be conveyed in different manners. For example, the QoS information may indicate whether the processor is allowed to enter a given low power state and/or by a QoS budget that defines how many deep low power state entries can take place in a unit of time. In another embodiment, QoS information can also be obtained from an energy performance bias value to indicate a high level direction on what type of performance/power balance a user is expecting.
Referring now to
As shown in
Next control passes to block 335 where the EBET may be determined. Although in some embodiments this value may be dynamically calculated, it is also possible to use a cached value of this EBET for performing the analysis. As an example, a fixed EBET may be stored in a configuration storage, or a cached EBET previously calculated by the logic can be accessed.
Still referring to
Referring now to the right branch of method 300, it may begin at block 360 by receiving performance metric information of a cache, e.g., of a MLC. Although the scope of the present invention is not limited in this regard, one performance metric that may be monitored is a miss rate which in an embodiment may be measured as an mpki value. Next control passes to diamond 365 to determine whether this miss rate is less than a threshold value. In an embodiment this threshold may correspond to a threshold value of misses that indicates relatively low utilization of an LLC coupled to the MLC. If this miss value is less than this threshold value, control passes to block 370 where a counting of an evaluation duration may be performed, either from an initial value or continuing of the evaluation duration for which the miss rate is below this threshold. In an embodiment this evaluation duration is greater than the EBET value.
Next control passes to diamond 375 to determine whether the performance metric information (e.g., miss rate) remains below this threshold. If not, the timer can be reset and monitoring may begin again at block 360.
Otherwise if at diamond 375 it is determined that the miss rate remains below the threshold level, control passes to diamond 390 where it can be determined whether the evaluation duration is greater than the EBET value. If not, control proceeds back to block 370 for further counting of an evaluation duration. If instead it is determined that the evaluation duration is greater than the EBET value, control passes to block 350, discussed above, where the LLC may be flushed and power gated. Although shown at this high level in the embodiment of
As more computer systems are used in Internet, text, and multimedia applications, additional processor support has been introduced over time. In one embodiment, an instruction set architecture (ISA) may be associated with one or more computer architectures, including data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I/O).
In one embodiment, the ISA may be implemented by one or more micro-architectures, which include processor logic and circuits used to implement one or more instruction sets. Accordingly, processors with different micro-architectures can share at least a portion of a common instruction set. For example, Intel® Pentium 4 processors, Intel® Core™ and Intel® Atom™ processors from Intel Corp. of Santa Clara, Calif., and processors from Advanced Micro Devices, Inc. of Sunnyvale Calif. implement nearly identical versions of the x86 instruction set (with some extensions that have been added with newer versions), but have different internal designs. Similarly, processors designed by other processor development companies, such as ARM Holdings, Ltd., MIPS, or their licensees or adopters, may share at least a portion a common instruction set, but may include different processor designs. For example, the same register architecture of the ISA may be implemented in different ways in different micro-architectures using new or well-known techniques, including dedicated physical registers, one or more dynamically allocated physical registers using a register renaming mechanism (e.g., the use of a register alias table (RAT), a reorder buffer (ROB) and a retirement register file). In one embodiment, registers may include one or more registers, register architectures, register files, or other register sets that may or may not be addressable by a software programmer.
Referring now to
A processor including core 600 may be a general-purpose processor, such as a Core™ i3, i5, i7, 2 Duo and Quad, Xeon™, Itanium™, XScale™ or StrongARM™ processor, which are available from Intel Corporation. Alternatively, the processor may be from another company, such as a design from ARM Holdings, Ltd, MIPS, etc. The processor may be a special-purpose processor, such as, for example, a network or communication processor, compression engine, graphics processor, co-processor, embedded processor, or the like. The processor may be implemented on one or more chips, and may be a part of and/or may be implemented on one or more substrates using any of a number of process technologies, such as, for example, BiCMOS, CMOS, or NMOS.
As shown in
Coupled between front end units 610 and execution units 620 is an out-of-order (OOO) engine 615 that may be used to receive the micro-instructions and prepare them for execution. More specifically OOO engine 615 may include various buffers to re-order micro-instruction flow and allocate various resources needed for execution, as well as to provide renaming of logical registers onto storage locations within various register files such as register file 630 and extended register file 635. Register file 630 may include separate register files for integer and floating point operations. Extended register file 635 may provide storage for vector-sized units, e.g., 256 or 512 bits per register.
Various resources may be present in execution units 620, including, for example, various integer, floating point, and single instruction multiple data (SIMD) logic units, among other specialized hardware. For example, such execution units may include one or more arithmetic logic units (ALUs) 622, among other such execution units.
Results from the execution units may be provided to a retirement unit 640 including a reorder buffer (ROB). This ROB may include various arrays and logic to receive information associated with instructions that are executed. This information is then examined by retirement unit 640 to determine whether the instructions can be validly retired and result data committed to the architectural state of the processor, or whether one or more exceptions occurred that prevent a proper retirement of the instructions. Of course, retirement unit 640 may handle other operations associated with retirement.
As shown in
While shown with this high level in the embodiment of
Referring now to
In general, each tile 710 may include a private cache hierarchy including two or more cache levels in addition to various execution units and additional processing elements. In turn, the various tiles may be coupled to each other a ring interconnect 730, which further provides interconnection between the cores, graphics domain 720 and system agent circuitry 750. In one embodiment, interconnect 730 can be part of the core domain. However in other embodiments the ring interconnect can be of its own domain.
As further seen, system agent domain 750 may include display controller 752 which may provide control of and an interface to an associated display. As further seen, system agent domain 750 may include a power control unit 755. As shown in
As further seen in
Embodiments may be implemented in many different system types. Referring now to
Still referring to
Furthermore, chipset 890 includes an interface 892 to couple chipset 890 with a high performance graphics engine 838, by a P-P interconnect 839. In turn, chipset 890 may be coupled to a first bus 816 via an interface 896. As shown in
Using an adaptive local LLC and associated power management as described herein, significant power savings may be realized, especially when per core utilization is low (e.g. less than approximately 30%), which is relatively common for certain workloads such as server-based workloads. Using an embodiment of the present invention having private cache organization and dynamic LLC power management, power consumption can be reduced by ˜50% with 10% loadline, and ˜40% with 20% loadline. In a multicore processor having a private or clustered cache organization, maximum processor cache power saving opportunities can be realized while maintaining performance requirements by flushing and powering off local LLC portions when appropriate, based on run-time conditions.
Embodiments may be used in many different types of systems. For example, in one embodiment a communication device can be arranged to perform the various methods and techniques described herein. Of course, the scope of the present invention is not limited to a communication device, and instead other embodiments can be directed to other types of apparatus for processing instructions, or one or more machine readable media including instructions that in response to being executed on a computing device, cause the device to carry out one or more of the methods and techniques described herein.
Embodiments may be implemented in code and may be stored on a non-transitory storage medium having stored thereon instructions which can be used to program a system to perform the instructions. The storage medium may include, but is not limited to, any type of disk including floppy disks, optical disks, solid state drives (SSDs), compact disk read-only memories (CD-ROMs), compact disk rewritables (CD-RWs), and magneto-optical disks, semiconductor devices such as read-only memories (ROMs), random access memories (RAMs) such as dynamic random access memories (DRAMs), static random access memories (SRAMs), erasable programmable read-only memories (EPROMs), flash memories, electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, or any other type of media suitable for storing electronic instructions.
While the present invention has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom. It is intended that the appended claims cover all such modifications and variations as fall within the true spirit and scope of this present invention.
This application is a continuation of U.S. patent application Ser. No. 13/715,613, filed Dec. 14, 2012, the content of which is hereby incorporated by reference.
Number | Name | Date | Kind |
---|---|---|---|
5163153 | Cole et al. | Nov 1992 | A |
5522087 | Hsiang | May 1996 | A |
5590341 | Matter | Dec 1996 | A |
5621250 | Kim | Apr 1997 | A |
5931950 | Hsu | Aug 1999 | A |
6748546 | Mirov et al. | Jun 2004 | B1 |
6792392 | Knight | Sep 2004 | B1 |
6823516 | Cooper | Nov 2004 | B1 |
6829713 | Cooper et al. | Dec 2004 | B2 |
6996728 | Singh | Feb 2006 | B2 |
7010708 | Ma | Mar 2006 | B2 |
7043649 | Terrell | May 2006 | B2 |
7093147 | Farkas et al. | Aug 2006 | B2 |
7111179 | Girson et al. | Sep 2006 | B1 |
7194643 | Gonzalez et al. | Mar 2007 | B2 |
7272730 | Acquaviva et al. | Sep 2007 | B1 |
7412615 | Yokota et al. | Aug 2008 | B2 |
7434073 | Magklis | Oct 2008 | B2 |
7437270 | Song et al. | Oct 2008 | B2 |
7454632 | Kardach et al. | Nov 2008 | B2 |
7529956 | Stufflebeam | May 2009 | B2 |
7539885 | Ma | May 2009 | B2 |
7596662 | Makineni et al. | Sep 2009 | B2 |
7730340 | Hu et al. | Jun 2010 | B2 |
8738860 | Griffin et al. | May 2014 | B1 |
8799624 | Griffin et al. | Aug 2014 | B1 |
20010044909 | Oh et al. | Nov 2001 | A1 |
20020147957 | Matsui et al. | Oct 2002 | A1 |
20020194509 | Plante et al. | Dec 2002 | A1 |
20030061383 | Zilka | Mar 2003 | A1 |
20040064752 | Kazachinsky et al. | Apr 2004 | A1 |
20040098560 | Storvik et al. | May 2004 | A1 |
20040139356 | Ma | Jul 2004 | A1 |
20040268166 | Farkas et al. | Dec 2004 | A1 |
20050022038 | Kaushik et al. | Jan 2005 | A1 |
20050033881 | Yao | Feb 2005 | A1 |
20050114850 | Chheda et al. | May 2005 | A1 |
20050132238 | Nanja | Jun 2005 | A1 |
20060050670 | Hillyard et al. | Mar 2006 | A1 |
20060053326 | Naveh | Mar 2006 | A1 |
20060059286 | Bertone et al. | Mar 2006 | A1 |
20060069936 | Lint et al. | Mar 2006 | A1 |
20060117202 | Magklis et al. | Jun 2006 | A1 |
20060184287 | Belady et al. | Aug 2006 | A1 |
20070005995 | Kardach et al. | Jan 2007 | A1 |
20070016817 | Albonesi et al. | Jan 2007 | A1 |
20070079294 | Knight | Apr 2007 | A1 |
20070106827 | Boatright et al. | May 2007 | A1 |
20070150658 | Moses et al. | Jun 2007 | A1 |
20070156992 | Jahagirdar | Jul 2007 | A1 |
20070214342 | Newburn | Sep 2007 | A1 |
20070239398 | Song et al. | Oct 2007 | A1 |
20070245163 | Lu et al. | Oct 2007 | A1 |
20080028240 | Arai et al. | Jan 2008 | A1 |
20080059707 | Makineni et al. | Mar 2008 | A1 |
20080250260 | Tomita | Oct 2008 | A1 |
20080307240 | Dahan et al. | Dec 2008 | A1 |
20090006871 | Liu et al. | Jan 2009 | A1 |
20090150695 | Song et al. | Jun 2009 | A1 |
20090150696 | Song et al. | Jun 2009 | A1 |
20090158061 | Schmitz et al. | Jun 2009 | A1 |
20090158067 | Bodas et al. | Jun 2009 | A1 |
20090172375 | Rotem et al. | Jul 2009 | A1 |
20090172428 | Lee | Jul 2009 | A1 |
20090235105 | Branover et al. | Sep 2009 | A1 |
20090327609 | Fleming et al. | Dec 2009 | A1 |
20100115309 | Carvalho et al. | May 2010 | A1 |
20100146513 | Song | Jun 2010 | A1 |
20100191997 | Dodeja et al. | Jul 2010 | A1 |
20100250856 | Owen et al. | Sep 2010 | A1 |
20110154090 | Dixon et al. | Jun 2011 | A1 |
20110161627 | Song et al. | Jun 2011 | A1 |
20110219190 | Ng et al. | Sep 2011 | A1 |
20120079290 | Kumar | Mar 2012 | A1 |
20120131370 | Wang et al. | May 2012 | A1 |
20120166731 | Maciocco et al. | Jun 2012 | A1 |
20120210032 | Wang et al. | Aug 2012 | A1 |
20120233377 | Nomura et al. | Sep 2012 | A1 |
20120246506 | Knight | Sep 2012 | A1 |
20120254643 | Fetzer et al. | Oct 2012 | A1 |
20130080813 | Tarui et al. | Mar 2013 | A1 |
20130111121 | Ananthakrishnan et al. | May 2013 | A1 |
20140040676 | Solihin | Feb 2014 | A1 |
20140181410 | Kalamatianos et al. | Jun 2014 | A1 |
20140223104 | Solihin | Aug 2014 | A1 |
20140237185 | Solihin | Aug 2014 | A1 |
Number | Date | Country |
---|---|---|
1 282 030 | May 2003 | EP |
2336892 | Jun 2011 | EP |
2012136766 | Oct 2012 | WO |
Entry |
---|
Intel Developer Forum, IDF2010, Opher Kahn, et al., “Intel Next Generation Microarchitecture Codename Sandy Bridge: New Processor Innovations,” Sep. 13, 2010, 58 pages. |
SPEC-Power and Performance, Design Overview V1.10, Standard Performance Information Corp., Oct. 21, 2008, 6 pages. |
Intel Technology Journal, “Power and Thermal Management in the Intel Core Duo Processor,” May 15, 2006, pp. 109-122. |
Anoop Iyer, et al., “Power and Performance Evaluation of Globally Asynchronous Locally Synchronous Processors,” 2002, pp. 1-11. |
Greg Semeraro, et al., “Hiding Synchronization Delays in a GALS Processor Microarchitecture,” 2004, pp. 1-13. |
Joan-Manuel Parcerisa, et al., “Efficient Interconnects for Clustered Microarchitectures,” 2002, pp. 1-10. |
Grigorios Magklis, et al., “Profile-Based Dynamic Voltage and Frequency Scalling for a Multiple Clock Domain Microprocessor,” 2003, pp. 1-12. |
Greg Semeraro, et al., “Dynamic Frequency and Voltage Control for a Multiple Clock Domain Architecture,” 2002, pp. 1-12. |
Greg Semeraro, “Energy-Efficient Processor Design Using Multiple Clock Domains with Dynamic Voltage and Frequency Scaling,” 2002, pp. 29-40. |
Diana Marculescu, “Application Adaptive Energy Efficient Clustered Architectures,” 2004, pp. 344-349. |
L. Benini, et al., “System-Level Dynamic Power Management,” 1999, pp. 23-31. |
Ravindra Jejurikar, et al., “Leakage Aware Dynamic Voltage Scaling for Real-Time Embedded Systems,” 2004, pp. 275-280. |
U.S. Appl. No. 13/285,465, filed Oct. 31, 2011, entitled, “Dynamically Controlling Cache Size to Maximize Energy Efficiency,” by Avinash N. Ananthakrishnan, et al. |
U.S. Appl. No. 13/225,677, filed Sep. 6, 2011, entitled, “Dynamically Allocating a Power Budget Over Multiple Domains of a Processor,” by Avinash N. Ananthakrishnan, et al. |
U.S. Appl. No. 13/600,568, filed Aug. 31, 2012, entitled, “Configuring Power Management Functionality in a Processor,” by Malini K. Bhandaru, et al. |
Intel Corporation, “Intel Core 2 Due Mobile Processor, Intel Core 2 Solo Mobile Processor and Intel Core 2 Extreme Mobile Processor on 45-nm Process,” Chaper 2, Low Power Features, May 2009, 15 pages. |
International Searching Authority, “Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority,” mailed Oct. 24, 2013, in International application No. PCT/US2013/048486. |
Number | Date | Country | |
---|---|---|---|
20140173206 A1 | Jun 2014 | US |
Number | Date | Country | |
---|---|---|---|
Parent | 13715613 | Dec 2012 | US |
Child | 13785228 | US |