Claims
- 1. A method for managing data, the method comprising the steps of:
maintaining a plurality of persistent data items on persistent storage accessible to a plurality of nodes, the persistent data items including a particular data item stored at a particular location on said persistent storage; assigning exclusive ownership of each of the persistent data items to one of the nodes, wherein a particular node of said plurality of nodes is assigned exclusive ownership of said particular data item; when any node wants an operation to be performed that involves said particular data item, the node that desires the operation to be performed ships the operation to the particular node for the particular node to perform the operation on the particular data item as said particular data item is exclusively owned by said particular node; in response to a failure that involves a set of persistent data items exclusively owned by a single node, performing the steps of:
assigning, to each of two or more recovery nodes, exclusive ownership of a subset of the set of persistent data items that were involved in the failure; and each recovery node of the two or more recovery nodes performing a recovery operation on the subset of persistent data items that were assigned to the recovery node.
- 2. The method of claim 1 wherein the failure is a media failure of a persistent storage device that stores said set of persistent data items.
- 3. The of claim 1 wherein:
the failure is a failure of the node that has exclusive ownership of said set of persistent data items; and the step of assigning includes assigning, to each of two or more recovery nodes, exclusive ownership of a subset of the persistent data items that were exclusively owned by the failed node.
- 4. The method of claim 3, wherein:
the two or more recovery nodes include a first recovery node and a second recovery node; and a least a portion of the recovery operation performed by the first recovery node on the subset of data exclusively assigned to the first recovery node is performed in parallel with at least a portion of the recovery operation performed by the second recovery node on the subset of data exclusively assigned to the second recovery node.
- 5. The method of claim 3 further comprising:
organizing the plurality of persistent data items into a plurality of buckets; and establishing a mapping between the plurality of buckets and the plurality of nodes, wherein each node has exclusive ownership of the data items that belong to all buckets that map to the node; and determining which data items need to be recovered based on said mapping.
- 6. The method of claim 5 further comprising:
performing a first pass on said mapping to determine which buckets have data items that need to be recovered; performing a second pass on said mapping to perform recovery on the data items that need to be recovered; and after performing the first pass and before completing the second pass, making available for access the data items that belong to all buckets that do not have to be recovered.
- 7. The method of claim 3 wherein each recovery node of the two or more recovery nodes performs the recovery operation based on recovery logs, associated with the failed node, on the persistent storage.
- 8. The method of claim 7 further comprising the step of a recovery coordinator scanning the recovery logs associated with the failed node and distributing recovery records to the two or more recovery nodes.
- 9. The method of claim 7 wherein each of the two or more recovery nodes scans the recovery logs associated with the failed node.
- 10. The method of claim 3 wherein:
the step of each recovery node of the two or more recovery nodes performing a recovery operation includes applying undo records to blocks; and the method further comprises the step of tracking which undo records have been applied.
- 11. The method of claim 5 further comprising the step of, prior to the failure, the failed node storing, within redo records that are generated by the failed node, bucket numbers that indicate to which buckets the data items associated with the redo records belong.
- 12. The method of claim 3 wherein recovery of the failed node involves various tasks, the method further comprising the steps of:
a recovery coordinator determining that a first set of one or more tasks required for recovery of said failed node should be performed serially, and that a second set of one or more tasks required for recovery of said failed node should be performed in parallel; and performing the first set of one or more tasks serially; and using said two or more recovery nodes to perform said second set of one or more tasks in parallel.
- 13. The method of claim 12 wherein the step of determining that a second set of one or more tasks required for recovery of said failed node should be performed in parallel is performed based, at least in part, on the size of one or more objects that need to be recovered.
- 14. The method of claim 12 wherein:
ownership of data items involved in said second set of one or more tasks is passed from the recovery coordinator to the two or more recovery nodes to allow said two or more recovery nodes to perform said second set of one or more tasks; and after performance of said second set of one or more tasks and before completion of the recovery of said failed node, ownership of data items involved in said second set of one or more tasks is passed back to said recovery coordinator from said two or more recovery nodes.
- 15. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 3.
- 16. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 4.
- 17. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 5.
- 18. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 6.
- 19. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 7.
- 20. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 8.
- 21. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 9.
- 22. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 10.
- 23. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 11.
- 24. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 12.
- 25. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 13.
- 26. A computer-readable medium carrying one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform the method recited in claim 14.
Parent Case Info
[0001] This application claims the benefit of priority from U.S. Provisional Application Ser. No. 60/492,019 entitled “Shared Nothing on Shared Disk Hardware”, filed Aug. 1, 2003, which is incorporated by reference in its entirety for all purposes as if fully set forth herein.
[0002] This application also claims benefit as a Continuation-in-part of application Ser. No. 10/665,062 filed Sep. 17, 2003 and application Ser. No. 10/718,875, filed Nov. 21, 2003, the entire contents of which are hereby incorporated by reference as if fully set forth herein.
[0003] This application is related to U.S. application Ser. No. ______, (Attorney Docket No. 50277-2323) entitled “Dynamic Reassignment of Data Ownership,” by Roger Bamford, Sashikanth Chandrasekaran and Angelo Pruscino, filed on the same day herewith, and U.S. application Ser. No. ______, (Attorney Docket No. 50277-2326) entitled “Partitioned Shared Cache,” by Roger Bamford, Sashikanth Chandrasekaran and Angelo Pruscino, filed on the same day herewith; both of which are incorporated by reference in their entirety for all purposes as if fully set forth herein.
Provisional Applications (1)
|
Number |
Date |
Country |
|
60492019 |
Aug 2003 |
US |
Continuation in Parts (2)
|
Number |
Date |
Country |
| Parent |
10718875 |
Nov 2003 |
US |
| Child |
10831413 |
Apr 2004 |
US |
| Parent |
10665062 |
Sep 2003 |
US |
| Child |
10831413 |
Apr 2004 |
US |