The present disclosure relates to processing of documents, and in particular, to aggregating documents based on relationships of attributes associated with the documents and/or correcting determined inconsistencies between aggregated documents.
In nearly any relatively large organization, whether it be a corporate organization, governmental organization, educational organization, etc., document management is important but very challenging for a myriad of reasons. To begin, in many organizations the sheer number of electronic documents is challenging. In many situations, organizations employ document management systems and related databases that may provide tools to organize documents. Various attributes of a document may be identified at the creation of the document. For example, a user may name the document, and store the document in a file structure that implicitly relates the document with other documents, which may be based on any number of relational and/or hierarchical characteristics including the type of document, a project, the creator of the document, etc. However, at creation, it is quite possible that none or few of these attributes may be associated with a document. Documents may also be categorized during a procurement phase that occurs after the initial document is created. Overall, whether at creation or during a later procurement, organizations often expend great resources reviewing and/or categorizing documents so that that those documents can be discovered in a search or otherwise identified at a later time based on information associated with each document.
In the majority of situations, however, document organization is a manual process. For example, many organizations manually associate, whether at creation, when uploaded into a system, or at some point later, attributes or metadata with each document that describe particular aspects of the stored electronic document. These manually applied attributes serve to aid end users in grouping and organizing information and identifying related documents. However, this process of manual attribution is often incomplete for a variety of reasons including a user having an incomplete understanding of the document necessary for proper definition, attribution tools being insufficient for proper and complete attribution, simple lack of prioritization, human error, and any number of other issues. In even a high functioning environment, a user may simply have insufficient knowledge about a document, or the information may simply not yet be knowable.
In addition and depending on the project, accessing the database to identify the correct documents for a particular search, let alone properly analyzing each document, can be burdensome due to the errors common in manual attribution of the documents. For example, in a complicated transaction, there may be many documents related to the transaction, and additional documents created over time. It would not be uncommon for a document or documents related to the transaction to be mis-labeled, stored incorrectly, simply not labeled, be correctly but insufficiently labeled, etc. Hence, when a user attempts to search for documents related to the transaction, not all documents are retrieved due to any one or more of the above issues or other issues. Further complicating document organization, each document may be organized uniquely and use different attribute terms, even when they pertain to the same topic, adding to the difficulty in properly aggregating the documents.
It is with these observations in mind, among others, that aspects of the present disclosure were concerned and developed.
Embodiments of the disclosure concern document management systems and methods. A first implementation includes a method for aggregating related documents comprising the operations of accessing, by a processor and based on receiving an attribute associated with aggregating related documents, a plurality of electronic documents and receiving, after receiving the attribute and by a trained machine learning model, a plurality of values each corresponding to one or more categories related to the content of a text of the plurality of electronic documents. The method may further include associating, by the processor, the received attribute with a subset of the plurality of electronic documents, wherein each of the subset of the plurality of electronic documents is associated with at least one of the plurality of values corresponding to the received attribute and generating, by the processor, a graphical user interface. The graphical user interface may include a first portion displaying a portion of each of the subset of the plurality of electronic documents and a second portion displaying one or more of the plurality of values common to the subset of the plurality of electronic documents.
Other implementations may include a system for aggregating related documents comprising a processor and a memory comprising instructions that, when executed, cause the processor to perform operations. Such operations may include accessing, based on receiving an attribute associated with aggregating related documents, a plurality of electronic documents, receiving, by a trained machine learning model and after receiving the attribute, a plurality of values each corresponding to one or more categories related to a content of a text of the plurality of electronic documents, and associating the received attribute with a subset of the plurality of electronic documents, wherein each of the subset of the plurality of electronic documents is associated with at least one of the plurality of values corresponding to the received attribute. The operations may also include generating a graphical user interface including a first portion displaying a portion of each of the subset of the plurality of electronic documents and a second portion displaying one or more of the plurality of values common to the subset of the plurality of electronic documents.
Yet another implementation may include one or more non-transitory computer-readable storage media storing computer-executable instructions for performing a computer process on a computing system. The computer process may include the operations of accessing, by the computing system and based on receiving an attribute associated with aggregating related documents, a plurality of electronic documents and receiving, after receiving the attribute and by a trained machine learning model, a plurality of values each corresponding to one or more categories related to the content of a text of the plurality of electronic documents. The method may further include associating, by the computing system, the received attribute with a subset of the plurality of electronic documents, wherein each of the subset of the plurality of electronic documents is associated with at least one of the plurality of values corresponding to the received attribute and generating, by the computing system, a graphical user interface. The graphical user interface may include a first portion displaying a portion of each of the subset of the plurality of electronic documents and a second portion displaying one or more of the plurality of values common to the subset of the plurality of electronic documents.
The foregoing and other objects, features, and advantages of the present disclosure set forth herein should be apparent from the following description of particular embodiments of those inventive concepts, as illustrated in the accompanying drawings. The drawings depict only typical embodiments of the present disclosure and, therefore, are not to be considered limiting in scope.
Aspects of the present disclosure involve systems and methods for automated analysis of documents to obtain attributes associated with those documents, and using the attributes to organize, relate, and/or aggregate documents. Attributes, or generally features of a document generated by the system, can be applied to the document or inferred from the document based on a machine learning model. One or more of either of these types of attributes can be used to relate documents together, join them together, or aggregate them with their associated metadata into a composite result. The aggregation of attributes may include rules for how attributes are to be aggregated.
In one implementation, a document management system may receive a collection of documents and, in some instances, scan the documents to create a corresponding image for use in aggregating the documents. Attributes, also referred to herein as “key values” or “key attributes”, for generating an aggregation of documents may be received at the document management system. In one instance, the key attributes may be received via a graphical user interface (also referred to herein as a “user interface”). An artificial intelligence or machine learning technique may then be applied to the collection of documents to extract or otherwise determine attributes or data from the documents. In an example in which the documents include a contract, such determined attributes may include names of parties to the contract, agreement numbers, expiration dates, initiation dates, particular provisions of the documents, auto-renewal indicators and the like. The artificial intelligence or machine learning techniques may also interpret portions of the electronic documents to infer one or more attributes of the documents. The received key attributes may be compared to these extracted or inferred attributes to determine if a match between the key attributes and the extracted or inferred attributes is present. Documents that have extracted or inferred attributes that match the provided key attributes may be included in an aggregation of related documents. Additional documents may also be included in the aggregation based on the extracted or inferred attributes. For example, an analysis of the attributes associated with the documents in the aggregation may generate a document profile for documents within the aggregation. Other documents managed by the document management system that do not include the provided key attributes but nonetheless include attributes common to the other documents in the aggregation may also be selected for inclusion in the library of related documents.
Using artificial intelligence data extractions and interpretations of document content to infer the presence of attributes in the documents brings at least two benefits: 1) allowing for the data attribution process to be automated (and only augmented by human intervention) thereby creating a higher degree of accuracy and completeness in the attribution process and 2) allowing for documents to be aggregated across many different types of attributes beyond those known a priori. Additionally, the aggregated documents may be displayed via a user interface in a manner as to indicate a current active “state” of a particular relationship, contractual obligation, or otherwise of the related documents. Displaying document attribute instances of key data across all of the participating documents in an aggregate may provide a clear understanding of what key data is to be considered in making a decision or interpretation on the aggregate of related documents.
In still another instance, the document management platform may provide for identification and/or correction of conflicts within an aggregate of related documents. For example, a particular provision of an aggregation of contract documents may be extracted from the documents and compared. Instances in which the compared portions from the aggregated documents do not match, a potential conflict between the documents may be displayed on a user interface. Conflict between any aspect of the aggregated documents may trigger a conflict alert on the user interface. Further, the conflict between the documents may be resolved by the document management platform by altering a document or altering an extracted portion of the document to match a controlling version of the portions in conflict. The controlling version of the portion in conflict may be selected by a user of the interface in one example. In another, the document management platform may select a controlling version of the portion based on other document information, such as execution date or document type, and correct the documents in the aggregated collection based on the selected controlling version.
Generally, the system may receive a document as an image file (e.g., PDF, JPG, PNG, etc.), and the system extracts text from the image file. In some embodiments, the system may receive one or more images of, for example, oil and gas documents. In some cases, the received image document may have been pre-processed to extract the text and thus includes the text level information. Text extraction can be done by various tools available on the market today falling within the broad stable of Optical Character Recognition (“OCR”) software. The extracted text may be associated or otherwise linked with the particular location in the document from which it was found or extracted.
Extracted text may then be fed into a trained machine learning model. The trained machine learning model may be trained on sample data and/or previously received documents so that it can identify categories and subcategories, which it associates with particular sections of text. Thus, even if a document does not include particular section titles, spacing key-words or other identifiers, extracted text may still be associated with an appropriate category. Having identified categories, which may further include subcategories, associated with particular sections of the text, the particular locations associated with the particular sections of text can then also be associated with the identified categories and subcategories as well. These categories and subcategories may then be used to build an aggregated collection of related documents based on one or more key attributes of the stored documents.
In the example illustrated, an image 122 of the electronic document is stored in a system database or other memory provided by a document management platform 106. It should be recognized that the document, when first loaded to or accessed by the system, may be in the form of image. It is often the case, for example, that final documents of some form of transaction are in image form, e.g., a PDF file or the like. The system, however, may work natively with other electronic forms of documents, such as those generated from word processing programs. The database can be a relational or non-relational database, and it will be apparent to a person having ordinary skill in the art which type of database to use or whether to use a mix of the two. In some other embodiments, the document may be stored in a short-term memory rather than a database or be otherwise stored in some other form of memory structure. Documents stored in the system database may be used later for training new machine learning models and/or continued training of existing machine learning models through utilities provided by the document management services platform 106. The document management services platform 106 can be a cloud platform, locally hosted, locally hosted in a distributed enterprise environment, distributed, combinations of the same, and otherwise available in different forms.
The document management services platform 106 may provide systems and methods for relating and aggregating the stored electronic documents 102. In particular, the document management services platform 106 may aggregate one of more of the documents 102 based on one or more key attributes identified through a user interface 113 executed on a user device 114.
Beginning in operation 202, the document management platform 106 may receive one or more key attributes for building an aggregate of related documents. In one instance, the attributes may be received via the user interface 113. One particular example of a user interface for defining key attributes for aggregating related documents is illustrated in
Returning to
A general description of the techniques for extracting attributes from the documents or otherwise inferring the document attributes is illustrated in the method 400 of
Machine learning models may be applied to the text to identify categories and subcategories for the text in operation 404. In one specific example system implementation, machine learning services utilize the storage and machine learning support 108 to retrieve trained models 121 from a remote device 110, which may include a database or other data storage facility such as a cloud storage service. The machine learning models 121 may identify categories and subcategories based on learned ontologies which are taught to the models through training on batches of text from previous documents received by the system and from training data, which may be acquired during the initial deployment of the system or otherwise. A learned ontology can allow a machine learning model 121 to identify a category or subcategory based on relationships between words, key words, and other factors determined by the machine learning algorithm employed, and will identify concepts and information embedded in the syntax and semantics of text. For example, in some oil and gas lease agreements, there is a specific legal concept referred to as a lot description. Thus, where a simple key word search of extracted text may not be capable alone of identifying a key word of “lot description” unless the exact term is present, machine learning can be used to analyze the document, including but not limited to the extracted text, and identify the “lot description” based on other criteria besides the use of the exact term such as using a previously identified location of the lot (e.g., via the lot state, lessor state, applicable laws state, etc.) to identify probable formats for the lot description and/or other qualities of the text (e.g., proximate categories, such as lessor name or related categories, such as state, and the like). In another example, a legal concept typical in various oil and gas industry documents is a “shut-in” provision. However, it is often the case that there is not a specific provision heading explicitly titled “shut-in” and in many instances the specific term “shut-in” is not used in what would otherwise be considered a shut-in provision as the document section is describing a shut-in provision but without explicitly using the term “shut-in.” Thus, the machine learning models may process the extracted words, along with other document attributes, to identify if a portion of the extracted text is a “shut-in” provision based on the use words typical of shut-in (e.g., “gas not being sold”), the use of sets of similar words being used in proximate locations (e.g., “gas not being sold,” “capable of producing,” and “will pay”) to identify a category. Such named entity resolution techniques may be applied to any identified text in a document. The machine learning algorithm employed may be a neural network, a deep neural network, support vector machines, a Bayesian network or networks, a combination of multiple algorithms, or any other implementation that will be apparent to a person having ordinary skill in the art.
In operation 406, the extracted text and automated category and/or subcategory identifications may be associated with locations in the respective document pertaining to the text of such categorizations. The category, subcategory, and location information may be stored in the document management platform 106 or remote devices 110 for use in matching to one or more key attributes for aggregated related documents. In some instances, the extracted text may be a hash value of a portion of a document. For example, a particular clause of a contract or other section of words or paragraph of a document may be transformed into a hash value for comparison to a key attribute. In this example, the document management platform 106 may similarly determine a hash value for a key attribute, such as a provision of a document, for comparison to the hash value of the document attribute. In one particular implementation, the method 400 of
The categories and/or subcategories of the text of the stored documents may be used to aggregate related documents. For example, a portion of a document may be identified as a “Company Name” category or subcategory and an extracted document attribute may be associated with the document as a “Company Name” value. Similarly, an agreement number, such as 123456, may be identified in the document as an agreement number and associated with an “Agreement No.” category or subcategory. To aggregate documents, the document management platform 106 may utilize the categorization of the extracted or inferred attributes associated with the documents to identify documents to include in the aggregation. For example, a “Company Name” key attribute may be identified for aggregation and, in response, the document management platform 106 may identify portions of stored documents categorized as a company name for comparison to a received company name value and determine if the document is to be included in the aggregation of documents. In this manner, categories or subcategories of portions of the stored electronic documents may be identified through the machine learning or artificial intelligence techniques and such categories may be used to aggregate documents as corresponding to a received key attribute.
Returning to
Aggregation of documents may be based on any data obtainable or inferred from the document. For example, related documents associated with a vendor name may be aggregated based on a key attribute identifying that vendor name. Also, as described above, aggregation may occur on entire clauses of the documents. For example, a user may provide a termination clause of a contract as a key attribute through the user interface. In one example, the document management platform 106 may generate a repeatable hash value based on the provided termination clause. Other techniques for converting a clause into a searchable form may also be utilized to compare a provided clause to clauses of other documents, including but not limited to frequency-inverse document frequency (tf-idf) technique or a trained machine learning based embedding model technique. The platform 106 may also extract similar termination clauses from one or more documents stored with the platform and generate a hash value or other searchable value for the extracted clauses or paragraphs using the same techniques. A comparison of the provided clause hash value to the extracted clause hash values may determine if other documents in the stored documents 102 include the same or a similar termination clause. In this manner, an aggregation of documents may be generated based on an entire clause of a document.
In some instances, documents with extracted or inferred document attributes that do not match the provided key attribute may also be included in the aggregation of related documents in operation 208. For example, the document management platform 106 may determine that a document does not include the provided key attribute as a document attribute. However, several other document attributes may match document attributes for other documents included in the aggregation, such as vendor name, vendor address, site address, date of execution, date of expiration, etc. A correlation of document attributes for the aggregated collection of documents may yield some document attributes that are common to some or all of the documents. The document management platform 106 may then use these common attributes to identify other documents that may be related to the key attribute while not specifically including the key attribute itself. This technique may also be used to gather documents with errors into an aggregation. For example, a document may include an incorrect agreement number but should otherwise be included in an aggregation of related documents as belonging to a particular agreement. Through an analysis of the documents already included in the library, the document management platform 106 may identify a vendor name and expiration date that are common to all or most of the aggregated documents. Other documents stored with the platform 106 may also include the same vendor name and expiration date, but not include the agreement number key attribute. These other documents may also be included in the aggregation as possible related documents by the document management platform.
In operation 210, the aggregated documents based on the key attributes may be presented or otherwise displayed by the user device 114 on the user interface 113.
Each of the documents illustrated in the user interface 500 may be selectable through an input device to the user interface for display of the document. For example, selection of a thumbnail of a page of Document A 506 may expand the thumbnail to show the full page within the user interface. An expanded page of a document of the aggregated collection of related documents is illustrated in the user interface 600 of
The second portion 604 of the user interface 600 may include an image 610 of the selected document or document page. For example, the image 610 may be a scan of a received document or a page or other portion of an electronic document. The second portion 604 may include a title or name 614 of the displayed document for reference by a user of the interface, along with other information, such as a type of document, the number of pages in the document, and the like.
A third portion 606 of the user interface 600 may display document attributes or other data or attributes extracted or inferred from the selected document. For example, the third portion 606 may include a company name, such as Company A, extracted from an analysis of the document. The document may be included in the aggregation of documents because this document attribute matches the key attribute 612 used to build the library. Other document attributes or information associated with the document are also noted in the third portion 606, such as a state of the agreement, an agreement type, an activation date, a deal type, etc. The information included in the third portion 606 may be obtained from the machine learning and artificial intelligence analysis of the documents to build ontologies of the document information to determine the various information contained in the document.
In a similar manner and returning to the user interface 500 of
Through the systems and methods described above, extracted or inferred attributes of electronic documents may be utilized to relate or aggregate documents. A current or active state of the aggregation of documents may be automatically provided without a manual attribution of data to the documents. Rather, artificial intelligence or machine learning techniques may be executed to extract and/or interpret the content of the documents to infer the document attributes. This process allows for a higher degree of accuracy and completeness of the aggregation process across many different types of documents. Such a system may also be utilized to identify potential conflicts of text or information within aggregated documents and provide a mechanism through which such conflicts may be corrected. One particular implementation for identifying and correcting potential conflicts within the documents of the aggregated library is described below.
Referring now to
The ordering system 710 may output a chronologically ordered set of documents 705. The ordered documents 705 may be organized differently than they are first received. For example, the original contract 704B may be sorted to the front of the received documents (thus denoting an earlier date), even though it was received after addendum 740A. As can be seen, the received documents are organized such that original contract 704B precedes addendum 704C, which precedes addendum 704A, which precedes addendum 704D. The document processing system 106, discussed above, can then receive the ordered documents in their correct sequence. However, where in some embodiments document processing system 106 may perform the document ordering as described herein.
In addition to the document ordering system 710, the document management system 106 or another system may detect conflicts and, in some instances, rectify detected conflicts in the aggregated collection of related documents. A user interface for displaying aggregated documents associated with a key attribute and an identified conflict between at least two documents is illustrated in
Returning to
In another instance, the document management platform 106 or other system may update the aggregated documents automatically. For example, the conflict check system 750 may provide the documents to a document markup system 752 for correction. The document markup system 752 may update one or more of the documents based on an ordering of the documents as determined by the document ordering system 710. In general, the document markup system 752 may determine the provision or text that is most recent in the ordered documents 705 and update the remaining documents with the most recent text or provision in operation 906. In another example, the document markup system 752 may determine the most often used provision or text for a document feature and select that version of the conflicting text for updating the documents in the aggregation. Regardless of the technique used by the system to select a controlling version of the conflicting information, one or more of the documents of the aggregated collection of documents may be corrected with the controlling version. The corrected documents may be displayed in the user interface in a manner similar to above in operation 908.
The document management system 106 described herein may include many such automated techniques and features. For example, the document management system 106 may incorporate a rule base that requires a certain number of documents to be within an aggregate of documents to limit the number of documents or for expanding the number of documents which may apply to the key attribute defining the aggregate. A similar rule for a certain type of documents may also be applied by the document management platform 106. In another example, the document management platform may include a rule set that defines relationships and hierarchy between types of documents (i.e., amendments supersede contracts) when determining a state of the aggregated documents. In another rule, a requirement that any defined aggregate of documents must include certain types of attributes and/or a certain number or type of each of the required attributes. Through these various rule sets, configurations or parameters of the aggregated collection of documents may be established and enforced by the document management platform 106.
The computer system 1000 can further include a communications interface 1018 by way of which the computer system 1000 can connect to networks and receive data useful in executing the methods and system set out herein as well as transmitting information to other devices. The computer system 1000 can include an output device 1016 by which information is displayed, such as the display 300. The computer system 1000 can also include an input device 1020 by which information is input. Input device 1020 can be a scanner, keyboard, and/or other input devices as will be apparent to a person of ordinary skill in the art. The system set forth in
In the present disclosure, the methods disclosed may be implemented as sets of instructions or software readable by a device. Further, it is understood that the specific order or hierarchy of steps in the methods disclosed are instances of example approaches. Based upon design preferences, it is understood that the specific order or hierarchy of steps in the methods can be rearranged while remaining within the disclosed subject matter. The accompanying method claims present elements of the various steps in a sample order, and are not necessarily meant to be limited to the specific order or hierarchy presented.
The described disclosure may be provided as a computer program product, or software, that may include a computer-readable storage medium having stored thereon instructions, which may be used to program a computer system (or other electronic devices) to perform a process according to the present disclosure. A computer-readable storage medium includes any mechanism for storing information in a form (e.g., software, processing application) readable by a computer. The computer-readable storage medium may include, but is not limited to, optical storage medium (e.g., CD-ROM), magneto-optical storage medium, read only memory (ROM), random access memory (RAM), erasable programmable memory (e.g., EPROM and EEPROM), flash memory, or other types of medium suitable for storing electronic instructions.
The description above includes example systems, methods, techniques, instruction sequences, and/or computer program products that embody techniques of the present disclosure. However, it is understood that the described disclosure may be practiced without these specific details.
While the present disclosure has been described with references to various implementations, it will be understood that these implementations are illustrative and that the scope of the disclosure is not limited to them. Many variations, modifications, additions, and improvements are possible. More generally, implementations in accordance with the present disclosure have been described in the context of particular implementations. Functionality may be separated or combined in blocks differently in various embodiments of the disclosure or described with different terminology. These and other variations, modifications, additions, and improvements may fall within the scope of the disclosure as defined in the claims that follow.
This application is related to and claims priority under 35 U.S.C. § 119(e) from U.S. Pat. Application No. 63/275,801 filed Nov. 4, 2021, entitled “System and Method for Building Document Relationships and Aggregates”, the entire contents of which is incorporated herein by reference for all purposes.
| Number | Date | Country | |
|---|---|---|---|
| 63275801 | Nov 2021 | US |