What problems is DocuPipe solving and how is that benefiting you?
We were faced with the prospect of taking thousands of large, dense, unstructured legal documents with wildly inconsistent formatting and terms, and dissecting them for numerous data points which would give us a feel for the landscape of a legal practice area in a particular state. This could in theory be done manually, but it would take hundreds of person-hours with no economy of scale: every document takes approximately as long as the last, per page, to the upper limit of the person reading the document, stripping data out of it, and inputting it into a spreadsheet. In the end, you get your output, but the next time you have to do it - and we anticipate this being an ongoing process - you're making another huge investment of person-hours. If the end product was vital, requiring perfect accuracy in the dataset, that investment of person-hours would unfortunately be necessary: nothing LLM-powered is accurate enough yet that I'd be willing to stake the fate of anything seriously important on one's performance. But for a project like this which provides some real benefit to our decision-makers, but isn't so critical that we *absolutely must* ensure the accuracy of every data point, we were willing to consider automation. Our team was willing to accept around 95% accurate output, but even that low bar was unattainable by most of the services we trialed, and I was looking for at least 99% on integer fields. DocuPipe not only managed to meet and exceed those requirements, it "showed its work" by citing sources within the documents to back up its claims. Having the dataset processed through DP also opened up the possibility of interrogating the dataset itself in a RAG-like fashion, but we haven't dived into those features yet. Review collected by and hosted on G2.com.