[Webcast Transcript] Getting AI Right in eDiscovery: Quality, Validation, and Results

Editor’s Note: An AI tool can return an answer in seconds, but the harder work begins when someone asks how it got there and whether they can trust the result. In this HaystackID® webcast, moderated by Todd Haley, Young Yu and Jeff Fleming examined the decisions behind defensible AI use in eDiscovery, from where sensitive data goes and who can access it to how teams test review results before moving forward. Along the way, the panel addressed vendor terms that deserve closer scrutiny, AI features embedded in existing software, and the difference between a human who checks the work and one who clicks approve. Their discussion connected governance to the daily work of document review, where clear responsibilities, carefully developed prompts, and documented validation give teams a basis for explaining their decisions. Drawing on experience with FTC requests, Yu also discussed demands for precision and recall metrics, prompt disclosure, and issue-level reporting, along with the practical challenges those requests create and opportunities to seek modifications. Read the transcript below to learn why AI shifts more preparation to the beginning of a review, where human judgment remains essential, and what teams should document before they trust the results across an entire document population.


Expert Panelists

+ Jeffrey Fleming
Managing Director, HaystackID

+ Todd Haley [Moderator] 
Executive Vice President, Operations, HaystackID

+ Young Yu 
Senior Vice President, Advanced Analytics and Strategic Solutions, HaystackID


[Webcast Transcript] Getting AI Right in eDiscovery: Quality, Validation, and Results

By HaystackID Staff

The review moved quickly, the AI assigned its calls, and the team started to see a path through a document population that once looked unmanageable. Then came the questions: what did the model miss, who checked its decisions, and what could the team show if someone challenged the results?

Those questions drove this HaystackID® webcast, “Getting AI Right in eDiscovery: Quality, Validation, and Results,” hosted by EDRM, where moderator Todd Haley joined Young Yu and Jeff Fleming to examine the work behind a defensible AI process.

Each speaker approached that work from a different part of the operation. Haley, HaystackID’s Executive Vice President of Operations, connected the discussion to the decisions teams faced when introducing AI into legal workflows. Fleming, Managing Director of Advisory Services, examined where data went, what vendor agreements actually permitted, and who took responsibility when a system produced a wrong answer. Yu, Senior Vice President of Advanced Analytics and Strategic Solutions, brought the discussion into document review, explaining how teams developed prompts, measured performance, and validated results, with practical observations from engagements involving the FTC and DOJ.

The conversation got specific about details that teams could easily overlook: a vendor switching on an AI feature inside software they already used, a model update changing the answers, or a reviewer approving output without examining it. Yu also described requests he had encountered for precision and recall metrics, prompt disclosure, and reporting by individual issue, along with the challenges those demands created. Across these examples, the speakers returned to the preparation that made faster review possible: understanding the tool’s purpose, assigning responsibility, testing against human-reviewed samples, and keeping a record of how the process improved.

Read the transcript below for the full discussion, including where human review remains necessary, why prompt development requires upfront investment, and how legal and data experts can work together to meet validation demands.


Transcript

Mary Mack
Thank you for joining today’s HaystackID webcast, “Getting AI Right in eDiscovery: Quality, Validation, and Results,” hosted by the Electronic Discovery Reference Model. I’m Mary Mack, CEO and Chief Legal Technologist at the EDRM. Today’s expert panel is led and moderated by Todd Haley, and we have Young Yu and Jeff Fleming with us as well. We’re recording today’s webcast for future on-demand access, and as with all HaystackID webinars hosted by EDRM, the recording will remain available on the EDRM Global Webinar channel through the next quarter to support your continued learning and reference needs. And before we begin the discussion, a few notes about the Webinar Council. If you look at the top of your screen, you’ll see the HaystackID logo, which you can click on to learn more about HaystackID. You’ll also see an option to contact the team at HaystackID directly, along with the speaker bios where you can learn more about today’s presenters. And moving down to the Q&A box, this is where you can type your questions for today’s faculty, and we highly encourage you to do so. We’ll be answering questions both during and after the webinar. And below the Q&A, you’ll find today’s resources, including the slide deck, a link to HaystackID’s AI governance services, and a link to the next HaystackID webcast, “Trade Fraud Investigation Best Practices,” that’ll be happening October 7th at 12:00 Eastern, and we’d love to have you join us again. Lastly, you’ll see emojis down at the bottom, and we encourage you to use those as well. And our normal disclaimer: all opinions here are those of the faculty, not their clients and organizations. This webinar is educational in nature, and no legal advice is being provided. And our expert faculty will be moderated, as I said, by Todd Haley. He is the Executive Vice President of Operations at HaystackID, and he plans, directs, coordinates, and oversees the operational activities at HaystackID, ensuring the development and implementation of efficient operations and cost-effective systems to meet the current and future needs of the eDiscovery lines, as well as client engagements. And without further ado, Todd, would you please take the reins?

Todd Haley
Thank you, Mary. I really appreciate it. As Mary said, my name is Todd Haley. I’m the Executive Vice President of Operations, but today I am here as the moderator to talk a little bit about how quality and validation of AI is absolutely essential as we move into the AI solutions, uh, uh, in very different areas of the law and other, uh, investigations. Um, with me is Young Yu, the Senior Vice President of Advanced Analytics and Strategic Solutions at HaystackID. Young Yu is the head of our Legal Data Intelligence group. He is responsible for ensuring that, with the operations around both generative AI and analytics are understood and handled by both our clients and our internal personnel, and he is a leading expert in managing and talking about AI when it comes to working with the DOJ and the FTC. Also with us today is Jeff Fleming, the Managing Director of Advisory Services at HaystackID. He’s an experienced cybersecurity professional and also an expert in AI governance. So he will be talking today a little bit about his knowledge and expertise around AI and the best ways to handle, uh, the governance of such. Um, with that, I’m gonna cover a little bit about the agenda. First, we’re gonna talk a little bit about the framing questions, which I will do. I will then turn it over to Jeff to talk about AI governance, and then I’ll turn it over to Young to talk about AI data, uh, and the name normalization FTC information. So let’s talk about the framing questions. A couple things that we wanna talk about is, are you regularly using an AI tool at any point along the EDRM and how you are using it and how it gets validated once it’s used? Can you easily explain the process that the engine uses in arriving at its results? Can you also manage the data prior to it entering that engine in order to ensure that the governance of the AI is correct? And then most importantly is how are those re- results validated? This session is really about defensibility, not just capability. And so we’re looking to have these experts give you a little bit of guidance on how that will work. And with that, I’m gonna turn it over to Jeff to talk a little bit about AI governance.

Jeffrey Fleming
Excellent. Thanks, Todd. Uh, good day everybody. So everything I’m gonna talk about really comes down to, uh, two higher questions is first is where does your data go and who’s accountable when the answer is wrong? And so this is about where the model lives, uh, and how it’s built, because those are generally going to be the two things that, uh, open you up to the most risk versus anything that vendor marketing may have or other things that you’re doing. So today we’re gonna talk, we’re not gonna get into code, we’re not gonna discuss complex math. Uh, I’m gonna give you an overview of vocabulary, some deployment choices, uh, and if you’re not sure where to start, a few things you could do as early as later this afternoon or tomorrow morning, uh, and really kind of generally set the stage, and then I’ll turn it over to Young to discuss AI and EDRM specifically and the FTC ruling that, uh, Todd had already mentioned. So first and foremost, when people say AI, what do they mean? Uh, and as we take a look at this, there’s three kind of instances in terms of model control that exist out there. So from the left, you have the least control, which are public tools that you’re very familiar with, the ChatGPTs, the Clauds, uh, different things like that. They’re free because as with most things in software that are, you are the product. So anything you put in there is going to be retained, they’re gonna use it to train the model, uh, and make it better and, and more valuable for you, for you to use. The private business instances are gonna be those same names you recognize, but the, uh, the data stays in your tenant. You have some more controls. There’s some other things you’ll see in there kind of in that bottom bullet, specifically legal tools, uh, that may sit on top of those models or run in that type of instance, all the way to the most control of models you run. Takeaways here are, especially in that middle one, you have to very much pay attention to the terms and conditions with every one of those you sign on for. Each one of those is gonna be different and each one of those is going to potentially be different based on the tier of service you’re contracting with them. So your procurement department, your contracting and legal review teams really need to look at those, uh, agreements on paper and make sure that those meet the requirements that you have based upon what you’re going to do in, uh, in those models, data types and things of that nature. And then you also need to work with your IT folks and your security team to make sure that you can actually validate, uh, and monitor that those controls are being upheld. All the way over the most control, uh, we’ll talk about those here in a little bit, but you own everything with those. You see some of those models there, those are on your servers, uh, but then you’re also more responsible for making sure the models are updated, a lot more drift monitoring and some other things that we’ll talk about in a little bit on a further slide. Key question here, especially as we talk about privileged data and, and data, uh, for legal review and things of that nature is does your data train the model? The answer here is gonna be an it depends. It’s gonna depend on the tier and the settings that you have with your provider. So what does it mean to say, is it trained on? Everything you put into that model from the prompt to the data uploads is going to be used by that model provider to refine the model and make it better. Uh, if you’ve ever noticed in ChatGPT and, and Claude and some of those, it’ll give you a little thumbs up, thumbs down or a comment box so you can, uh, critique the response you got. It’s gonna take that entire conversation and it’s gonna use that to, to work on making the model better at a larger scale. Some of the agreements, it may retain the data, but not use it for training. And you can see there it’s potentially for logs, uh, abuse review. And so it would still be discoverable, potentially still, uh, found in a, you know, when there’s a breach of, of that model. And so you need to know how long is it gonna retain it and who can see it. And then, and the easiest ones, as soon as it runs the data, uh, and you get your answer back, it goes ahead and gets rid of the rest of it. Uh, but one of the key things there is, is a lot of times you’ll see them saying, “We don’t train on your data,” but that is not the same as them saying, “No one else can touch your data.” So going back to the contracting thing and terms and conditions, not only how are they processing it, but do they use sub-processors potentially for any of the actions as well that may give them access to your data? So a few different words and things you’ll hear today. I’ll leave that up there for, for you to take a look at. But one of the key things that you need to understand with AI is if I feed the same model the exact same prompt multiple times, I am always going to get a different variation of that answer. It’s not a fault of AI. It doesn’t mean it’s going off the rails. It’s just simply how it works. So, uh, the overall, uh, synopsis of the answer should generally be the same, but the very specifics may be different. This becomes a critical point when we’re talking about explainability and defensibility, um, and some of the things that Young is gonna get into later. But, uh, you’ll hear me say it a few times: if you can’t, in plain language, understand how you got the answer you got based on the inputs you put into the model you used, that is going to cause you problems, uh, if it, uh, ever gets questioned in a legal setting or something of that nature. So you need to understand, regardless of the prompt you put in, how you got to the answer that came out of whatever, uh, AI tool or enabled tool that you were using. So data, we’re talking about sensitive data, we’re po- potentially talking about privilege or confidential data, maybe regulated data. What are the, at, at a very high level, these are the two different types of, of data locations you’re gonna have. It could be, be up in the vendor cloud, and you can see some advantages there, makes it highly scalable. It’s also going to potentially shift the security burden more to the provider, but as it’s been mentioned, it’s gonna depend on your tier and your contract. So paying attention to those terms and conditions is extremely critical. Also, again, the IT folks need to monitor that and make sure if the data’s going there, what data’s going there, and making sure it stays there. If you keep it in your environment, and this would apply to if you run your own model or potentially the enterprise private models, depending on the agreements you have, same thing. Potentially scalability issues if you have, uh, it could get expensive; you may have more burden for the security, um, but you’re potentially gonna keep your data there. So really paying attention to the terms and conditions and understanding who has the data, and then also who has the responsibility to secure the data. AI is coming from inside the house. It’s, it’s likely already in use in your organizations today, whether it’s sanctioned or unsanctioned. So, understanding what AI is there and how to find it. So you can see on the left, everybody knows people can sign up for their own. Lots of things are now getting browser extensions and meeting note takers. It’s starting to show up everywhere. Two key things though to pay attention to are that second bullet on the left there is where AI gets switched on inside something you’re already using for where it may not have been on before. How do you monitor for that? What are your contractual protections and things? And then same thing, uh, on the bottom left there is maybe it’s not an AI feature of the program, but that company’s using AI somewhere on the back end. So need to be able to monitor for that, have contractual things in place to protect, uh, the organization, but also just understand because you need to be checking that. Sometimes they’ll bury it in quarterly- quarterly releases. Sometimes they’ll use it for QA, QC. You need to know if you’re sending your data there, what they’re doing with it, because at the end of the day, uh, you as that organization are gonna be held liable. You can’t simply point at the vendor and say it was on them. That is a pre- precursor. So why is governance a differentiator as we go forward? The regulatory environment is already evolving rapidly. There’s the EU AI Act, lots of other different countries are starting their own. Multiple states here in the US have their own regulations going. There’s the executive order attempting to streamline everything here. There’s already litigation in place. Uh, there’s class actions, there’s fines, reputational damage. Uh, it’s making the headlines pretty quickly when it happens. So you need to have a governance program stood up. It needs to be something that is defensible, uh, and that will help carry you through as you, uh, use AI in the world today. So, all right, what can we do to stand up additional governance? Name an owner. Somebody has to be accountable. An AI governance committee is great, or assigning AI to another committee is great, but it’s, uh, it’s, it’s the old Spider-Man GIF of everybody’s point at everybody else. If there’s not somebody accountable, whether it’s the chair of that committee or a C-suite executive or some other, uh, managing partner, somebody needs to be the owner for that. They can have support, and they should from different groups, like legal, IT, uh, procurement, and things, but somebody has to be, uh, the responsible owner for that. You need to understand what AI is in use, potentially what shadow AI is in use, and what each system is used for. So you may allow your employees to use public ChatGPT to figure out the best restaurant to go to for lunch, but you probably don’t want them putting proprietary client data into that model. It needs to go somewhere else. So define what types of AI can go where, write down the rules. Uh, you can have their list of, “Hey, this data should never go into AI if you want to.” And then now once you have those processes and policies in place, you need to have the technical ability to check and validate to make sure that people aren’t putting data that shouldn’t go into places into those places so that you can prove, uh, that your policies and procedures are working. So along those lines, the bare minimum, I said it earlier, but I will say it again, if you can’t explain why and how the model gave you the response it did in plain English, everything else you have is, it’s gonna invalidate a defensible program. So you have to have that. You can see some things here on the slide to get covered. Uh, another key thing is: what is the incident response or disaster recovery or business continuity plan for when AI goes wrong? You can see at the, uh, bottom right there; you need a clear way to report an AI mistake or an AI going off the rails. But what is the process for that? What are you gonna do? Is it turn the AI off and revert to manual processes? Is it add more humans in the loop along the way? Depends on what we’re using it for. But that is another thing you need to have in place is, is be prepared for when something goes wrong with it so that your business can continue running. But all those things there are definitely bare minimum requirements for something that is defensible. It does not need to be 90 pages of, of documentation to have these policies around. These can be very short, very quick and concise documents. And then have the proof, right? Have the policies, have the process, have the proof. Same as you do for anything else you’re gonna audit. Add this to your annual audit cycle and, and it should just fall into that type of cadence. It’s not just something, though you can stand up and walk away from. This is something that requires monitoring and maintenance. Some of the key things is you need to, uh, as mentioned earlier, know what AI is being used and are employees are bringing new shadow AI into the environment. Are your vendors turning on AI features? Are they using AI subpro – or sub-processors using AI things, um, on your data? Also need to watch for model updates. If, as an example, you’re using Private Claude and it goes from, you know, Sonnet to Opus as the default, you’re gonna start getting some markedly different answers pretty quick and you need to be able to understand why that change happened because the model changed or make sure that models, you know, you revert to the old model that better suits the needs of the organization and the output. Especially if you’re running your own model, you need to make sure you’re maintaining and monitoring fairness testing, uh, any sort of mistakes, hallucinations, inaccuracies, and things that may need, uh, your involvement to make sure you, you know, revert to an older model or pause the AI process and get it back in check before you continue, uh, using the locally hosted one. And then, as we mentioned, the legal landscape- gotta monitor that. That’s always changing, especially if you’re in multiple geographic areas. You have to keep an eye on that because that may trigger new, uh, requirements, uh, or new things that you need to do, new reporting requirements, things of that nature. So –

Todd Haley
Hey, Jeff, we have a question. Yes, sir. The question is, you’re talking a lot about humans, uh, monitoring and looking at this, uh, information. Are there tools or technologies out there to assist with the monitoring of AI, uh, when managing this ongoing maintenance and monitoring?

Jeffrey Fleming
Yeah, certainly. There are, there are lots of tools, uh, in popping up to do these sorts of things. Um, some of them are, if you’re using some of the private enterprise setups, they’re very, they’re native to it. M365, I know, has some, uh, security suite and operation suite tools already nested in the M365 environment that simply need to be, uh, you need to ensure turned on and configured. Your IT security teams can use, uh, the tools that they’re already using to monitor the network, whether they’re internal or, uh, managed service providers. And there’s different things they can turn on in the tool suites that they’re already using. So it doesn’t necessarily require a whole new tool set. Uh, there are some vendors out there that, uh, do offer things like monitoring for shadow AI and, and tagging and logging and, and some additional features that go on if you wanna, uh, pursue a, an, an additional add-on solution for that. Uh, but yes, there are lots of different things in there, um, and, and ways for companies to go ahead and do that.

Todd Haley
Perfect. Thank you.

Jeffrey Fleming
So thank you for that. Uh, and then one more thing I’ll add, uh, as we go here to the summary on, on human in the loop. One of the key things is, is if your, uh, if your process requires a human in the loop, do you ensure that that human is actually stopping and doing the review and not simply a person just clicking okay, uh, as more of a, a, a widget factory to go through, because that is not human in the loop. That is not, uh, there are s- several cases out there pending that I’ve seen, uh, waiting for the results to come out, but where the, uh, the person just clicking okay is not the, uh, I’ll say due diligence or best standard for quote-unquote human in the loop. If you’re gonna have a human in there, they have to actually be doing the review. And when they click okay, that human becomes responsible for whatever the AI answer was. Um, and there’s a few, uh, especially as it pertains to case filings by attorneys even, where the attorneys were, are being put before their respective state bars because they just said, “Yep, looks good. Let’s file this document with the court.” Um, and it turns out that there were some hallucinations in either made-up cases or made-up quotes, uh, or precedents in, uh, actual cases. So human-in-the-loop needs to be true doing review and paying attention because that human is going to be held accountable. So, uh, there are some, some five steps to your defensible program. Uh, as I said, it’s not a project. It has to be an enduring program. You have to have somebody focused on it, um, and make sure you’re doing regular check-ins to keep it going right. You’re going to need AI to compete and stay relevant in the world today. It’s just the way things are going. Uh, it may take over small parts of industries; it may take over large parts of industries, but AI is making things easier. It is helping folks digest large corpuses of information in really rapid fashion when it’s used and tuned correctly. So to stay relevant in the world today, you’re gonna have to be able to use AI, uh, and wield it as a tool. But you need to have the proper governance in place. So don’t be scared of it, you just need to be responsible with it when you do it. And that’s where the governance comes in and helps. Um, but wanted to set the stage, uh, and make sure everybody’s kind of on the same level with some common terms and phrases, but right now I’m gonna turn it over to Young, uh, and he’s gonna get more into AI and EDRM today.

Young Yu
Oh, thank you very much for that. And, um, it’s important to note that for, for context here in the EDRM, many of the, uh, the offerings, the AI offerings, they’re purpose-driven, right? So, um, what I mean by that is you have AI that will help you determine privileged documents from non-privileged documents, or let’s say responsive documents from non-responsive documents. So these are purpose-driven. Um, I would say the vast majority of these are closed models. So they are not learning from your inputs. So the, uh, you know, the responsiveness criteria or the privilege criteria that you’re entering, it’s not learning from any of that or the text that’s being submitted at a document level. It’s not learning from your outputs. It is important to know that, um, understanding this will go a long way in, in providing you a pathway and sort of defensibility here all around, right? Because let’s say you have a model that learns on your inputs and outputs and, um, you know, you’re doing, let’s say, commercial litigation. Uh, the next time that you’re faced with a different type of litigation, all that training, if it does train, will affect the, uh, the calls for responsiveness or privilege made on the other side, right? So you really don’t want it to learn. Um, the other thing here is, uh, sort of with these closed models, um, we’re simply adding to the overall prompt, right? So these are purpose-driven. There’s going to be a layer where each provider has a piece of that prompt, and it’s asking it, you know, “Hey, how are you checking, or how confident are you in this determination?” So with that, um, this first slide here is really just an overview of the EDRM. I’m sure everybody’s familiar with this slide here. And, you know, sort of breaking into each phase here, you can see where we’re, where you can deploy AI. Um, I think it’s important to know where you decide to implement; just like Jeff was saying, you need to have a governance process and then also a layer of validation. So the validation layer here, um, and, and, and sort of when we talk about each of these sections, um, they can be more or less stringent, right? So if you’re using it to pre-classify, uh, let’s say privileged markers for, you know, correspondence, uh, you know, that’s a, a feature set within Copilot, you’re not omitting, you’re not screening anything, you’re just pre-flagging, right? So there may be less of a burden to, uh, you know, to validate those markers because everything is still getting put out, um, you know, during, uh, collection and processing. Um, but if you were to apply it, let’s say, later down the line during review, there is likely going to be a more stringent, um, series of validations that you’ll need to take. Um, sort of more granularly here, this breakout shortest, it, it shows in terms of where you can deploy purpose-driven AI, right? So these are solutions that are offered by your vendors, by Relativity, um, by other providers here in the market. And, um, you know, ECA is one of them where, you know, let’s say traditionally we would have used search terms, right? Now, as we talk about validation, I would say that is, uh, it’s very unlikely that search terms require validation. Most of the times, um, search terms are agreed to by both parties, and then you just sort of move forward, right? And you never know what’s being left behind. Um, I’d, I’d say with AI, that requirement is, is, is much higher, right? They’re gonna wanna know, okay, how did you train the AI, or what did you use as your prompt criteria or how did you get this corpus, um, you know, out of ECA? What did you do there? Um, with purpose-driven solutions, I think understanding, um, sort of how they work is important. Um, you know, it’s, you want to be over-inclusive, right? And a lot of the solutions are, but understanding that will help you determine, um, your next phase of workflow. Um, relevance and issues, right, this is really during review, right? So you’ve now collected your data, you’ve processed your data, you’ve applied some layer of filtering, whether it’s search terms, um, let’s say you’ve used, uh, like, you know, email threading or something like that to reduce that population, and now you’re in review. Um, relevance and, uh, issues identification, so, uh, sort of replacing, let’s say, first-level contract review or supplementing. Um, this area here does require a fair bit of validation, and also an understanding. As Jeff said, uh, earlier in the session, if you ask AI a question 100 times, you’re likely gonna see some variance. Understanding that variance helps you understand where you’re making improvements as you change criteria for your prompt, for your definitions of responsiveness. And it, it’s good to understand or, or have context. Uh, and what I’ll say here is typically we see about 1% variation. What that means is if you, uh, submit a thousand documents, um, 10 times, five times, whatever it might look like, your calls for responsiveness will change about 1% in terms of, uh, recall and/or precision. And what that really means is that that model understands the language, but sometimes it’s gonna get a handful of documents correct or incorrect. And what that really allows us to do is, “Hey, we made this, you know, we ran this the first time, let’s say we got 90% recall. The next, we made an adjustment, and the next time we ran it, we got 91% recall.” Well, we know there’s 1% that’s, uh, deviation, right? That variance. So potentially you couldn’t, like, there might not have been a material difference to that language change. You want to understand that. You need to understand that. Um, at the same time, um, you know, 90% recall is great. Uh, the flip side to that is precision, right? So recall is, of all the responsive documents, how many did you, uh, actually capture? And precision is, of all the documents that you’ve marked as responsive or not responsive, how many times, uh, are the calls correct, or what is the percentage of calls that are correct? And, you know, it’s a fine balance that you need to play here. Many of these terms are, you know, used in TAR, active learning. They’re the same, you know, definitions here. Um, in terms of privilege analysis, again, another purpose-driven solution, um, typically these, uh, these models or, or these solutions, what they require is that you identify, uh, uh, individuals or corporations that, uh, generate privileged content or pert – perhaps break privilege content. And this is going to analyze, uh, your documents to see if they are indeed privileged. Some of them will, um, say this is a protected or privileged act. Um, and a lot of the work here has been done for you. And what I mean by that, um, earlier in the session, Jeff had mentioned, hey, how do you check for hallucinations? How do you make sure that the AI is, is explaining itself? Many of these solutions are gonna provide you a reasoning or a rationale as to why those determinations were made. Now, I’m not saying if you submit a million documents, you need to read every single one of these. Um, you would like to do some sampling and, you know, how large that sample is, it’s really incumbent upon, um, let’s say counsel or, you know, you as the, uh, the user of AI to, to make that determination. If it’s your first sort of go at this, a larger sample; if you’re getting used to the process and you can understand where things may go wrong, a smaller sample may be, uh, a better place to check. But it’s, it’s fairly easy to see with the solutions that are out there, um, where the AI is getting things wrong or, or things that don’t line up. And these things can also be used to adjust your prompt language, right? Um, PII, PHI identification, again, during that review process, I think, um, you know, sort of in today’s day and age, with the amount of data that we have and, you know, especially sort of with consumerism being what it is, I think a lot of that data is present within the larger corpus. And traditionally, we’d identify those by using, let’s say, regular expressions, um, or, you know, searches for, let’s say, you know, date of birth or, you know, Social Security number terms like that, right? Um, generative AI does a better job of sort of compiling that information together and saying, “Hey, this exists, or it doesn’t,” or maybe it gives you a density. So, um, you know, this document has 30, 40%, you know, PII, PHI. It lets you make a determination faster as to whether you’re going to redact this document or withhold this document, right? And then, uh, the last piece there is, uh, natural language interrogation. And what that really means is it’s like a chatbot; you can ask questions freehand. And, and understand that, um, you know, you can use that almost, um, pretty much in every phase here, right? Let’s say during ECA, if you wanna get a better understanding of the data corpus, you know, during review, you can ask, you know, questions sort of on topic or off topic, um, you know, during that production phase, if you’re using it for, you know, depo prep, witness prep, um, and that’s all in line and getting ready to, to present, you know, your, your evidence or, or those documents how, in whatever light you want to, right? And with that comes a, a framework that you need to establish. And when you generate, um, results out of AI, you, you, you need to understand sort of what box you’re putting those into. And things you need to consider there would be, let’s say, state or federal, uh, rules that apply to your litigation or, or your use case, right? Um, the documentation and ability to describe what it is you’re doing, um, for that separate piece, um, identifying who the human in the loop is, right, and, and, and really saying, “Hey, this is the person, or these are the people that are making the changes and, and guiding the AI to find, um, you know, the relevant or privilege or whatever set of data, um, data sampling and verification, right? So how did you measure success, how did you identify success, and then measure success, right? And then what metrics are you gonna put out to say,” Hey, we did a wholesome process, “right? So that’s very similar to TAR, right? Precision, recall, maybe illusion, and, and all these things really tie up into, you know, how you are validating, right? Um, and then the last piece is a two-tiered review workflow, right? So what, what that means is – what percentage of documents are you, uh, QCing or, or using to say, “Hey, we didn’t only just screen for this. We did a wholesome check or a representative wholesome check.” There are also gonna be documents that fall out of the, uh, eligible population for generative AI. And, you know, how did you review those, right? How did you handle anything that’s an outlier, right? Or maybe there’s a pocket of documents that’s not conducive to this process, right? So, you know, how did you review those? All this to, to set up the, uh, you know, you, you, you take all of that as a whole, and how do you find the validation and defensible, like, strategy, right? So one, you’re, you’re gonna have to figure out upfront and, and maybe, you know, those guidelines will be forthcoming by, you know, regulatory agency or let’s say, you know, you have meet and confer with opposing. It’s, it’s really about how much you want to disclose. Um, I, we’ve seen some that are fairly crazy, like, uh, we want you to share your prompt language or show us all the, uh, human versus AI overturns. Um, this is an area where you can really sort of hold firm or, or, or, or see sort of what reasonable means, right? Um, in terms of just, like, sort of disclosing prompt language, I mean, the alternative to that is, all right, why don’t you give us the prompt language, right? Um, so it’s really a, a give and take here, but understanding, look, um, a prompt language could or, you know, could be considered attorney work product or, you know, maybe it’s not. It’s, it’s like search terms, right? Do you share search terms? Yeah, absolutely, right? Is that work product? Arguably, yes. Um, it’s, it’s about transparency in the process without giving away sort of everything, um, maybe setting goals for recall. Um, I’m not, you know, we’ve seen as low as 75, we’ve seen high as, uh, as 90. And, and it’s really sort of how much effort is gonna go into the process, right? And I think that’s, like, sort of the biggest takeaway here. AI is great. It’s, it’s easy to use. It doesn’t take away any of the lift in terms of fine-tuning, in terms of drafting the language. Um, everything it takes to draft a coding protocol or a privilege review protocol or everything that you’re considering when you’re drafting search terms- all that energy and time is refocused into drafting the problem language. There is an investment there, um, to, to guide the AI, and then potentially to fine-tune the AI, right? So, you know, that time is, is, is more upfront than it is sort of in the middle or down the line as new things come up. And sort of figuring out what, how, how you can represent all that, right? So, you know, you take a sample, you have that sample reviewed, and then you wanna prompt against that sample. Um, you get a baseline for metrics. If you make improvements in the language, did those metrics go- did they get better? Did they get worse? Which ones worked? Which ones didn’t? And then, you know, whether you decide, hey, let’s have a, let’s do a control set so we can get a better read, or let’s do illusion testing to see what we’ve missed. You know, these are all just sort of different me- methodologies in terms of how you want to validate, right? And while being transparent in the process, I do think it’s probably most important to be transparent in the validation method, right? And also figuring out how you’re gonna go about this process. So, framework, again, who are the humans in the loop? How are they attacking the data and the problem? What’s the language and the metrics that, um, we’re going to use or disclose? Um, and then how do you show improvement? And then how do you prove that you didn’t leave too much behind, right? Uh, all this, again, is, is sort of wrapped up in the court or jurisdiction that you’re in, who the other side is, whether it’s a regulator or opposing counsel type of litigation. And, you know, a lot of this really ties up to your document corpus itself, right? Um, you have 10 million documents, and that’s what you need to review. You may wanna figure out just randomly sampling what percentage of this huge population is actually relevant. Do we wanna throw AI at everything or do we want a hybrid workflow? And, and that leads to a little bit more of a, uh, I’d say involved validation and, and, and one that’s, uh, you know, a little harder to tackle, but it can’t be done. Um, implementation and defensibility, right? So the governance aspect, again, you need that framework, you need a process, you need to understand, um, what’s going into it, what’s coming out of it, right? If it is a purpose-driven solution, um, you need to be able to define that and explain sort of what it is a human is entering and reviewing at the end, uh, your method of validation. Um, what I would say is if you’ve used TAR active learning before, um, much of the, uh, the metrics, you, you know, they, they are used for, for generative AI. They’re not the only ones that are used, but, you know, you can certainly refer to those. Um, I would say that the standard may be slightly higher for generative AI, at least, uh, you know, let’s say in the present, they may be, uh, you know, a little more loosened, uh, down the line here. Human oversight: understanding who your subject matter experts in, who has, uh, that feedback loop with AI, identifying those individuals, right, and figuring out where those decision points are and how those sort of decisions are made into your next phase of workflow. And then the underlying reporting in terms of, okay, here was our starting point, here’s how we made improvement, here’s our validation process, and potentially here’s what we’ve left behind. And, and, and really, if you think about this sort of in that context of, of document review and how you’ve used it, uh, how have used TAR before, it, it, it’s very much in that same vein. It’s just a little more time upfront, a little more investment upfront, but in terms of sort of analyzing your document population, it’s much faster than, let’s say, a traditional, uh, first-level review team would be. Um, sort of comes to FTC and DOJ AI guidance. Um, we were, uh, <laugh> sort of met with some challenges in, in terms of what the FTC has been asking for. Um, there have been some changes to the definition of, uh, AI, which does encapsulate, uh, let’s say TAR and CAL, uh, including generative AI. They’ve made that sort of whole now. There is definitely a preference for active learning if you are gonna use TAR. In terms of generative AI, um, there are a bunch of criteria that, uh, have been requested by the FTC. Um, we can sort of jump into this now and, you know, there’s a minimum threshold for, for recall, which has always been there, right? But now there’s a minimum threshold for precision. Um, the FTC is asking for 75% precision. What that means is, uh, you know, out of all the documents you produce, three out of every four documents are actually responsive or, or, you know, related to one of the issues. Um, sh- sharing any human versus AI-authored documents, uh, that’s been another request. Um, tracking issue by issue. So if, uh, you know, you are pr – you know, if the, if the FTC is requiring, let’s say, 35 different, like, specs, um, you’re going to have to show precision recall for each one of those specs, um, which is, it, it, it can be very challenging. It makes the workflow sort of very rigid and the process very difficult to achieve that standard, right? Um, and this is not just for generative AI, it’s also for TAR. Um, there is a minimum cutoff for TAR. Uh, the FTC has requested a minimum s – cutoff of 50, which, you know, doesn’t sound bad until you factor in some of the, uh, tools out there, right? They’re, you know, traditional active learning, let’s say, within Relativity scores all uncertain docs at 50, um, which means if you’re using 50 as your cutoff, you know, there are gonna be a ton of documents there unless you clear them all out, um, through review one way, shape, or form. If you’re recording issues at a, uh, you know, at a, let’s say, at an issue level, right? So precision and recall for each issue, I’m not sure if it means you need to reach 75% recall and 75% precision for each issue, but under that assumption, um, you may s – be stuck in a never-ending loop of, of iteration at an issue level. Um, you know, ag- again, there are modifications you can seek with the FTC. I think this is a good segue into that. Um, the FTC would like you to disclose the language of the prompts, um, and also identify, um, whether you’re using it for privilege, uh, analysis. I mean, it, it’s sort of a little odd, right? Because, um, you know, if you’re, let’s say you’re using, uh, AI for privilege analysis and then having humans review all those calls and either agree or overturn, do you need to disclose because you had humans re- review all those docs anyhow? And then if you had disagreements, do you need to now sort of say, “Okay, here are all the human overturns of privilege.” I mean, I, I, I feel like that’s a very tall ask and maybe a little onerous on, uh, for, for all parties. It does seem like the, the guidance here by the FTC, um, really i- is trying to, to limit what you produce to the FTC and sort of give everything over wrapped in a bow. It looks like there is a question here.

Todd Haley
I got it. Yeah. Um, are we at the stage in using AI where human reviewers are no longer necessary, for example, to accomplish first-pass relevant screening? Put another way, is genera – is generative AI being used or will soon be used in a way to el – that eliminates the need for human reviewers?

Young Yu
I kind of think that’s a trick question, but let’s say this. Uh, I think at, at the front, the front end of this, you will always need a human to generate the language, uh, for, let’s say, relevant documents, right? So, responsive documents. You’re gonna need a human there. You’re still gonna need a human to code documents for your samples that you are measuring against. So someone needs to review those documents. Let’s say you’re using 500 documents, 1,000 documents to train the model. Those documents will have to be reviewed, um, you know, and coded. And then whatever validation you use, if you’re using a control set, you know, anywhere from, let’s say, 385 to 3,000 documents, and the illusion sample, let’s call that 1,500 documents. So there’s still a need for some human review. Um, there’s also gonna be documents that don’t fit neatly into a generative AI workflow. Most of these models do have a context window. What that means is you can only shove so much text in there before, you know, it, it, it, it just doesn’t take anymore. Those documents that don’t fit will have to be reviewed in one way, shape, or form. So, you know, those documents typically undergo human review outside of generative AI analysis. I would say for the most part, um, you can opt to remove first-pass review, but that would increase the burden on your 2L, right? So whoever you’re, you’re, let’s say you’re using contract reviewers for first pass, uh, and you’re using subject matter experts for your second pass, I think you’re, you’re shifting that burden of first review sort of onto your second pass. I wouldn’t say all of it. Um, I would say you are increasing the burden, um, but it’s probably at a ratio of one to three, one to five.

Todd Haley
Yeah, and Young, I’m gonna actually add to that. I think one of the things to think about is the days when review was done on paper and when technology got introduced, they thought reviewers were going to go away. They did not. I would argue that going from t – from that technology to this technology may change what the reviewers are doing or what they’re focused on, but I think it will actually raise the bar of value in that review and may, uh, still allow reviewers to exist at a first-level review, but maybe at a higher calling around more responsive documents. So I’ll add that. Um, and with that, we actually have a second question. Can you share the site or source where FTC requires 75% and also the definition of AI, including TAR?

Todd Haley
Unfortunately, the model second request, uh, by the FTC is still the same as it was, uh, last year or two years ago. Um, this was seen through a modified CID and, um, a second request, um, issued by the FTC this year. Um, nothing that we can’t share because it does have, uh, you know, <laugh> our, their issue to our clients, not to, to, to HaystackID directly. Um, but if you do receive a spec or a request from the FTC, I would expect to see it there.

Young Yu
All right. Uh, just practical takeaways. Um, I, I think understanding the technology, um, understanding purpose-driven solutions, um, will inform you what to ask vendors, right? Um, it is important to, to understand how each one of these solutions actually works, um, and what the goal of each one of these solutions are. Um, you know, we, we talked about where you would drop these in, in terms of, uh, areas in the EDRM, so I would not use, let’s say, a purpose-driven solution for ECA to produce documents, right? It’s a s – it’s a, it’s, it’s, you know, your front-end screen. It’s gonna be over inclusive. Um, what to document internally, build a process, understand the process, understand where humans are part of, you know, the workflow, where the acceptance criteria is, and who is making, uh, that decision to go forward, um, you know, and, and push the go button across, you know, your entire corpus. Who’s doing the validation, right? So who’s reviewing those documents, right? And, and, and what to monitor going forward. Um, the FTC usually has a comments period after issuing, uh, new guidance. Um, there is always room for modification, uh, with the FTC and DOJ. So if you are issued one of these, uh, modified COD, CIDs, or second requests, certainly have conversations a- and seek modifications to either lower criteria or, you know, perhaps, you know, negotiate away, you know, let’s say 75% precision, um, or that cutoff score, whatever that looks like, right? And, and, and look for updated guidance, um, on the FTC and DOJ websites, right? You know, the model second request is still there. Those don’t reference sort of what we’ve spoken about here. This is all coming from, from practical experience, uh, on the service provider side in terms of what we’ve seen come through our door. I think our clients are receiving these as well. I’m pretty sure that, uh, you know, if you’re part of any of those forums, you’ll see this, you know, spoken about pretty heavily there. Um, Todd, uh, anything you want to, uh, add here?

Todd Haley
The only thing I would say from a practical takeaway based on what we’ve been working on together is that a combination of a data expert with a legal expert has done really well at talking about modifications to these regulations and these modified CIDs. Um, so I’m recommending working with someone who has worked with the FTC and DOJ and has the experience to talk through why they are trying to ask for what they’re asking for and where there may be better ways you can give them what they ask for without giving it exactly the way that they ask. I think that’s an important takeaway as well. And with that, I think –

Young Yu
I agree with you 100% there.

Todd Haley

Perfect. Uh, with that, are there any other questions? Uh, if so, please remember to put them in the Q&A box, and we’ll try to answer them now.

Jeffrey Fleming
Hey, Todd, while we’re waiting for questions to come in, um, just back to that first question we had on potentially phasing out, uh, it’s gonna ultimately come down to, I think, two things. One is, if you’re gonna roll your own model and set, how confident are you that it’s gonna produce acceptable results based on thresholds, legal precedents, things of that nature? Or, can you explain it? And two, if you’re gonna use a third-party tool, again, what is your confidence level in that third-party tool to produce the same level of accuracy? Because at the end of the day, whatever is found either to be included or excluded as filings go forward, documents go forward, somebody ultimately is gonna have to sign that, and that person is gonna be held accountable for whatever comes from that AI use. So, uh, it’s, it, it just becomes a risk acceptance problem at that point, is if that person’s comfortable with it, great. But if it’s not producing a high enough level of confidence, I, I think it’s gonna be a little bit before, uh, we could get to a full phase-out.

Todd Haley
Okay. I am not seeing any more questions at this time. Um, I will turn it now back to Mary.

Mary Mack
Thank you, Todd. And, uh, before closing, please mark your calendar for HaystackID’s next webcast, “Trade Fraud Investigation Best Practices,” happening on October 7th at 12:00 PM Eastern. You can find the registration link in today’s resources, and we hope to see you there. Thank you again for joining today’s HaystackID webcast. Thank you to our panelists for sharing their expertise. And on behalf of EDRM, sincere appreciation is extended for your p- participation today. Wishing everybody a productive day. Thank you.


Expert Panelists

+ Jeffrey Fleming
Managing Director, HaystackID

Jeffrey Fleming is an experienced cybersecurity professional with a proven track record of delivering excellence in client relations and operational success. With over a decade of experience in contract oversight and strategic leadership, he specializes in transforming challenges into solutions that yield exceptional results. Fleming’s ability to navigate complex contractual landscapes and build strong client relationships has been instrumental in driving organizational growth and success. As an adjunct professor, he is passionate about sharing his knowledge and expertise in cybersecurity with the next generation of professionals. By providing hands-on guidance and mentorship, Fleming empowers students to excel in navigating the ever-evolving cybersecurity landscape and contribute meaningfully to the industry. In addition to his expertise in contract management and cybersecurity, Fleming brings a wealth of knowledge in cloud solutions architecture. Leveraging his background in cloud technologies, he integrates innovative solutions to optimize operations, enhance efficiency, and drive strategic initiatives forward. By staying abreast of the latest advancements in cloud computing, Fleming ensures that organizations are equipped with the tools and resources needed to thrive in today’s digital age.

+ Todd Haley [Moderator] 
Executive Vice President, Operations, HaystackID

In 2008, Todd M. Haley joined HaystackID and its associated acquisitions and is currently the global data operations group’s Executive Vice President and General Manager. As Executive Vice President of Operations, Todd Haley plans, directs, coordinates, and oversees operational activities at HaystackID, ensuring the development and implementation of efficient operations and cost-effective systems to meet the current and future needs of the eDiscovery service line. Mr. Haley also consults on HaystackID’s processes, procedures, and reports necessary to ensure the handling of all client eDiscovery and data management projects. He brings his experience as Chief Technology Officer, directing both IT and litigation support in a nationally recognized consulting firm, to the forefront.

+ Young Yu 
Senior Vice President, Advanced Analytics and Strategic Solutions, HaystackID

Young Yu joined HaystackID in 2018 as a director and is currently the Senior Vice President of Advanced Analytics and Strategic Solutions. In this role, Young is the primary strategic and operational adviser to HaystackID clients, focusing on the planning, execution, and management of eDiscovery activities. Young brings extensive experience to his position, having previously worked at IPRO Tech as a Professional Services Consultant/Product Manager for Analytics, Wilmer Cutler Pickering Hale & Dorr LLP as a Team Lead/Litigation Support Coordinator, Chadbourne & Parke LLP as a Project Manager, and Ikon Office Solutions in various roles, including Data Engineer and Global Database Administrator. He holds certifications as a Brainspace Certified Admin, Analyst, and Specialist and is affiliated with Agile, Scrum, and Six Sigma.


HaystackID® solves complex data challenges related to legal, compliance, regulatory, and cyber requirements. Core offerings include Global Advisory, Cybersecurity, Core Intelligence AI™, and ReviewRight® Global Managed Review, supported by its unified CoreFlex™ service interface and eDiscovery AI® technology. Recognized globally by industry leaders, including Chambers, Gartner, IDC, and Legaltech News, HaystackID helps corporations and legal practices manage data gravity, where information demands action, and workflow gravity, where critical requirements demand coordinated expertise, delivering innovative solutions with a continual focus on security, privacy, and integrity. Learn more at HaystackID.com.

Assisted by GAI and LLM technologies.

SOURCE: HaystackID

Advisory Note: AI governance becomes more difficult to manage when policies sit apart from the teams building, buying, and using the technology. HaystackID® AI Governance Services bring those responsibilities together, translating expectations for transparency, oversight, and accountability into workflows that organizations can use every day. The services connect product, engineering, and operations teams with legal, privacy, and risk stakeholders to establish clear decision rights, validation routines, and documentation throughout an AI system’s use. That work gives organizations a repeatable way to assess third-party tools, test performance, and capture evidence as requirements evolve. Reusable assurance materials and standardized controls can also reduce procurement friction, streamline customer and partner due diligence, and give teams a stronger basis for approving AI deployments. To learn more about how HaystackID® AI Governance Services can support defensible oversight and responsible AI adoption across your organization, connect with our team of experts.