[Podcast] HaystackID® in the EDRM Illumination Zone: Esther Birnbaum
Editor’s Note: Generative AI (GenAI) has proved its value in document review, but legal teams now face a bigger question: How should the technology change the discovery process itself? In this episode of EDRM‘s Illumination Zone podcast, HaystackID‘s Esther Birnbaum joined Mary Mack and Holley Robinson to discuss what comes next, explaining why adding AI to old workflows limits its value and calling on legal teams to rethink discovery with today’s technology in mind. The conversation explored how Case Insight™ gives teams an early view of their data and how CaseBot™ follows facts across the record. Birnbaum also drew a firm line between finding evidence and exercising legal judgment: AI can uncover and connect facts, but lawyers must decide what those facts mean. Her message is clear: Use AI to find the story sooner and give legal professionals more time to build strategy, assess risk, and make sound decisions.
Rebuilding Discovery for a World Where AI Already Exists
By HaystackID Staff
For most of the past two years, the questions hanging over legal technology have been narrow and mechanical: Does GenAI work for document review? Can it classify documents accurately? Can the technology hold up to scrutiny in court?
Those questions mattered, and largely, the industry has answered them. But according to Esther Birnbaum, executive vice president of data intelligence at HaystackID, answering them was never the real goal, and treating them as the finish line risks missing what needs to change.
“The biggest change is that the conversation has moved from does GenAI work for document review to where does GenAI fit into legal work and how do we use it responsibly at scale,” she said.
Birnbaum returned to EDRM’s Illumination Zone podcast recently for a conversation with Mary Mack and Holley Robinson, roughly eighteen months after her first appearance in March 2025. Back then, the conversation centered on bringing GenAI into her in-house experience and applying it to document review. This time, the framing was different. The industry, she said, has moved past asking whether the technology works.
The harder question now is where it belongs, and whether legal teams are still building discovery processes as if the technology didn’t exist yet.
Starting From the Technology, Not the Old Process
The limiting mindset in legal tech right now is familiar: teams keep asking how to slot a GenAI feature into a workflow designed before GenAI existed. First-level review, privilege review: these are places where the technology has proven itself.
“We’ve proven that GenAI can work there,” Birnbaum said. But proving that a feature works inside an existing process is a different exercise than asking whether the process itself still makes sense.
That distinction matters more than it sounds. Treating AI as an add-on to a pipeline built before the technology existed puts a ceiling on what it can do, because the pipeline was designed around a set of constraints—how much a team could realistically read, how information had to move linearly from collection to review to production—that no longer apply the same way. Slot GenAI into that structure, and it can make individual steps faster. It won’t change the shape of the process itself. The more consequential work is reimagining discovery around the assumption that this technology is simply part of the toolkit from day one, not a tool bolted onto a process that predates it.
“What we’re really trying to do is reimagine a world where our discovery process is built in a world where GenAI exists,” Birnbaum said, “rather than just where you can insert a feature into a process that already exists.”
That’s a meaningfully different design question. Inserting a feature asks: where in our current process can AI save time? Reimagining the process asks: if we were building discovery today, knowing what technology can do, what would we no longer need to do the old way? The first question tends to produce incremental gains inside a fixed structure. The second tends to produce something closer to what HaystackID has built over the past year: tools that don’t map neatly onto any single step of the traditional EDRM funnel, because they weren’t designed to.
That distinction sounds abstract until you follow it through the specific tools and decisions behind it—because nearly all of them trace back to the same underlying instinct: stop asking where AI fits, and start asking what discovery should look like if you were designing it today.
Understanding a Data Set Before You Review a Single Document
The clearest expression of that shift is Case Insight™, HaystackID’s tool for evaluating a data set as a whole before review ever begins. Traditional review looks at documents individually, or at most within a family. Case Insight instead classifies every document in a collection contextually, against everything else in that same collection, without predefined categories. The categorization is generated entirely by the AI, based on what’s present in the data.
The practical effect shows up earliest in early case assessment, where matters routinely involve five or ten million documents. Case Insight gives teams a substantive read on the data up front, identifying what should move forward and what shouldn’t. That’s a genuine departure from the traditional model, where teams historically had to wait for a linear review process to develop any real understanding of what a data set contained.
The numbers behind that shift are concrete. In one matter, Case Insight classified 92% of a 276,897-document population as not relevant, cutting review time by 72% and cost by 60% compared to a traditional continuous active learning approach. In a different matter, a 1.29-million-document criminal production review was narrowed to less than 9% requiring attention using Case Insight.
Case Insight also generates a narrative—key people, key events, key documents—shaped by whatever case background a user provides. But the tool includes a deliberate safeguard against tunnel vision: it doesn’t limit its findings strictly to the issue it was pointed at. Birnbaum illustrated this with a synthetic data set her team constructed around an insider trading scenario, into which she buried a separate, unrelated FCPA issue.
When Case Insight analyzed the data for the insider trading narrative, it still surfaced the buried project as a potential concern worth investigating on its own. That’s the tool doing exactly what it’s designed to do: flag what else matters in the data, so a team knows where to look next, rather than trying to resolve every issue itself. A deeper investigation into that second issue is a job for a different tool—one built specifically for that kind of drill-down.
Turning a Flag into an Investigation
That’s where CaseBot™ comes in as the tool built for the moment a team needs to go deeper into something Case Insight surfaced. HaystackID introduced the next-generation version of CaseBot, technology from eDiscovery AI®, around ILTACON, moving the tool from conversational search into what Birnbaum called agentic retrieval. Rather than answering a single search pass, CaseBot works iteratively: deciding what information it needs, searching the record, reading what it finds, identifying gaps, and repeating that cycle until it has enough to answer.
The reasoning behind that iteration gets at something core to how legal fact-finding works in practice. An email references a meeting. The meeting leads to a custodian nobody had prioritized. That custodian’s files contain an attachment that contradicts a statement made months later. Following that chain requires the kind of persistent, adaptive searching a single-pass system was never built to do, but it’s exactly what an attorney does when reconstructing what happened.
“Understanding what happened requires following all those threads,” Birnbaum said. Single-pass retrieval, she noted, works well for discrete factual questions, but it wasn’t designed for the layered, iterative fact-finding that most legal work demands.
Built on top of that retrieval approach are Specialized Skill Agents, each paired with a methodology designed for a specific kind of legal task rather than a generic prompt. A chronology skill doesn’t just pull documents containing dates; it identifies discrete events, places them in context, connects related evidence, flags inconsistencies, and clearly marks what the record supports and what it doesn’t. A fact narrative skill produces a structured memo, often running three to twelve pages, with every claim tied to a direct citation. A record check skill, built in direct response to a client request, lets a team upload a witness interview or deposition transcript and verify its claims against the underlying data.
None of these skills exist because someone decided that a chatbot needed more features. They grew out of specific client conversations: practitioners described a real bottleneck, and we met it with a methodology built to address it. That’s the throughline connecting Case Insight’s wide-angle classification to CaseBot’s close-up investigation: both exist because the underlying question was never “what can a generic AI assistant do here,” but “what does discovery require?”
Purpose-Built Tools Solve Problems General Models Can’t
That distinction, purpose-built versus general-purpose, is one Birnbaum returned to often, and for good reason: why not just use a tool like Claude or Harvey for this kind of work? The answer isn’t a knock on those tools. It’s a point about fit.
Discovery involves duplicate documents, custodian mapping, and data volumes that a general-purpose assistant simply isn’t built to process the way a tool designed specifically for eDiscovery is. A model that can’t distinguish duplicates from originals, or that only reads a fraction of a multimillion-document set, can’t reliably build a chronology or verify a witness statement against the full record, no matter how capable it is in other contexts.
“It’s not built to take a million documents and understand what we need in discovery to build a chronology,” Birnbaum said. “It’s going to give you duplicate documents; it’s only going to read X amount.”
That gap matters most when the output must survive scrutiny. Every fact that CaseBot surfaces carries a direct, clickable citation back to its source document. Quality control layers run behind the scenes specifically to catch hallucinations before a user ever sees them. And critically, CaseBot is built to find facts, not to offer legal judgment.
“I am the biggest proponent of GenAI that you’ll find,” Birnbaum said, “but you’re still the lawyer.”
Defensibility, in this framing, isn’t one fixed standard. It shifts depending on the task. Document review still leans on established metrics like recall, precision, and elusion testing, and Case Insight’s own validation work bears that out.
In one post-review test of 131,000 already-coded documents, Case Insight’s classifications held up against human review at recall levels of 88% to 93% and precision between 85% and 89%, an overall F-score of roughly 89%. Narrative work product headed toward trial preparation or an investigation needs a different kind of validation; one centered on sourcing and an honest account of the record’s limitations. What stays constant across both is the requirement that a team understands the tool it’s relying on: what’s behind it, what it’s built to do, and where its edges are. Standing behind AI-assisted work product means knowing those boundaries well enough to defend them.
The Work That’s Left When the Facts Are Found
Case Insight’s early read on a data set, CaseBot’s iterative digging, the skill agents built around specific legal tasks: each traces back to the same instinct: don’t ask where AI fits into discovery as it currently exists. Ask what discovery would look like if it were built today, with this technology already in hand. That question doesn’t have a single answer, and it isn’t finished being asked. The same thinking that produced Case Insight and CaseBot is already pushing further back into the process, into how teams hold context across a long matter and how they collect data in the first place.
What hasn’t moved is where the tools stop. CaseBot doesn’t decide what a fact means. It decides where a fact lives, how it connects to the rest of the record, and whether the record supports it. Case Insight doesn’t decide which flagged issue deserves a deposition or a motion. It decides what’s there to flag in the first place. Birnbaum put a name to that boundary during the conversation, describing what these tools ultimately let a team do as “investigate your record.” Investigation surfaces what’s true. It doesn’t decide what to do about it.
That’s the actual measure of what’s changed. Not that the work got automated, but that the work got sorted; in one matter, a 2.2-million-document investigation that would traditionally have taken months to work through instead surfaced its key findings in 48 hours, and the time that opens up goes toward the part no tool was ever built to do: building a theory, weighing a risk, deciding how to argue a case. Reimagining discovery around the technology that exists today doesn’t mean handing the process to that technology. It means finally giving the experts running it enough time to do the parts only they can.
More About Esther Birnbaum
Esther Birnbaum is the Executive Vice President of Data Intelligence at HaystackID, where she leads strategic initiatives integrating advanced AI technologies and data solutions across legal, compliance, and governance workflows. Esther brings extensive experience from both law firm and in-house environments. She began her career as an eDiscovery attorney at top-tier law firms, where she developed deep expertise in managing complex discovery processes for high-stakes litigation. She later served as Associate General Counsel at Interactive Brokers LLC, where she founded and scaled the company’s eDiscovery program from the ground up, pioneering AI-driven workflows that set new standards for operational efficiency in financial services litigation support. In her corporate role, she also spearheaded data intelligence initiatives across compliance investigations, regulatory response matters, and cross-functional governance programs, developing innovative approaches to leverage enterprise data for risk management and strategic decision-making. A recognized thought leader in the legal technology community, Esther is a sought-after speaker on the transformative applications of Generative AI in practice.
![[Podcast] HaystackID® in the EDRM Illumination Zone: Esther Birnbaum](https://haystackid.com/wp-content/uploads/2026/09/2026.09.22-HaystackID-EDRM-Illumination-Zone-Esther-Birnbaum-Web.jpg)
The podcast is available on your favorite listening app, including Spotify, Apple Podcasts, and Google Play. The podcast is also available on the EDRM website and is provided below for convenience.
Join HaystackID’s experts as they share actionable insights on today’s most material topics—from how GenAI is reshaping legal data strategies to the latest approaches in digital forensics. Explore our full library of EDRM Illumination Zone podcast episodes.
About the Electronic Discovery Reference Model
Empowering the global leaders of e-discovery, the Electronic Discovery Reference Model (EDRM) creates practical global resources to improve e-discovery, privacy, security, and information governance. Since 2005, EDRM has delivered leadership, standards, tools, guides, and test datasets to strengthen best practices throughout the world. EDRM has an international presence in 136 countries, spanning six continents. EDRM provides an innovative support infrastructure for individuals, law firms, corporations, and government organizations seeking to improve the practice and provision of data and legal discovery with 19 active projects. Learn more at EDRM.net.
About HaystackID®
HaystackID® solves complex data challenges related to legal, compliance, regulatory, and cyber requirements. Core offerings include Global Advisory, Cybersecurity, Core Intelligence AI™, and ReviewRight® Global Managed Review, supported by its unified CoreFlex™ service interface and eDiscovery AI® technology. Recognized globally by industry leaders, including Chambers, Gartner, IDC, and Legaltech News, HaystackID helps corporations and legal practices manage data gravity, where information demands action, and workflow gravity, where critical requirements demand coordinated expertise, delivering innovative solutions with a continual focus on security, privacy, and integrity. Learn more at HaystackID.com.
Assisted by GAI and LLM technologies.
Source: HaystackID
Advisory Note: As organizations face growing pressure to extract meaning from data before review begins, understanding a data set at the outset, rather than discovering it document by document, has become a strategic priority. HaystackID® Core Intelligence AI Case Insight™ delivers that early, whole-dataset understanding by classifying and contextualizing every document in a collection against everything else in it, without predefined categories, surfacing key people, events, and issues before a single document moves to manual review. Deployed across matters ranging from regulatory inquiries to criminal investigations, Case Insight has reduced review populations by more than 90% in some engagements, cut review time and cost by up to 72% and 60% respectively, and surfaced critical findings in enterprise-scale investigations within 48 hours. Validation testing against fully human-coded datasets has shown recall between 88% and 93% and precision between 85% and 89%, giving legal teams a defensible, quantifiable basis for relying on AI-driven early case assessment. To learn more about how HaystackID Core Intelligence AI Case Insight can help your organization manage discovery obligations more effectively, connect with HaystackID’s team of experts.