Discover Docket
← The DocketPERSPECTIVE

JILL vs. the generic chatbot: what changes when AI is built for litigation

The gap between a clever answer and a defensible one comes down to what the model is allowed to stand on.

CWChris WatersJune 9, 2026 · 8 min read

There is an honest version of why generic AI chatbots — ChatGPT, Gemini, and the rest — are popular among lawyers. They are good at language. They are conversational. They produce plausible-looking output on a wide range of legal questions in seconds. For tasks that don't require precision, they are genuinely helpful.

There is also an honest version of why lawyers who actually try to use generic chatbots for litigation work consistently encounter problems. The chatbot doesn't know the local rules of the court the matter is in. It cannot read the case file. It cannot validate its own citations. It cannot maintain an audit trail of what it did. It cannot recognize when it has been asked to do something it doesn't have the information to do well. The very generality that makes it useful for the easy cases is what fails it on the hard ones.

JILL — the AI built into Discover Docket — is a different kind of tool. Not in degree, in kind. This article walks through the substantive differences and what they mean for the work.

What "general-purpose" means and why it limits litigation work

A general-purpose chatbot is a foundation model accessed through a generic interface. The model has been trained on a large corpus of internet text, books, code, and other public data. It can answer questions about an enormous range of topics, write in dozens of styles, and produce structured output (lists, tables, code) on request.

The generality is the product. The chatbot is not designed for any specific use case. It is designed to be helpful across many use cases simultaneously.

For litigation work, generality is a structural limitation in three specific ways.

No jurisdictional awareness. A litigation matter sits in a specific court, in a specific jurisdiction, under specific procedural rules. The Federal Rules of Civil Procedure differ from the California Code of Civil Procedure. The local rules of the Eastern District of New York differ from those of the Southern District of New York. The standing orders of one judge in San Diego Superior Court differ from those of another judge in the same building. A general-purpose chatbot, asked to draft a motion to compel, will produce a motion to compel. It will not know what jurisdiction the matter is in unless told. Even when told, it will not have the specific local rules and standing orders in working memory at the level of detail those documents require.

No access to the case file. The chatbot can answer questions about general legal concepts. It cannot read your client's deposition transcript. It cannot review the opposing party's discovery responses. It cannot analyze the medical records that are central to your case theory. The case file lives outside the chatbot. The chatbot operates on the prompt, and the prompt can only contain what fits in the input. Realistic litigation work involves dozens of documents and hundreds of pages of material. The chatbot doesn't have any of it.

No validation. This is the Mata v. Avianca problem. The chatbot can produce a citation. The chatbot has no idea whether the citation is real. The chatbot cannot check, because the chatbot has no live connection to an authoritative case database. The chatbot cannot tell you which paragraphs of its output it is confident about and which it is uncertain about, because the chatbot was not built with that kind of metadata-aware output structure.

These three limitations are not flaws in the chatbot. They are features of what the chatbot was built to do, which is to be helpful across a very wide range of use cases. The chatbot was not built to be a litigation tool. Asking it to be one is asking it to operate outside its design.

What JILL is built around

JILL is not a chatbot. The interface is conversational because conversation is the right interaction model for many tasks, but the architecture underneath is fundamentally different.

JILL operates inside a litigation-specific data model. When a lawyer asks JILL to draft a motion, JILL knows:

  • The matter the request is associated with
  • The court the matter is venued in
  • The judge assigned to the matter
  • The procedural rules of that jurisdiction
  • The local rules of that court
  • The standing orders of that judge
  • The current state of the case file (filings, discovery, depositions, expert reports)
  • The deadline state of the matter (what's coming up, what just happened)
  • The matter team and ethical wall configuration

None of this is information the lawyer has to provide in the prompt. It is information that lives in the data model and is automatically available to JILL when she's working on the matter. The lawyer can ask "draft a motion to compel responses to Form Interrogatory 17.1" and JILL will produce a draft that knows it's a San Diego Superior Court Dept. 73 matter, applies the relevant CCP sections and CRC rules, follows the Dept. 73 standing order's formatting requirements, includes the meet-and-confer language Judge Wesley prefers, and incorporates the actual interrogatories from the matter file.

The lawyer can also ask JILL to do things that no chatbot can do at all: "Identify contradictions in the depositions of the plaintiff's three medical experts and produce an impeachment outline keyed to specific transcript pages." JILL can do this because the transcripts are in the matter file, the depositions are structured objects in the data model, and the workflow is supported by the platform architecture. The chatbot can't do this even in principle, because the chatbot doesn't have the transcripts.

The architectural difference is roughly the same as the difference between asking a brilliant generalist for help and asking a senior associate who has been working on the matter for six months. Both can be useful. They are useful for different tasks.

Validation as a structural feature, not an add-on

Beyond access to the matter, the second structural difference is validation.

Every citation JILL produces passes through validation against current authoritative case databases for the jurisdiction the matter sits in. If a case doesn't exist, the citation doesn't appear in JILL's output. If a case has been overruled or distinguished in a way that affects its applicability to the matter, the citation is flagged or annotated. The lawyer reading JILL's output sees citations that have been validated, with current good-law status, alongside the analysis they support.

The validation is not a separate verification step the lawyer remembers to run. It is the gate between JILL's reasoning and JILL's output. JILL cannot return a citation that has not been validated, because the architecture does not allow it.

Confidence scoring works the same way. Every paragraph of a JILL-drafted document carries a numeric confidence score that reflects the strength of the underlying authority, the consistency of supporting sources, and the freshness of the data. The lawyer always knows which paragraphs were drawn from strong, current, validated authority and which are operating on thinner ground. The information is built into the output, not delegated to the lawyer's intuition.

A general-purpose chatbot cannot do this in any meaningful way. The chatbot has no live database connection, no per-claim metadata structure, and no architectural concept of validation. Some chatbots have, in recent versions, added the ability to cite sources from web searches. This is closer to validation than nothing, but it is not the same as architectural verification against an authoritative legal database, and it does not produce the kind of paragraph-level confidence signal that defensible litigation work requires.

Audit and defensibility

The third structural difference is the audit trail.

When a lawyer uses a general-purpose chatbot, the only record of the interaction is whatever the lawyer remembers or screenshots. The provider's logs may exist, but they are not in the lawyer's control, are not signed, are not chained, and are not designed to be produced as evidence in a court proceeding.

When a lawyer uses JILL, every action — every prompt, every retrieval, every output, every confidence score — is recorded in the DDEAS audit log: append-only, cryptographically signed, chained into a tamper-evident sequence. The lawyer can produce a complete, evidentiary record of any work session on demand.

This matters because, in the post-Mata v. Avianca, post-Park v. Kim world, lawyers using AI need to be able to demonstrate that they used it responsibly. Generic chatbot use cannot be demonstrated; it can only be testified to. JILL use can be demonstrated, because the demonstration is built into the system.

What a general-purpose chatbot is still useful for

This is not an argument that lawyers should never use general-purpose chatbots. There are legitimate use cases for them, and most lawyers will find them useful for some part of the day.

General-purpose chatbots are useful for general legal questions, drafting non-substantive language, brainstorming arguments, writing client-facing emails (where the language matters more than legal precision), summarizing documents that the lawyer will personally verify, and learning about unfamiliar areas of law. They are useful for the kinds of tasks where the cost of getting it slightly wrong is low and the benefit of getting it quickly is high.

What they are not useful for is the substantive work of litigation — the drafting of court filings, the citation of authority, the analysis of evidence, the preparation of motions. These tasks require the validation, the matter file access, and the audit infrastructure that general-purpose chatbots structurally do not have.

The right mental model is the same as for any other professional tool. A general-purpose chatbot is what you ask a smart friend on a coffee break. JILL is what you have a senior associate do. The two are not substitutes. They are different tools for different work.

The market is going to bifurcate. On one side, general-purpose AI will keep improving — the consumer chatbots will be better at general legal questions in 2027 than in 2026. On the other side, specialized AI built for specific domains will deliver capabilities that general-purpose AI cannot, because the specialized AI operates inside data models, validation frameworks, and audit infrastructure that generality cannot replicate.

For law, the specialized side of the bifurcation is going to be the structurally consequential one. The lawyers who can defensibly use AI for substantive work — the work that produces filings, advice, and decisions — will be the ones using AI that was built for the work. The lawyers using general-purpose AI for the same work will be in the Mata position, except increasingly without the "I didn't know" defense.

We built JILL because that's the direction the work is going. Generic AI will be part of the day; it has its uses. The substantive litigation work, increasingly, will require the litigation-specific tool.

Learn about JILL →

Read the DDEAS framework →

Continue reading

AI SANCTIONS

Mata v. Avianca: What every lawyer needs to know about AI hallucinations

In May 2023, attorney Steven Schwartz filed a routine personal injury brief citing six federal cases that did not exist. The AI had fabricated all of them. What followed became the canonical cautionary tale of generative AI in legal practice.

Chris Waters · June 9, 2026 · 9 min read

PERSPECTIVE

The case for ethical AI in litigation

The legal profession is in the middle of the largest expansion of the competence duty in a generation. The question is not whether lawyers will use AI. The question is whether the lawyers using it will be the ones who can defend the work afterward.

Chris Waters · June 9, 2026 · 9 min read

Stop running your firm on fifteen tools.

Discover Docket replaces case management, research, AI, depositions, billing, and communications in one platform. California and Federal first, 52 jurisdictions on day one.