← Back

The Source of Truth for Law

We just shipped a complete, versioned library of US law. It includes statutes, regulations, constitutions, executive orders, and agency guidance for all 50 states, D.C., and the federal government. This data now lives alongside our database...

We just shipped a complete, versioned library of US law. It includes statutes, regulations, constitutions, executive orders, and agency guidance for all 50 states, D.C., and the federal government. This data now lives alongside our database of case law, PACER integration, and the citator that binds it all together.

What We Shipped

There are two datasets that rival ours—Lexis’s and Westlaw’s. Those took a century to assemble, are behind expensive paywalls, and don’t integrate well into ChatGPT, Codex, Claude, Cowork, or Perplexity. Lexis and Westlaw almost never license their data to other companies. As a result, most attorneys overpay to access it, and the vast majority of Americans don’t even bother to ask their everyday legal questions because getting them answered is too expensive.

Our ambition has always been to change the status quo by building the source of truth for law. Every single user-facing feature we’ve built to date—the interfaces we’ve iterated on, the platforms we’ve plugged into, and the internal benchmarks we’ve spent hundreds of hours building—has been released in service of this goal.

Our ambition has always been to change the status quo by building the source of truth for law.

With today’s release, we’re one step closer to our goal:

  • Our new primary law dataset spans 192 collections and includes ~2.7 million statute sections, ~2 million regulation sections, ~16,000 constitutional provisions, ~283,000 agency guidance documents. We also carry historical versions where available.
  • Our case law corpus houses more than 14 million decisions from 3,397 courts—published and unpublished, from all federal courts, all state appellate courts, and a growing number of state trial courts—all connected by citation relationships between opinions.
  • Our PACER integration connects you to every single federal docket entry, with on-demand retrieval of docket sheets and filings—many of them free of charge. This means your agent can do things like cite check a brief against record cites, draft a statement of facts pulling from a complaint, and more.

On Quantity And Quality

We’re proud of the size of our dataset and that new data gets ingested and goes live within 7 hours of release. We’ll continue to build the largest commercially available corpus of legal data as we expand outside of the US to include coverage for Europe, the UK, Canada, and Australia. But when it comes to legal data, the numbers tell only half the story. The truth is that it’s easy to pad your stats: Count duplicate opinions, scrape 100+ year old inactive courts, fold in niche agency guidance documents from 50 years ago that are useful to a small handful of people, or count cases based on metadata instead of the actual opinion documents, and suddenly you can make grand claims about your enormous legal dataset. More and more companies are doing just this. Be cautious about them: a big dataset isn’t necessarily a useful one.

And so we’re even prouder that our offering is more than just a pile of data. It’s a high-quality and hyper-organized corpus engineered with our end-users (litigators and the people who build products for them) in mind. That means opinions are standardized, formatted, deduplicated, and enriched with all the metadata fields that matter. It means that they carry pin cites and treatments indicating whether they’re still good law. And it means that appellate court proceedings are linked to their trial court predecessors—and vice versa—even when the two don’t cite each other explicitly.

Our standardized opinion format in Midpage. For a side-by-side against a raw version, visit https://www.midpage.ai/data-quality.
The Midpage version on the left, compared to a raw opinion on the right. For more side-by-sides, visit https://www.midpage.ai/data-quality.

Our high bar also requires us to close gaps in our coverage even if they’re small. For example, we have detected and backfilled 352 gaps of greater than 4 weeks across 84 courts between 2020 and 2025 in sources that claim complete coverage. The end result of our diligence is that we have cases that you would otherwise only find on the most expensive platforms—and some users have even told us that they have surfaced opinions in Midpage that they could not find on Westlaw or Lexis.

For example, we have detected and backfilled 352 gaps of greater than 4 weeks across 84 courts between 2020 and 2025 in sources that many startups rely on.

Statutes and regulations are versioned and soon will be cross-linked. While statutes, regulations, and agency materials are (for the most part) available online, they’re hard to search across because they come from hundreds of sources and in a variety of formats. We standardize them all so they’re easy to navigate and readable for humans and agents alike. The example below comes from a 274-page scanned document.

Every week, a new entrant claims it solved hallucinations, created a zillion-document dataset, captured the entirety of the law in a knowledge graph, or outright fixed the American justice system. What separates Midpage is our singular obsession with building the best legal dataset on the planet. All the other stuff is downstream of that.

Build On It

Our users and our data customers notice the difference. Hundreds of thousands of visitors read cases on Midpage every month. More than 300 law firms use our products—from solos to midsized boutiques to big law firms. 5 multibillion-dollar organizations who service more than 100,000 attorneys each—including Perplexity, Noxtua, and other large legal tech players with valuations in the billions—trust us as their legal data supplier. Michael Showalter, from Showalter PLLC, told us that he’s “now able to draft devastating appellate briefs without even opening Westlaw” and called Midpage “a must-have for every litigator in 2026.”

The reason more attorneys and legal tech companies are choosing Midpage is simple: We have a comprehensive and high-quality dataset and we want you to build on top of it.