Loading market data...

EngramLab and Harvey Open Source 100M-Token Synthetic Legal Dataset

EngramLab and Harvey Open Source 100M-Token Synthetic Legal Dataset

EngramLab and Harvey have released a synthetic law firm dataset containing more than 100 million tokens, an open-source trove built to train legal AI without touching real client files. The announcement surfaced this week via Crypto Briefing.

A 100M-token training set

The dataset is synthetic, meaning it's generated rather than pulled from actual law firm records. That's a big deal for anyone building legal AI. Real legal data is expensive to license, slow to clean, and often locked behind confidentiality walls. This release gives developers a large, ready-to-use corpus without the usual legal headaches.

At 100 million tokens, it's not small. For context, that's enough text to train a serious language model from scratch, or fine-tune an existing one for legal tasks. The companies say the goal is to make training scalable and cost-effective, which could lower the barrier for startups and research labs that can't afford proprietary legal data.

The confidentiality angle

None of the data comes from real clients. That's the core pitch. By using synthetic examples, the dataset sidesteps the biggest risk in legal AI: leaking privileged information. A model trained on this won't have memorized a real person's case, a settlement amount, or a confidential memo. That's a meaningful step for a field where privacy isn't just a nice-to-have — it's a legal obligation.

It also means the dataset can be shared freely. No NDAs, no redaction checks, no worrying about which jurisdiction's rules apply. For researchers who've had to tiptoe around data-sharing agreements, that's a real convenience.

Open source, available now

The release is open source, so anyone can download it and start experimenting. That includes law schools, legal tech startups, and in-house AI teams at firms. The companies didn't say whether they plan to update the dataset regularly or add new versions, but the initial drop is live.

For now, the onus is on the community to put it to use. The dataset is out there, and the next step is seeing what people build with it.