Testing embedders on Honduran case law
We tested eleven embedding models using citations from official Honduran Supreme Court rulings as ground truth. This article shares our findings and everything you need to reproduce the benchmark.
Start with the dataset
Honduras’ Supreme Court has an open portal with more than 20,000 official, public rulings. We saw an opportunity to use that data to test legal retrieval on real cases.
For this first experiment, we chose to limit our scope to the Honduran Traffic Act, passed in 2005 and amended six times. Small enough for a first benchmark, but complex enough to be useful.
Building the evaluation corpus
Only 23 of the 20,000-plus rulings explicitly cite articles from the Traffic Act. We excluded three*, leaving 20 cases.
For each case, we extracted the text surrounding the citation. We limited each excerpt so models with smaller context windows could be evaluated fairly.
We replaced people’s names with [PERSONA-X] placeholders. We also marked
the article numbers and direct quotations from the law so they could be masked during
evaluation. Otherwise, a model could simply match the quoted text instead of
understanding the passage.
The complete dataset, including the original spans, is published on Hugging Face.
We also reconstructed every version of the Traffic Act from the official gazettes. This allowed us to evaluate each ruling against the law as it existed on the date of the decision.
That turned out to matter. The copy published by the National Police in 2025 still omits amendments from 2008, 2022 and 2025. For this benchmark, the gazette is the source of truth.
The benchmark
We tested eleven embedders, from 278 million to 8 billion parameters.
Hardware: NVIDIA RTX PRO 6000 Blackwell, 96 GB VRAM
Total runtime: 89 seconds
| # | Case | Expected | N-8B | N-1B | Gemma | Q-4B | Q-8B | Q-0.6B | e5 | BGE | Arctic | Granite | Para |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | CC-41-12 | 104 | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 11º | 2º | 5º | 6º |
| 2 | CC-153-12 | 26, 27 | 1º | 1º | 1º | 1º | 1º | 1º | 2º | 1º | 1º | 12º | — |
| 3 | CP-106-14 | 1, 79 | 1º | 1º | — | 4º | 3º | — | — | 12º | — | — | — |
| 4 | CA-161-15 | 1, 98 | — | — | — | 12º | — | 17º | 8º | 18º | — | 3º | — |
| 5 | CP-230-13 | 98 | 4º | 6º | 5º | 2º | 8º | 11º | 5º | 6º | 10º | 7º | 7º |
| 6 | CP-259-14 | 60 | 12º | — | — | — | — | — | — | — | — | — | 12º |
| 7 | CP-447-14 | 40, 68, 99 | 2º | 8º | 5º | 2º | 5º | 2º | 2º | 2º | 16º | 6º | 9º |
| 8 | CP-269-16 | 4 | 9º | 9º | 18º | — | — | 9º | — | 12º | 8º | 7º | 7º |
| 9 | AA-469-21 | 19 | 1º | 1º | 1º | 1º | 1º | 2º | 1º | 3º | 11º | 3º | — |
| 10 | AP-1173-21 | 19, 98 | 1º | 1º | 1º | 7º | 12º | 5º | 3º | 1º | 3º | 2º | 7º |
| 11 | AC-785-21 | 104 | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 5º | 1º | 1º | — |
| 12 | CA-202-22 | 117 | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 3º | 2º |
| 13 | CL-203-21 | 45, 98 | 1º | 1º | 2º | 1º | 1º | 2º | 2º | 1º | 6º | 8º | 15º |
| 14 | CL-234-22 | 112 | — | — | 1º | 9º | 4º | 1º | 2º | — | — | 9º | 2º |
| 15 | AA-1614-22 | 19 | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 1º | 5º | 11º | — |
| 16 | AA-335-23 | 26 | 4º | 6º | 14º | 6º | 16º | 6º | 6º | 8º | 5º | 4º | 11º |
| 17 | AP-824-22 | 11 | 3º | 3º | 5º | 1º | 1º | 3º | 2º | 4º | 2º | 3º | 8º |
| 18 | AC-1050-22 | 104 | 1º | 1º | 1º | 4º | 2º | 2º | 2º | 1º | 1º | 5º | 13º |
| 19 | CC-223-19 | 98 | 6º | 6º | 3º | 9º | 11º | 14º | 4º | 3º | 3º | 1º | 1º |
| 20 | AA-639-24 | 98, 116 | 4º | 6º | 4º | 9º | 2º | 1º | 3º | 9º | 4º | 3º | — |
What surprised us
The 8B Nemotron model won, but the more interesting result was how close the small models came.
- Nemotron-3-1B outperformed every other model family, including the 4B and 8B models from other families.
- The three Qwen3 models finished within 0.009 MRR of each other despite a 13× difference in size.
- EmbeddingGemma, at only 300M parameters, placed third overall and is small enough to run on a phone.
More than the leaderboard, we wanted to share the process: court-authored ground truth, historical versions of the law and span-based masking. We think it can be reused in many other legaltech experiments.
Reproduce it
ollama pull embeddinggemma
pip install numpy
python run_eval.py
What’s next
Next we want to test rerankers, build larger corpora and run the full benchmark on-device.
This is the first study published by DVSGlobal Lab. More soon.
* We excluded three rulings: one cited a nonexistent article, one contained only procedural boilerplate and one cited nine articles together, which distorted any-hit scoring.