DVSGlobal

Testing embedders on Honduran case law

We tested eleven embedding models using citations from official Honduran Supreme Court rulings as ground truth. This article shares our findings and everything you need to reproduce the benchmark.

Start with the dataset

Honduras’ Supreme Court has an open portal with more than 20,000 official, public rulings. We saw an opportunity to use that data to test legal retrieval on real cases.

For this first experiment, we chose to limit our scope to the Honduran Traffic Act, passed in 2005 and amended six times. Small enough for a first benchmark, but complex enough to be useful.

Building the evaluation corpus

Only 23 of the 20,000-plus rulings explicitly cite articles from the Traffic Act. We excluded three*, leaving 20 cases.

For each case, we extracted the text surrounding the citation. We limited each excerpt so models with smaller context windows could be evaluated fairly.

We replaced people’s names with [PERSONA-X] placeholders. We also marked the article numbers and direct quotations from the law so they could be masked during evaluation. Otherwise, a model could simply match the quoted text instead of understanding the passage.

The complete dataset, including the original spans, is published on Hugging Face.

We also reconstructed every version of the Traffic Act from the official gazettes. This allowed us to evaluate each ruling against the law as it existed on the date of the decision.

That turned out to matter. The copy published by the National Police in 2025 still omits amendments from 2008, 2022 and 2025. For this benchmark, the gazette is the source of truth.

The benchmark

We tested eleven embedders, from 278 million to 8 billion parameters.

Hardware: NVIDIA RTX PRO 6000 Blackwell, 96 GB VRAM
Total runtime: 89 seconds

#CaseExpectedN-8BN-1BGemmaQ-4BQ-8BQ-0.6Be5BGEArcticGranitePara
1CC-41-121041º1º1º1º1º1º1º11º2º5º6º
2CC-153-1226, 271º1º1º1º1º1º2º1º1º12º—
3CP-106-141, 791º1º—4º3º——12º———
4CA-161-151, 98———12º—17º8º18º—3º—
5CP-230-13984º6º5º2º8º11º5º6º10º7º7º
6CP-259-146012º—————————12º
7CP-447-1440, 68, 992º8º5º2º5º2º2º2º16º6º9º
8CP-269-1649º9º18º——9º—12º8º7º7º
9AA-469-21191º1º1º1º1º2º1º3º11º3º—
10AP-1173-2119, 981º1º1º7º12º5º3º1º3º2º7º
11AC-785-211041º1º1º1º1º1º1º5º1º1º—
12CA-202-221171º1º1º1º1º1º1º1º1º3º2º
13CL-203-2145, 981º1º2º1º1º2º2º1º6º8º15º
14CL-234-22112——1º9º4º1º2º——9º2º
15AA-1614-22191º1º1º1º1º1º1º1º5º11º—
16AA-335-23264º6º14º6º16º6º6º8º5º4º11º
17AP-824-22113º3º5º1º1º3º2º4º2º3º8º
18AC-1050-221041º1º1º4º2º2º2º1º1º5º13º
19CC-223-19986º6º3º9º11º14º4º3º3º1º1º
20AA-639-2498, 1164º6º4º9º2º1º3º9º4º3º—
The rank at which each of the eleven models first retrieved one of the articles cited in the ruling. Blue = top 3, amber = top 10, gray = below the top 10, — = not retrieved. Columns ordered by MRR; scroll inside the table for the full matrix. Every case links to the official ruling.
Nemotron-3-8BNemotron-3-1BEmbeddingGemma 0.3BQwen3-4BQwen3-8BQwen3-0.6B
R@1R@5R@1000.250.50.751
Recall@k measures how often a cited article appears in the first 1, 5 or 10 results.
phone-class (≤ 1.1B params) needs a GPU (4B–8B)
0 0.2 0.4 0.6 Nemotron-3-8B 0.597 Nemotron-3-1B 0.562 EmbeddingGemma 0.3B 0.540 Qwen3-4B 0.511 Qwen3-8B 0.507 Qwen3-0.6B 0.502 e5-large-instruct 0.6B 0.470 BGE-M3 0.6B 0.417 Arctic-2 0.6B 0.343 Granite 0.3B 0.284 Paraphrase 0.3B 0.158
MRR measures how high the first correct article appears and gives us an overall model ranking.

What surprised us

The 8B Nemotron model won, but the more interesting result was how close the small models came.

More than the leaderboard, we wanted to share the process: court-authored ground truth, historical versions of the law and span-based masking. We think it can be reused in many other legaltech experiments.

Reproduce it

ollama pull embeddinggemma
pip install numpy
python run_eval.py

What’s next

Next we want to test rerankers, build larger corpora and run the full benchmark on-device.

This is the first study published by DVSGlobal Lab. More soon.

* We excluded three rulings: one cited a nonexistent article, one contained only procedural boilerplate and one cited nine articles together, which distorted any-hit scoring.