feat(dict): build the corpus and word meanings from the Wikimedia viwiktionary dump

reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
This commit is contained in:
tiennm99 committed 2026-09-08 22:48:26 +07:00
1 parent 557de1af94
commit 80f216d56b
25 files changed
+2338 -757

No files matched your search

+4 -3
View File
@@ -137,9 +137,10 @@ jobs:
grep -qx "$required" files.txt || { echo "missing from the image: $required"; exit 1; }
done
# The upstream export must never reach the final image.
if grep -q '\.jsonl$' files.txt; then
echo "the upstream wordlist leaked into the image"
# The upstream dump, compressed or not, must never reach the final
# image.
if grep -Eq '\.(bz2|xml)$' files.txt; then
echo "the upstream dump leaked into the image"
exit 1
fi