mirror of
https://github.com/tiennm99/noitu.git
synced 2026-10-11 03:13:45 +00:00
feat(dict): build the corpus and word meanings from the Wikimedia viwiktionary dump
reader for both wikitext dialects; `meanings(word, ord, pos, gloss)` table; `meaning_count`/`words_with_meaning`/`source_pages` in meta, `source_rows` gone, builder_version 5; `--dump`/`--min-pages` replace `--kaikki`; attribution names the dump and the definition excerpts; 36,200 words, 96.9% with a meaning, every kaikki word kept.
This commit is contained in:
1 parent
557de1af94
commit
80f216d56b
25 files changed
+2338
-757
No files matched your search
@@ -137,9 +137,10 @@ jobs:
|
||||
grep -qx "$required" files.txt || { echo "missing from the image: $required"; exit 1; }
|
||||
done
|
||||
|
||||
# The upstream export must never reach the final image.
|
||||
if grep -q '\.jsonl$' files.txt; then
|
||||
echo "the upstream wordlist leaked into the image"
|
||||
# The upstream dump, compressed or not, must never reach the final
|
||||
# image.
|
||||
if grep -Eq '\.(bz2|xml)$' files.txt; then
|
||||
echo "the upstream dump leaked into the image"
|
||||
exit 1
|
||||
fi
|
||||
|
||||
|
||||
Reference in new issue
Block a user