Data sources and licenses
LexiBrain is built on open dictionaries, word lists and sentence collections made by many people. Thank you to all of them. The table lists what each source is used for and its license.
Share-alike: several sources are licensed CC BY-SA. The word data LexiBrain derives from them (the data/ files: meanings, readings, etymologies, example choices) is therefore also available under CC BY-SA 4.0, with the attributions below. The app’s own code and design are not covered by these licenses.
Dictionaries and word lists
| Source | Used for | License |
|---|---|---|
| Wiktionary (English edition), extracted by kaikki.org / wiktextract (Tatu Ylonen) | Etymologies and cognates (the Origins view), checks of gender, part of speech and meanings for Spanish, French and German | CC BY-SA 4.0 / GFDL |
| WikDict (Karl Bartel), bilingual dictionaries built from many language editions of Wiktionary via DBnary | Human translations used as references when writing meanings in each interface language, and to measure how often our meanings agree with a dictionary | CC BY-SA 3.0 |
| JMdict and KANJIDIC2, Electronic Dictionary Research and Development Group (EDRDG), via jmdict-simplified | Japanese words, readings, meanings (English, German, Russian, Spanish, French); kanji readings and meanings | CC BY-SA 4.0 |
| Basic Korean Dictionary (한국어기초사전), National Institute of Korean Language (국립국어원) | Korean words, vocabulary levels, hanja origins, pronunciations, and meanings in English, Japanese, French, Spanish, Arabic, Indonesian, Russian and Chinese | CC BY-SA 2.0 KR |
| CC-CEDICT (MDBG) | Checking Chinese pinyin and meanings; new Chinese words | CC BY-SA 4.0 |
| Make Me a Hanzi (Shaunak Kishore) | How Chinese characters are built (components, radicals) | Arphic Public License / LGPL |
| Unihan database, Unicode Consortium | Radicals, stroke counts and readings of CJK characters; the “readings across East Asia” table (Mandarin, Cantonese, Japanese on’yomi, Korean, Vietnamese, fanqie, and Tang-dynasty reconstructions by Hugh M. Stimson) | Unicode License v3 |
| HSK 3.0 word list (ivankra), from the Chinese Proficiency Grading Standards for International Chinese Language Education (2021) | HSK levels in the Chinese brain | MIT |
| Open Anki JLPT decks (Jamie Sinclair), based on Jonathan Waller’s JLPT lists | JLPT levels in the Japanese brain (unofficial: official lists have not been published since 2010) | MIT / CC BY |
| OpenCC | Simplified ↔ traditional characters and Japanese shinjitai, to match Chinese, Japanese and Korean words written with the same characters | Apache 2.0 |
| Glottolog 5.3 (Hammarström, Forkel, Haspelmath, Bank) | Language family classification | CC BY 4.0 |
| ECDICT (skywind3000) | English–Chinese meanings, word forms, roots and affixes | MIT |
| WordNet 3.0 (Princeton University) and the Open Multilingual Wordnet | English meanings, synonyms and antonyms; checks of translated meanings | WordNet License; CC BY and others (per wordnet) |
| CEFR-J Wordlist and Octanove Vocabulary Profile C1/C2 | CEFR levels of English words | Free for research and commercial use with attribution |
| wordfreq (Robyn Speer) | Word frequencies (which words are common) | Code Apache 2.0; data CC BY-SA 4.0 |
| GloVe (Stanford NLP) | Word vectors that place related meanings near each other in the brain | Public Domain Dedication and License (PDDL) |
| IE-CoR 1.2, Indo-European Cognate Relationships (Heggarty, Anderson, Scarborough et al.; CLDF edition) | “One meaning, many languages”: forms in 160 Indo-European languages with expert cognate judgements and reconstructed roots | CC BY 4.0 |
| NorthEuraLex 0.9 (Dellert et al.; CLDF edition) | “One meaning, many languages”: forms and IPA in 107 languages of northern Eurasia, including Mandarin, Japanese and Korean | CC BY 4.0 (CLDF edition; original CC BY-SA 4.0) |
| New General Service List project (Browne, Culligan & Phillips): NGSL 1.2, NAWL 1.2, TSL 1.2, BSL 1.20 | English word lists: general core, academic, TOEIC and business | CC BY-SA 4.0 |
Example sentences
Example sentences and their human translations come from Tatoeba, a collection of sentences written and translated by volunteers, licensed CC BY 2.0 FR. Each example links to its page on Tatoeba, where its authors are listed. Some Korean examples come from the Basic Korean Dictionary (CC BY-SA 2.0 KR); their English translations were drafted by AI.
Software
3D rendering: three.js (MIT). Fonts: Inter, Newsreader and Noto Serif SC (SIL Open Font License 1.1); icons: Material Symbols (Apache 2.0).
Usage statistics (countries): IP Geolocation by DB-IP, IP to Country Lite (CC BY 4.0). The IP address is only looked up on our server and is not stored.
AI-drafted content
Some content was drafted with AI and then checked against the sources above: meanings in languages a dictionary does not cover, some Chinese character explanations, and the placement of words into the 25 regions. Where a dictionary and a draft disagree, the dictionary wins. Example sentences that have no human translation in your language are translated by AI. These checks were also done by AI; no native-speaker review has been completed yet. If you find a mistake, please tell us.
数据来源与许可(中文)
词脑建立在许多人做的开放词典、词表和例句库之上,谢谢他们。下面列出每一项的用途和许可。
相同方式共享:其中几项是 CC BY-SA 许可,所以词脑从它们整理出的词汇数据(data/ 下的释义、读音、词源、例句选择等)也按 CC BY-SA 4.0 提供,并保留下面的署名。应用本身的代码和设计不在这些许可范围内。
- Wiktionary(维基词典英文版),由 kaikki.org / wiktextract 抽取:词源和同源词(语源脑),西、法、德语的性、词性和释义核对。CC BY-SA 4.0。
- WikDict(从多种语言版 Wiktionary 抽取的双语词典,经 DBnary):编写各界面语言释义时参照的人工对译,也用来统计我们的释义和词典的一致率。CC BY-SA 3.0。
- JMdict、KANJIDIC2(EDRDG,经 jmdict-simplified):日语词、读音、释义;汉字的读音和意思。CC BY-SA 4.0。
- 《韩国语基础词典》(国立国语院):韩语词、词汇等级、汉字原词、发音,以及英、日、法、西、阿、印尼、俄、中文对译。CC BY-SA 2.0 KR。
- CC-CEDICT(MDBG):核对中文拼音和释义,补充中文词。CC BY-SA 4.0。
- Make Me a Hanzi:汉字的构成(部件、部首)。Arphic 公共许可 / LGPL。
- Unihan 数据库(Unicode 联盟):部首、笔画、读音;"各语言读音"对照表(普通话、粤语、日语音读、韩语、越南语、反切,以及 Stimson 构拟的唐代读音)。Unicode 许可 v3。
- HSK 3.0 词表(ivankra 整理自《国际中文教育中文水平等级标准》):中文脑的 HSK 等级。MIT。
- JLPT 词表(open-anki-jlpt-decks,基于 Jonathan Waller 的词表):日语脑的 JLPT 级别(非官方:2010 年后官方不再公布词表)。MIT / CC BY。
- OpenCC:简繁转换、日本新字体,用来把写法相同的中日韩汉字词对上。Apache 2.0。
- Glottolog 5.3:语系分类。CC BY 4.0。
- ECDICT:英汉释义、词形、词根词缀。MIT。
- WordNet 3.0 和 开放多语 WordNet:英文释义、近义反义;译文核对。
- CEFR-J 词表、Octanove C1/C2 词表:英语词的 CEFR 等级。
- wordfreq:词频。数据 CC BY-SA 4.0。
- GloVe(斯坦福):词向量,让意思相近的词在脑里挨在一起。PDDL。
- IE-CoR 1.2(印欧语同源关系数据库,Heggarty、Anderson、Scarborough 等):"一词多语"里 160 种印欧语言的说法、专家判定的同源组和构拟词根。CC BY 4.0。
- NorthEuraLex 0.9(Dellert 等):"一词多语"里欧亚北部 107 种语言(含普通话、日语、韩语)的说法和国际音标。CC BY 4.0(CLDF 版;原版 CC BY-SA 4.0)。
- New General Service List 项目(Browne、Culligan、Phillips)的 NGSL、NAWL、TSL、BSL:英语通用核心、学术、TOEIC、商务词表。CC BY-SA 4.0。
例句和它们的人工译文来自 Tatoeba(志愿者写作和翻译的句子库,CC BY 2.0 FR),每个例句都链接到 Tatoeba 上的原句页面,那里有作者信息。部分韩语例句来自《韩国语基础词典》,它们的英文译文由 AI 起草。
软件:three.js(MIT);字体 Inter、Newsreader、Noto Serif SC(SIL OFL 1.1);图标 Material Symbols(Apache 2.0)。
使用统计里的国家:IP Geolocation by DB-IP,IP to Country Lite(CC BY 4.0)。IP 地址只在我们的服务器上查一下,不保存。
AI 起草的内容:词典没有覆盖的语言的释义、部分汉字说明、词在 25 个领域里的归类,是 AI 起草后再用上面的数据核对的;词典和草稿不一致时以词典为准。例句在你的语言里没有人工译文时,由 AI 翻译。这些核对也是 AI 做的,还没有完成母语者审校。发现错误请告诉我们。