Has anyone tried building a large language deck from public datasets? I’m planning one for Urdu

I’m thinking about building a pretty comprehensive Urdu Anki deck by combining a few existing datasets like CLE, Wiktionary/Kaikki, Makhzan, Universal Dependencies, and Wikimedia/Lingua Libre for audio.

The basic plan is:

  • Keep lemmas and inflected forms separate
  • Use corpus frequency to rank/tag words
  • Pull definitions, POS, gender, IPA, etc. only when they’re actually available in the source
  • Leave fields blank if the data is missing or sources disagree rather than trying to guess
  • Add human-recorded audio wherever available
  • Keep track of where each piece of data came from

I’d automate the data collection/merging with scripts, but I don’t want AI generating any of the actual language content, definitions, pronunciations, etc.

Has anyone done something similar with multiple dictionary/corpus datasets? Mainly wondering what problems I’m likely to run into when merging everything together, especially duplicate words and conflicting data.

  1. CLE (Center for Language Engineering): Urdu-specific wordlists and corpus frequency data.
  2. Kaikki / Wiktionary: Machine-readable definitions, parts of speech, and IPA phonetic transcriptions.
  3. Makhzan: Digital Urdu dictionary corpus for classical and modern definitions.
  4. Universal Dependencies: Treebank data for grammatical tags and part-of-speech validation.
  5. Wikimedia / Lingua Libre: Open-license human audio recordings