Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

A starting list. Check the licence on everything before you use it, and record where each file came from as you go — see Step 6.


Sanskrit

Digital text archives

Institutional projects

Large digitisation efforts are running, including a major collaboration in Chennai involving IIT Madras and Madras Sanskrit College, working through more than 110,000 rare manuscripts.

These projects usually publish their datasets and benchmarks openly. Check what they have released before you start collecting. They are doing the expensive part — digitisation — and you can build on it.

Existing models worth studying

Not for using directly, but for reading their data and tokenizer decisions:


Urdu


Multilingual collections that include both


These treat each language as one among many. That is exactly the gap described in the introduction, and exactly why their open data is useful to you while their models are not the last word.


Before you use any of it

  1. Check the licence. Free to read is not the same as free to redistribute or free to train on.

  2. Record the source. Every file, every time.

  3. Check the encoding. Older archives use a variety of transliteration schemes and legacy encodings.

  4. Deduplicate across sources — see Step 7. The same verse appears everywhere.