Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Goal: build the biggest, cleanest Sanskrit corpus you can, and find out honestly how big that actually is.


Why this step matters

This is not a preparation step. For a low-resource language, this is the project. Everything else in this book is easier than this.

It is also the step where you discover the single most important fact about your project. Do not skip it, and do not guess the answer.


What you do

1. Gather from every source you can find

See the corpora appendix for a working list.

2. Record where every single file came from

Source, date, licence, and how you got it. Do this while you collect, not at release time.

Reconstructing this later is miserable, and without it you cannot legally release anything.

3. Count your tokens

Use your Step 4 tokenizer, not a general-purpose one. The number will be different, and yours is the one that matters.

4. Compare against what you need

A rough rule of thumb: a useful model wants somewhere between 5 and 20 tokens of training data for every parameter it has.

So:

Model sizeTokens wanted
100 million parameters0.5 to 2 billion
500 million parameters2.5 to 10 billion
1 billion parameters5 to 20 billion
7 billion parameters35 to 140 billion

5. Face the number

You will almost certainly find that all the clean Sanskrit text in the world adds up to far less than the smallest row in that table. Perhaps a few hundred million tokens, and much of that repeated across sources.

6. Do the same for Urdu

You will find much more text, and much of it much dirtier. A different problem, needing different solutions.

7. Write down your answer to one question

What do you actually want this model to do?

Not “understand Sanskrit”. Something you could test:

Your answer changes your data mix, your evaluation, and your architecture. A vague answer here produces a vague model later.


Where people usually get stuck

Assuming more data exists and that they just have not found it yet.

Do the count. Trust the count. The count is the plan.


You are ready to move on when

You have a token count for both languages, you believe it, and you have written one clear sentence saying what your model is for.