I have to apply LDA (Latent Dirichlet Allocation) to get the possible topics from a data base of 20,000 documents that I collected.

How can I use these documents rather than the other corpus available like the Brown Corpus or English Wikipedia as training corpus ?

You can refer this page.

Edit
Report